HeadlinesBriefing favicon HeadlinesBriefing.com

Inside DeepSeek: Reverse Engineering AI Assistant

Hacker News •
×

I run manish.sh, writing about AI tools and how LLMs behave when pushed. This piece is part of the Inside LLMs series, where I interview a chat model about its thinking and then compare its answers to published research. The DeepSeek interview follows the earlier Kimi K2.6 entry.

When I asked what it actually knows, DeepSeek split its reply into observation, inference, and guess — a framing I hadn't expected. It identified itself as the latest version with a May 2025 knowledge cutoff, said it was not labelled as a reasoning model, and noted a 21 July 2026 export date. It also mentioned 256 experts and 671B/37B parameters, and listed innovations such as DeepSeek, MLA, Mo E, SFT, and RLHF.

The interview reveals hard limits: the model cannot see its own weights, routing, or attention maps, and its explanations are generated text, not direct inspection. I treat the post like a documentary — question, short reply, my reaction, diagram, takeaway, and hook — while reminding readers that chat output is not proof. For architecture numbers, check the arXiv paper; use the chat for behavior and prompting intuition.