HeadlinesBriefing favicon HeadlinesBriefing.com

DeepSeek-V4.1-Flash: Smaller, Faster AI Model Launches

Hacker News •
×

DeepSeek has introduced DeepSeek-V4.1-Flash, the smallest model in its new architecture family, featuring native visual understanding and designed for greater capability, faster inference, and higher throughput. The model uses an asymmetric architecture with 552B-parameter MoE and a new Causal Encoder–Decoder design, requiring only 8B active parameters for input and 16B for output. Compared to the previous generation, V4.1-Flash reduces KV cache needs to 1/4 HBM and 1/8 SSD storage, significantly lowering agent costs by compressing cache-hit charges.

The model is now live on the DeepSeek API with native multimodal support; users should set their model to deepseek-flash. V4-Flash and V4-Flash-Vision-Exp are retired, with temporary routing to V4.1-Flash for compatibility. DeepSeek reports lower API prices due to improved efficiency, passing savings to users, and maintains peak/off-peak pricing with off-peak rates at 50% of peak.

The company is collaborating with the open-source community on inference support and exploring deployment options, inviting inquiries for large-scale setups involving 2,000 GPUs and storage clusters. The model is available on Hugging Face as deepseek-ai/DeepSeek-V4.1-Flash.