HeadlinesBriefing favicon HeadlinesBriefing.com

Qwen3.8-Flash-Next: Alibaba Previews Qwen4 Architecture

Hacker News •
×

Alibaba has released Qwen3.8-Flash-Next, an open-weight multimodal mixture-of-experts model built on the Qwen4 architecture, before any Qwen4 model exists. The official Hugging Face repository went live on August 24, two days ahead of the scheduled Model Scope release on August 26. The model card confirms it as "an experimental preview of the architecture that will underpin Qwen4."

Key architectural changes include Gated Delta Network (GDN) and Qwen Sparse Attention (QSA), within a 48-layer hybrid of 36 linear-attention and 12 full-attention layers. The model has 125B total parameters with 6B activated per token, 512 experts, a separate 51B n-gram embedding table, and 262,144-token native context, extensible to 1,000,000. The weights are non-gated and fully spec'd.

Alibaba frames this as a technical commitment: preview the hard part first, so runtimes and applications are ready for the Qwen4 family. However, benchmark scores remain unconfirmed and independently unreproduced. The formal Model Scope drop is still scheduled for 2026-08-26 23:00 (UTC+08:00), but the weights have already beaten the timer.