HeadlinesBriefing favicon HeadlinesBriefing.com

Qwen3-Omni-Flash-2025-12-01: Next-Gen Multimodal Model

Hacker News •
×

Qwen3-Omni-Flash-2025-12-01 is a next-generation native multimodal large model with 30B parameters and 3B active parameters, succeeding the previous 7B omni model. It is positioned as a successor to Qwen3-Omni-30B-A3B-Instruct, though its weights are not currently available in open-source frameworks, making it difficult to use despite claims of strong performance. The model includes a 650M Audio Encoder, 540M Vision Encoder, 30B-A3B LLM, 3B-A0.3B Audio LLM, and 80M Transformer/200M Conv Net for audio token-to-waveform conversion.

A reasoning version exists that may pronounce thinking tokens during voice interaction, adding an amusing interactive element. While some speculate it may be a closed-weight model-as-a-service due to the 'Flash' designation, others note it may be an omni version of Qwen3-235B-A22B, given its benchmark performance against larger models. The model’s context window may extend to 200K+ tokens, but weights remain unavailable on Hugging Face or ModelScope as of the article’s writing.

Real-time conversation support is confirmed, though local hardware implementation remains unverified by the community.