HeadlinesBriefing favicon HeadlinesBriefing.com

Maple-Preview: 20B MoE LLM runs fast on iPhone

Hacker News •
×

We introduce Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM that achieves state-of-the-art performance in its weight class. It can solve IMO-level problems and runs at over 200 tokens/s on a Mac mini M4, significantly faster than models like Gemma 4 and Qwen3.5. On an iPhone, Maple-Preview achieves 127 tokens/s, outperforming Bonsai 27B by a large margin.

Our vision is a shift towards always-active, on-device AI assistants. We believe ultra-low precision is crucial for this, enabling models to be trained and run on everyday devices. Maple-Preview is natively trained at low precision, demonstrating that efficiency does not require compromise. Its hardware-aware architecture was optimized for inference speed.

Maple-Preview sets a new benchmark for memory-to-performance and speed-to-performance. Beyond raw reasoning, its light footprint enables on-device adaptation. In a demo, it learned a user's vegan preference and adapted its recommendations, a capability that existing models like Claude Sonnet 5 struggled with. Future work includes scaling agentic training and on-device learning methods.