HeadlinesBriefing favicon HeadlinesBriefing.com

Qwen 3.8 27B on Cerebras at 1500 tokens/s

Hacker News •
×

Qwen 3.8 27B is now available on Cerebras public endpoints, delivering 1500 tokens/s throughput. These endpoints are accessible via free trial and pay-as-you-go tiers, subject to rate limits and pricing. For dedicated capacity and production SLAs, Cerebras offers Dedicated Endpoints.

All models on Cerebras public endpoints are original, unpruned versions. The platform does not host pruned models on its shared API. While Cerebras researches pruning techniques like REAP (Router-weighted Expert Activation Pruning), these pruned models are shared on Hugging Face for research but not served through the API.

Cerebras uses selective weight-only quantization during storage to preserve quality. Weights are stored in partial 16-bit/8-bit/4-bit formats, with sensitive layers kept at full precision and dequantized on the fly. Activations, attention, and KV cache remain in full precision and unquantized.

Cerebras commits to serving original models without architectural changes. Future compression techniques would be offered as separate endpoints with specific names. REAP pruned models are available on Hugging Face for research, but not in production.