HeadlinesBriefing favicon HeadlinesBriefing.com

AI Race Awkward: Chinese Labs Lead Optimization

Hacker News •
×

If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse. Not a week goes by when Anthropic doesn’t release another article on how the Chinese are distilling their models. But it’s clear that the days of mindless distilling are over. The new game in town is adopting Chinese labs’ advances. Unlike the Western companies, the Chinese are pretty much giving away their recipes.

The latest one shamelessly copied without acknowledgement is the breakthrough in KV cache optimizations that Deep Seek has generously shared with the world. It dropped the KV cache footprint for certain use cases by a factor of roughly 437x compared with Deep Seek-V1. The latest Deep Seek-V4.1-Flash pushes it even further, bringing the global KV cache down to 890 bytes per token. For serving long-context models, one of the largest costs is the VRAM needed to hold this cache in GPU memory.

All this must mean the Western AI companies are now extremely inference-margin positive. The constraints on access to advanced GPUs forced Chinese labs to make performance optimization a number one goal. For once both Anthropic and OpenAI released models that are basically top-tier and are using these optimizations. Hence the silent releases for both Claude Opus 5.5 and GPT-6.1 Sol. The cache read costs dropped sharply: Opus 5.5 cut cache-read pricing by 60% versus Opus 5, while GPT-6.1 Sol cut it by 80% versus GPT-5.6 Sol’s late-July pricing.