A reproducible benchmark on GitHub shows that prompting LLMs to write in cablese—the terse, lowercase style of 19th-century telegraph messages—yields 40–49% token savings without losing downstream comprehension. Using a 50-passage, ~1,300-question benchmark, the study found that models like gemma-4-31b, qwen3.8-27b, and GLM-5.3-Flash all achieve recovery ratios between 0.99 and 1.10 when reading compressed records versus plaintext. The key is a single lowercase instruction: without it, models default to ALL CAPS, costing 14–19 accuracy points.
Notably, gpt-5-mini cannot disable reasoning, doubling its bill. The technique works across four model families, including those never trained on cablese, proving the capability is already in their weights. Compression is most effective when applied after content is settled, not during composition, avoiding a "register tax." The result is a verified, zero-cost optimization that nearly doubles effective context or halves output bills.
Source: Hacker News · Summarized by HeadlinesBriefing