HeadlinesBriefing favicon HeadlinesBriefing.com

Real-SWE Benchmark: AI Performance on Private Codebases

Hacker News •
×

September 2026 Real-SWE benchmarking evaluates frontier AI models on private, real-world enterprise codebases. Each task originates from a private production codebase licensed from a real company, presenting engineers with proprietary systems and business-critical problems. Results reveal significant performance gaps between leading models, with Claude Code achieving the highest resolution rate at 38.8%, followed by GPT-6 Astra Codex CLI at 33.8% and Gemini 3.8 Flash at 31.2%.

The evaluation measures model-and-harness combinations, requiring agents to navigate proprietary conventions and implement changes affecting business operations across multiple services. Real company tasks demand specific context, such as correct billing depending on business rules and external tax services. Tasks involve fixing invoice billing so each business charges the right tax while exempt customers remain untaxed.

The benchmark highlights that expert-generated or synthetic tasks differ from actual engineer work, as both the underlying coding artifact and instruction specificity add complexity challenging today's frontier models.