HeadlinesBriefing favicon HeadlinesBriefing.com

Real-SWE: निजी Codebases पर AI का प्रदर्शन

Hacker News •
×

सितंबर 2026 Real-SWE benchmarking निजी, वास्तविक-दुनिया के enterprise codebases पर frontier AI models का मूल्यांकन करता है। प्रत्येक task एक वास्तविक company से licensed निजी production codebase से आता है, जो engineers को proprietary systems और business-critical problems प्रस्तुत करता है। परिणाम अग्रणी models के बीच महत्वपूर्ण performance gaps को उजागर करते हैं, जिनमें Claude Code ने 38.8% की सबसे उच्च resolution rate प्राप्त की, उसके बाद GPT-6 Astra Codex CLI 33.8% और Gemini 3.8 Flash 31.2% पर रहे। यह evaluation model-and-harness combinations को मापता है, जिसमें agents से proprietary conventions को navigate करने और multiple services में business operations को प्रभावित करने वाले changes लागू करने की आवश्यकता होती है। Real company tasks में specific context की मांग होती है, जैसे सही billing का business rules और external tax services पर निर्भर होना। Tasks में invoice billing को ठीक करना शामिल है ताकि हर business सही tax charge करे, जबकि exempt customers untaxed ही रहें। यह benchmark इस बात को उजागर करता है कि expert-generated या synthetic tasks वास्तविक engineer work से अलग होते हैं, क्योंकि underlying coding artifact और instruction specificity दोनों जटिलता जोड़ते हैं, जो आज के frontier models के लिए चुनौतीपूर्ण है।