HeadlinesBriefing favicon HeadlinesBriefing.com

big-pickle Scores 50.8% on SWE Atlas Benchmark

Hacker News •
×

The free stealth model big-pickle on Open Code Zen achieved a 50.8% Task Resolve Rate (63/124) on Scale AI’s SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold, run on 2026-08-11 with official harnesses and judge model. In the official leaderboard (updated 2026-07-28), big-pickle outperforms all other Mini‑SWE‑Agent entries and tops Codex‑scaffold GPT models, with only the two Claude models on their native scaffold scoring higher.

Language‑specific performance shows Python at 55.2%, Go at 50.0%, and C# at 38.5%. Category breakdowns reveal Root‑cause analysis at 74.9% and Code Onboarding at 60.7%. The evaluation followed Scale’s protocol: 124 tasks, mini‑swe‑agent 2.4.6 scaffold, and the claude‑opus‑4‑5‑20251101 judge. Resources were reduced to 4 CPU / 8 GB, yet no timeouts or OOM kills occurred. The run used $0 model cost (free during stealth) and modest Modal compute (~$70) and Anthropic judging (~$25).

Full verifier logs are available for independent audit, and the run is reproducible using the provided configs. The model’s identity is unconfirmed, likely served by Deep Seek infrastructure, and may change without notice.