HeadlinesBriefing favicon HeadlinesBriefing.com

نتائج اختبار معيار SVG للبجع لـ Claude Fable 5.1

Hacker News •
×

Anthropic released Claude Fable 5.1 on September 1st, 2026, claiming a new standard for coding and knowledge work with a 52.6% score on Terminal-Bench-Science 0.1, up from 24.7% for Fable 5 and 22.4% for GPT-5.6 Sol. The author tested Fable 5.1's five reasoning levels (low, medium, high, xhigh, max) using the pelican SVG benchmark. At low and medium settings, the model appeared to skip reasoning entirely, producing similar outputs in ~23 seconds costing ~10 cents.

High effort added minimal reasoning (29.6 seconds, 13 cents). Xhigh effort produced dramatically different results: 36,767 tokens, 7 minutes 51 seconds, $1.83, with detailed reasoning about pelican anatomy and bike mechanics. Max effort delivered the best pelican from any Anthropic model: 65,927 tokens, 13 minutes 54 seconds, $3.30, featuring a tasteful background, proper leg positioning, a blue hat, and fish basket.

The reasoning trace revealed careful deliberation about helmet vs. crest, pedal visibility, and proportion adjustments. While impressive, the author notes it still lacks the flair of Gemini 3.7 Flash.