HeadlinesBriefing favicon HeadlinesBriefing.com

AI Labs Aren't Pelicanmaxxing, Study Finds

Hacker News •
×

A recent experiment investigated whether major AI labs are "pelicanmaxxing" – optimizing their models for a specific, informal benchmark: generating an SVG of a pelican riding a bicycle. This benchmark, popularized by Simon Willison, has become a notable informal test for new LLM releases.

The experiment involved generating 1,008 SVGs across seven frontier models, including GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and Deep Seek V4 Pro. Prompts varied animals and vehicles, with the pelican-bicycle combination being one of 48 tested pairs.

Analysis, using an LLM judge and Claude Fable 5, revealed no significant evidence of pelicanmaxxing. Pelicans were ranked sixth out of eight animals in terms of drawing quality, and bicycles ranked second to last out of six vehicles. Even when adjusting for inherent prompt difficulty, the specific "pelican on a bicycle" combination did not show a statistically significant boost across the tested models, suggesting labs are not specifically training for this popular benchmark.