HeadlinesBriefing favicon HeadlinesBriefing.com

Reasoning Prefill Study: GPT-5.5 Pro Influence

Hacker News •
×

This v1.1 reruns a reasoning-prefill experiment using GPT-5.5 Pro as the teacher. For each of 45 problems (15 STEM, 15 non-STEM, 15 synthetic puzzles), two responses were generated per target model: one unprefilled and one prefilled with the first 1% of GPT-5.5 Pro's hidden reasoning. The visible answer remained freely generated. Scores measured how much of the teacher's visible answer appeared in the first 100 tokens of the target's answer, using mean unigram, bigram, and trigram recall.

Results show Deep Seek V4 Flash had a delta of -1.17 percentage points, dropping from 27% to 26.13%. Inkling rose from 19.99% to 20.45% (+0.46 pp). Kimi K3 increased from 31.11% to 35.65% (+4.54 pp). Qwen3.8 A95B showed the largest shift, from 16.79% to 34.97% (+18.18 pp).

By category, Qwen's improvement was strongest in synthetic puzzles (+26.99 pp, from 19.26% to 46.24%), followed by non-STEM (+12.80 pp) and STEM (+14.75 pp). Overall, Qwen barely moved toward Opus 4.8 in an earlier experiment but moved significantly toward GPT-5.5 Pro here, suggesting it may have learned from GPT-5.5 Pro or a closely related model. Kimi K3 had the highest raw overlap with GPT-5.5 Pro both without and with prefill, though its prefill gain was modest.