Hacker News
Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
The experiment reran reasoning-prefill tests using GPT-5.5 Pro as the teacher, inserting the first 1% of its reasoning into each target model’s response and measuring overlap of the teacher’s visible answer within the first 100 tokens. Qwen’s overlap increased by +18.18 percentage points toward GPT-5.5 Pro, especially on private synthetic puzzles, while Kimi K3 showed the highest overall overlap (31.11% without prefill, 35.65% with prefill, a +4.54-point gain).