← October 6, 2026 briefing

P1 Research Paper arXiv

Study: fixing a response's opening tokens brings base models close to RL-trained models on reasoning

Announced October 5, 2026

What happened

The paper analyzes how training data links the opening tokens of a base model's response to the reasoning behavior that follows. Fixing specific opening-token cues made base models' math and coding performance comparable to RL-trained models. For example, the cue '.\n\nOkay' raised Olmo-3-7B's MATH-500 pass@1 from 42% to 78%, and the cue 'Alright,' raised Qwen3-14B from 72% to 87%. The authors also reported that RL makes these cues more likely to appear, and that fixing the cues recovers much of the gain from RL.

Why it matters

It supports the view that much of RL's reasoning gain comes from drawing out behaviors base models already have. This has implications for how post-training gains are evaluated.

Sources

More about arXiv