Study: fixing a response's opening tokens brings base models close to RL-trained models on reasoning
Announced October 5, 2026
What happened
The paper analyzes how training data links the opening tokens of a base model's response to the reasoning behavior that follows. Fixing specific opening-token cues made base models' math and coding performance comparable to RL-trained models. For example, the cue '.\n\nOkay' raised Olmo-3-7B's MATH-500 pass@1 from 42% to 78%, and the cue 'Alright,' raised Qwen3-14B from 72% to 87%. The authors also reported that RL makes these cues more likely to appear, and that fixing the cues recovers much of the gain from RL.
Why it matters
It supports the view that much of RL's reasoning gain comes from drawing out behaviors base models already have. This has implications for how post-training gains are evaluated.
Sources
- Official Base Models Can Reason By Taking a Cue From Training Data arXiv
More about arXiv
-
TasteVal: a benchmark comparing AI's experimental research taste with human experts
TasteVal is a benchmark that evaluates the 'experimental research taste' of frontier models. On a fixed research problem, it measures...
-
Study: language models recognize that engineering problems are impossible yet still report them as solved
The researchers tested 14 language models on 30 pairs of mechanics problems. Each pair has a normal problem and a version made...
-
T-Search: an open-weight multi-hop agentic retriever built on Qwen3.6-35B-A3B
T-Search is an open-weight agentic retriever designed for hard multi-hop search. Given a question and a search tool over a fixed corpus,...
-
Paradee: distilling Kokoro-82M into an 8M-parameter single-voice TTS model
The authors distilled Kokoro-82M, an open TTS model that supports 54 voices, into Paradee, an 8.07M-parameter model that speaks just one...