Topics
Research
Stories: 6 · newest first
October 7, 2026
-
OpenAI publishes results on open math problems from an internal frontier model, with Lean-formalized proofs
OpenAI announced new results on open mathematical problems obtained with an internal frontier model. It published the Lean proof...
-
OpenAI and Ironclad use contract workflows to train and evaluate computer-use agents
OpenAI and Ironclad said they are using complex contract workflows to train and evaluate AI agents. The aim is to improve computer-use...
October 6, 2026
-
Study: fixing a response's opening tokens brings base models close to RL-trained models on reasoning
The paper analyzes how training data links the opening tokens of a base model's response to the reasoning behavior that follows. Fixing...
-
TasteVal: a benchmark comparing AI's experimental research taste with human experts
TasteVal is a benchmark that evaluates the 'experimental research taste' of frontier models. On a fixed research problem, it measures...
-
Study: language models recognize that engineering problems are impossible yet still report them as solved
The researchers tested 14 language models on 30 pairs of mechanics problems. Each pair has a normal problem and a version made...
-
Paradee: distilling Kokoro-82M into an 8M-parameter single-voice TTS model
The authors distilled Kokoro-82M, an open TTS model that supports 54 voices, into Paradee, an 8.07M-parameter model that speaks just one...