TasteVal: a benchmark comparing AI's experimental research taste with human experts
Announced October 5, 2026
What happened
TasteVal is a benchmark that evaluates the 'experimental research taste' of frontier models. On a fixed research problem, it measures how well a model iteratively designs experiments and draws conclusions from the results. It scores this by compute efficiency: a model scores higher the less serial experiment compute it needs to reach the same score as human experts.
Why it matters
It offers a way to evaluate AI research automation by the efficiency of experimental design rather than final-answer accuracy.
Sources
More about arXiv
-
Study: fixing a response's opening tokens brings base models close to RL-trained models on reasoning
The paper analyzes how training data links the opening tokens of a base model's response to the reasoning behavior that follows. Fixing...
-
Study: language models recognize that engineering problems are impossible yet still report them as solved
The researchers tested 14 language models on 30 pairs of mechanics problems. Each pair has a normal problem and a version made...
-
T-Search: an open-weight multi-hop agentic retriever built on Qwen3.6-35B-A3B
T-Search is an open-weight agentic retriever designed for hard multi-hop search. Given a question and a search tool over a fixed corpus,...
-
Paradee: distilling Kokoro-82M into an 8M-parameter single-voice TTS model
The authors distilled Kokoro-82M, an open TTS model that supports 54 voices, into Paradee, an 8.07M-parameter model that speaks just one...