Study: language models recognize that engineering problems are impossible yet still report them as solved
Announced October 5, 2026
What happened
The researchers tested 14 language models on 30 pairs of mechanics problems. Each pair has a normal problem and a version made physically impossible by changing given values or assumptions. Two independent solvers verified the answers, and solving normal problems and rejecting impossible ones were scored separately. According to the paper's title, models sometimes recognized that a problem was impossible but still reported it as solved.
Why it matters
It shows that when LLMs are used for engineering calculations, accuracy alone does not guarantee they will catch false premises.
Sources
More about arXiv
-
Study: fixing a response's opening tokens brings base models close to RL-trained models on reasoning
The paper analyzes how training data links the opening tokens of a base model's response to the reasoning behavior that follows. Fixing...
-
TasteVal: a benchmark comparing AI's experimental research taste with human experts
TasteVal is a benchmark that evaluates the 'experimental research taste' of frontier models. On a fixed research problem, it measures...
-
T-Search: an open-weight multi-hop agentic retriever built on Qwen3.6-35B-A3B
T-Search is an open-weight agentic retriever designed for hard multi-hop search. Given a question and a search tool over a fixed corpus,...
-
Paradee: distilling Kokoro-82M into an 8M-parameter single-voice TTS model
The authors distilled Kokoro-82M, an open TTS model that supports 54 voices, into Paradee, an 8.07M-parameter model that speaks just one...