🧪 Test?View on arXiv
What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus
Author1, Author2, Author3, Author4, Author5
benchmarkingtask difficultyAI evaluationagent performance
2609.26826
Builder Relevance
1d ago70%
Abstract
This paper investigates the distinction between genuine task difficulty and factors leading to perceived difficulty in AI benchmarks.
Reality Card
Core Claim
The study identifies that only 78 out of 125 tasks with no honest pass are genuinely unsolved, highlighting the need for evidence behind all-fail tasks in benchmarks.
Method / Result
Out of 125 tasks with no honest pass, only 78 were certified as unsolved candidates.
Limitations
The certified-unsolved label does not guarantee intrinsic hardness or completeness of the verifier.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.