Papers/2609.26826
🧪 Test?View on arXiv

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

Author1, Author2, Author3, Author4, Author5

benchmarkingtask difficultyAI evaluationagent performance
2609.26826
Builder Relevance
70%
1d ago

Abstract

This paper investigates the distinction between genuine task difficulty and factors leading to perceived difficulty in AI benchmarks.

Reality Card

Core Claim

The study identifies that only 78 out of 125 tasks with no honest pass are genuinely unsolved, highlighting the need for evidence behind all-fail tasks in benchmarks.

Method / Result

Out of 125 tasks with no honest pass, only 78 were certified as unsolved candidates.

Limitations

The certified-unsolved label does not guarantee intrinsic hardness or completeness of the verifier.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers