📖 Read?View on arXiv
Agent Psychometrics: Task-Level Performance Prediction in Agentic Coding Benchmarks
coding-agentsbenchmarksevaluationswe-bench
2604.00594
Builder Relevance
Jul 2090%
Abstract
This work extends item-response theory with agent and task features to predict task-level performance in agentic coding evaluations.
Reality Card
Core Claim
Benchmark and agent teams can estimate task difficulty and success probability using LLM, scaffold, and rich task features instead of relying only on one aggregate solve rate.
Method / Result
The method evaluates across SWE-bench Verified, SWE-bench Pro, Terminal-Bench 2.0, and a software-optimization benchmark, while noting the cost of full agent-evaluation runs.
Limitations
Predictions are conditioned on the available evaluation data, feature extraction, benchmarks, and agent/scaffold coverage; they are not a replacement for production evaluations.