Papers/2604.00594
📖 Read?View on arXiv

Agent Psychometrics: Task-Level Performance Prediction in Agentic Coding Benchmarks

coding-agentsbenchmarksevaluationswe-bench
2604.00594
Builder Relevance
90%
Jul 20

Abstract

This work extends item-response theory with agent and task features to predict task-level performance in agentic coding evaluations.

Reality Card

Core Claim

Benchmark and agent teams can estimate task difficulty and success probability using LLM, scaffold, and rich task features instead of relying only on one aggregate solve rate.

Method / Result

The method evaluates across SWE-bench Verified, SWE-bench Pro, Terminal-Bench 2.0, and a software-optimization benchmark, while noting the cost of full agent-evaluation runs.

Limitations

Predictions are conditioned on the available evaluation data, feature extraction, benchmarks, and agent/scaffold coverage; they are not a replacement for production evaluations.

← Back to all papers