🧪 Test?View on arXiv
Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses
David, Laya, Jev
System-1 modelsdecision-makingagent harnessesmodel evaluation
2610.02267
Builder Relevance
1h ago70%
Abstract
This paper evaluates System-1 decision models for agent harnesses, highlighting their accuracy and limitations in decision-making tasks.
Reality Card
Core Claim
Jev outperforms Laya significantly on 9 out of 11 decision points, demonstrating the potential for improved accuracy in System-1 decision models.
Method / Result
Jev shows an accuracy improvement of +10.8 to +46.0 percentage points over Laya on various decision points.
Limitations
The models did not perform better than chance on zero-shot model routing, raising concerns about their reliability in certain contexts.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.