Papers/2610.02267
🧪 Test?View on arXiv

Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

David, Laya, Jev

System-1 modelsdecision-makingagent harnessesmodel evaluation
2610.02267
Builder Relevance
70%
1h ago

Abstract

This paper evaluates System-1 decision models for agent harnesses, highlighting their accuracy and limitations in decision-making tasks.

Reality Card

Core Claim

Jev outperforms Laya significantly on 9 out of 11 decision points, demonstrating the potential for improved accuracy in System-1 decision models.

Method / Result

Jev shows an accuracy improvement of +10.8 to +46.0 percentage points over Laya on various decision points.

Limitations

The models did not perform better than chance on zero-shot model routing, raising concerns about their reliability in certain contexts.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers