Papers/2610.02425
🧪 Test?View on arXiv

Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

Not provided

evaluationgame theoryLLMclosed-loop
2610.02425
Builder Relevance
70%
1h ago

Abstract

The paper introduces XiangqiBench, a benchmark for evaluating LLM agents in Chinese chess, emphasizing the importance of closed-loop outcomes over static evaluations.

Reality Card

Core Claim

The study demonstrates that static evaluations of LLM agents in chess do not correlate with actual game-winning success, highlighting the need for closed-loop evaluation metrics.

Method / Result

Only 13.9% of trials where the model played the correct first move resulted in a win.

Limitations

The reliance on static evaluations may misrepresent the true capabilities of LLM agents in dynamic game scenarios.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers