🧪 Test?View on arXiv
Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents
Not provided
evaluationgame theoryLLMclosed-loop
2610.02425
Builder Relevance
1h ago70%
Abstract
The paper introduces XiangqiBench, a benchmark for evaluating LLM agents in Chinese chess, emphasizing the importance of closed-loop outcomes over static evaluations.
Reality Card
Core Claim
The study demonstrates that static evaluations of LLM agents in chess do not correlate with actual game-winning success, highlighting the need for closed-loop evaluation metrics.
Method / Result
Only 13.9% of trials where the model played the correct first move resulted in a win.
Limitations
The reliance on static evaluations may misrepresent the true capabilities of LLM agents in dynamic game scenarios.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.