🧪 Test?View on arXiv
LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform
Vanellope Asher
literature reviewevaluationAI alignmentexpert judgment
2608.21374
Builder Relevance
2h ago70%
Abstract
The paper introduces LitReview Arena, a platform for evaluating literature review agents through expert comparisons, highlighting the challenges in aligning AI-generated reviews with human judgment.
Reality Card
Core Claim
LitReview Arena demonstrates that current literature review systems only win 23.0% of matches against human drafts in overall utility, while agentic LLMs significantly outperform base models.
Method / Result
Agentic LLMs like Sonar Deep Research outperform base language models by over 60%.
Limitations
Existing LLM-as-a-judge methods are misaligned with human expert evaluations, particularly on synthesis-heavy criteria.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.