Papers/2608.21374
🧪 Test?View on arXiv

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

Vanellope Asher

literature reviewevaluationAI alignmentexpert judgment
2608.21374
Builder Relevance
70%
2h ago

Abstract

The paper introduces LitReview Arena, a platform for evaluating literature review agents through expert comparisons, highlighting the challenges in aligning AI-generated reviews with human judgment.

Reality Card

Core Claim

LitReview Arena demonstrates that current literature review systems only win 23.0% of matches against human drafts in overall utility, while agentic LLMs significantly outperform base models.

Method / Result

Agentic LLMs like Sonar Deep Research outperform base language models by over 60%.

Limitations

Existing LLM-as-a-judge methods are misaligned with human expert evaluations, particularly on synthesis-heavy criteria.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers