LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization
Not specified in the provided content
Abstract
This study proposes LongNovel, a multi-scale benchmark for hallucination detection in long-context novel summarization, addressing the challenges of hallucinations in this domain.
Reality Card
LongNovel is introduced as a multi-scale bilingual benchmark for hallucination detection in long-context novel summarization, constructed from 29 Chinese novels and designed to explore hallucination variations with context length.
The benchmark includes 8 types of hallucinations and demonstrates significant challenges for existing models.
The main limitation is the potential difficulty in reproducing the manual revisions made to the test set for data reliability.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.