🧪 Test?View on arXiv
Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court
Author1, Author2, Author3, Author4, Author5
reasoningbenchmarkingclassificationlegal AI
2609.26945
Builder Relevance
2h ago70%
Abstract
This paper presents a sentence-level benchmark for evaluating large language models' ability to classify interpretive canons in judicial reasoning.
Reality Card
Core Claim
The study operationalizes interpretive canons for classification and provides a dataset of annotated decisions from the German Federal Constitutional Court, achieving mean F1 scores between 70.4 and 79.2 across models.
Method / Result
Mean F1 scores range from 70.4 to 79.2 across models for seven binary subtasks.
Limitations
GEPA-optimized prompts do not systematically outperform expert hand-written prompts, indicating potential limitations in prompt optimization techniques.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.