🧪 Test?View on arXiv
HakemBench: A Turkish Benchmark of Typed Decisions
Not specified in the provided content
benchmarkingdecision-makingnatural language processingTurkish language
2610.02293
Builder Relevance
1h ago70%
Abstract
HakemBench is a Turkish benchmark of typed decisions that evaluates models based on their ability to read text, answer questions, and provide probabilities for options.
Reality Card
Core Claim
HakemBench provides a comprehensive benchmark for evaluating decision-making models in various domains with a focus on Turkish language processing.
Method / Result
Version 1.0 includes 2,346 items and 4,275 questions across seven tracks, achieving a composite score of 0.888 for the leading model.
Limitations
Most gold labels are derived from AI model comparisons rather than human verification, raising concerns about the reliability of the dataset.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.