Papers/2610.02293
🧪 Test?View on arXiv

HakemBench: A Turkish Benchmark of Typed Decisions

Not specified in the provided content

benchmarkingdecision-makingnatural language processingTurkish language
2610.02293
Builder Relevance
70%
1h ago

Abstract

HakemBench is a Turkish benchmark of typed decisions that evaluates models based on their ability to read text, answer questions, and provide probabilities for options.

Reality Card

Core Claim

HakemBench provides a comprehensive benchmark for evaluating decision-making models in various domains with a focus on Turkish language processing.

Method / Result

Version 1.0 includes 2,346 items and 4,275 questions across seven tracks, achieving a composite score of 0.888 for the leading model.

Limitations

Most gold labels are derived from AI model comparisons rather than human verification, raising concerns about the reliability of the dataset.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers