Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms
Abstract
This paper evaluates large language models on their ability to generate culturally specific kinship terms in non-Western languages, revealing significant gaps between recognition and production capabilities.
Reality Card
The study demonstrates that large language models struggle to generate culturally specific kinship terms, achieving only 36.00% production accuracy compared to 90.67% recognition accuracy.
GPT OSS120B selects the correct term in 90.67% of valid cells but produces an accepted term in only 36.00% of attempts.
The evaluation format gap suggests that the results may not directly reflect lexical knowledge, raising concerns about reproducibility.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.