Papers/2609.26942
🧪 Test?View on arXiv

Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms

multilingualkinship termslanguage generation
2609.26942
Builder Relevance
70%
2h ago

Abstract

This paper evaluates large language models on their ability to generate culturally specific kinship terms in non-Western languages, revealing significant gaps between recognition and production capabilities.

Reality Card

Core Claim

The study demonstrates that large language models struggle to generate culturally specific kinship terms, achieving only 36.00% production accuracy compared to 90.67% recognition accuracy.

Method / Result

GPT OSS120B selects the correct term in 90.67% of valid cells but produces an accepted term in only 36.00% of attempts.

Limitations

The evaluation format gap suggests that the results may not directly reflect lexical knowledge, raising concerns about reproducibility.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers