When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha
Not provided
Abstract
This paper evaluates the safety risks of conversational AI systems used for mental health support among Generation Alpha, highlighting significant gaps in vocabulary comprehension and clinical risk calibration.
Reality Card
Conversational AI systems for mental health support understand 76-82% of vocabulary but only correctly calibrate 64-72% of clinical risk, leading to a significant vocabulary-comprehension gap.
A 10-14 percentage point vocabulary-comprehension gap was identified between AI models and human therapists.
The study's findings may be limited by the specific architectures evaluated and the context of the conversations used in the benchmarks.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.