Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures
Not provided in the content
Abstract
The paper investigates the unexpected degradation of language models under few-shot prompting, revealing that the effect is task-dependent and proposing a method to isolate the impact of prompt length from content.
Reality Card
The proposed metric 'content delta' effectively predicts the performance of few-shot prompting, showing that models benefiting from demonstration content restructure their representations more.
Content delta correlates with performance (rho = +0.65, p = 0.043), while raw shift does not (r = 0.20).
The study's findings may be limited by the specific tasks and models evaluated, which may not generalize across all applications.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.