What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis
Author1, Author2, Author3, Author4, Author5
Abstract
This paper analyzes the limitations of current systematic generalization tasks and introduces TranSGrid, a testbed that integrates various forms of reasoning.
Reality Card
Existing systematic generalization tasks fail to comprehensively measure the capability due to oversimplifications, as demonstrated by the significant performance drop of models on the TranSGrid testbed.
The largest model solved 79.6% of a standard test set but only 55.3% of TranSGrid, highlighting the inadequacy of current evaluation methods.
The study's reliance on a specific testbed may limit generalizability to other contexts or models.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.