🧪 Test?View on arXiv
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
Not provided in the abstract
benchmarkingGUI modelsmulti-step environmentscontextual consistency
2609.00048
Builder Relevance
1h ago70%
Abstract
This paper introduces GUI-CC, a benchmark for evaluating the contextual consistency of GUI world models as multi-step environments for agents.
Reality Card
Core Claim
GUI-CC demonstrates that current GUI world models often fail to maintain task-relevant context in multi-step interactions despite producing plausible single-step outputs.
Method / Result
Constructed 500 offline trajectory tasks and 200 emulator-verified online tasks across 30 mobile apps.
Limitations
The benchmark highlights that plausible single-step generation does not ensure reliable environment simulation, raising concerns about the reproducibility of multi-step interactions.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.