🧪 Test?View on arXiv
Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment
H. Velesaca, Author 2, Author 3, Author 4, Author 5
multimodalreasoningaction quality assessment
2609.19354
Builder Relevance
2h ago70%
Abstract
This paper evaluates the capability of Vision-Language Models to perform zero-shot action quality assessment on Olympic diving videos.
Reality Card
Core Claim
The proposed ensemble regression framework improves performance in action quality assessment, achieving a Spearman correlation of 0.67 with a four-model configuration.
Method / Result
Achieved a Spearman correlation of 0.67, significantly higher than the standalone VLMs' correlation below 0.32.
Limitations
The reliance on open-source VLMs may introduce variability in performance across different implementations.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.