Papers/2609.19354
🧪 Test?View on arXiv

Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment

H. Velesaca, Author 2, Author 3, Author 4, Author 5

multimodalreasoningaction quality assessment
2609.19354
Builder Relevance
70%
2h ago

Abstract

This paper evaluates the capability of Vision-Language Models to perform zero-shot action quality assessment on Olympic diving videos.

Reality Card

Core Claim

The proposed ensemble regression framework improves performance in action quality assessment, achieving a Spearman correlation of 0.67 with a four-model configuration.

Method / Result

Achieved a Spearman correlation of 0.67, significantly higher than the standalone VLMs' correlation below 0.32.

Limitations

The reliance on open-source VLMs may introduce variability in performance across different implementations.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers