Papers/2609.01725
🧪 Test?View on arXiv

From Visual Cues to Spoken Narration: Rethinking Audio Description

Author1, Author2, Author3, Author4, Author5

multimodalaudio descriptionvideo localization
2609.01725
Builder Relevance
80%
2h ago

Abstract

The paper presents Cue2Narrate, a novel two-stage pipeline for generating audio descriptions in movies that predicts both what and when to narrate, improving accessibility for visually impaired audiences.

Reality Card

Core Claim

Cue2Narrate outperforms existing baselines in audio description generation for long-form movie clips by 5-12 points in average mAP.

Method / Result

Cue2Narrate establishes the first benchmark for multi-segment audio description generation on long-form clips.

Limitations

The reproducibility of results may be limited by the availability of the LongLSMDC benchmark data and the complexity of the model.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers