Papers/2608.14700
🧪 Test?View on arXiv

Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait Synthesis

Chaolong Yang, Xiaoyang Li, Yinghao Zhang, Yong Zhang, Jingyi Yu

emotion synthesisaudio-driventalking headsmachine learning
2608.14700
Builder Relevance
80%
1h ago

Abstract

Xemo-Talker improves emotion control in audio-driven talking heads by using explicit emotion-related losses and a novel Tri-Loss approach.

Reality Card

Core Claim

Xemo-Talker achieves state-of-the-art emotion classification accuracy while maintaining competitive lip synchronization and high inference efficiency.

Method / Result

Achieves performance approaching that measured on real videos.

Limitations

Training with explicit emotion-related losses poses significant difficulties due to the trade-off between lip synchronization and emotion control.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers