Papers/2609.35791
🧪 Test?View on arXiv

FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

Not provided in the content

streamingsemantic detectionfull-duplexaudio processing
2609.35791
Builder Relevance
80%
1h ago

Abstract

FD-VAD introduces a method for semantic endpoint detection in full-duplex voice interaction, eliminating the need for ASR.

Reality Card

Core Claim

FD-VAD outperforms existing streaming and non-streaming semantic turn classifiers, achieving the highest EOT recall of 0.853 on the TurnBench dev set in a zero-shot setting.

Method / Result

Achieved the highest EOT recall of 0.853 at FP<=0.10.

Limitations

The paper does not specify the authors or provide extensive details on the dataset used for evaluation.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers