Papers/2609.09185
🧪 Test?View on arXiv

Integrating Unimodal and Vision-Language Representations in Latent Space for Multi-Label Chest X-Ray Classification

Not provided

multimodalchest X-rayclassificationlatent space
2609.09185
Builder Relevance
70%
1d ago

Abstract

This study proposes a framework that combines unimodal and vision-language representations for improved multi-label chest X-ray classification.

Reality Card

Core Claim

The proposed framework achieves a mean AUROC of 0.840 and an mAP of 0.467 by effectively integrating RAD-DINO and BioViL-T representations.

Method / Result

The best-performing model achieves a mean AUROC of 0.840.

Limitations

The study has only been evaluated internally on MIMIC-CXR-JPG, raising concerns about generalizability to other healthcare data.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers