MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models
Not specified in the provided content
Abstract
MoR-MLLM introduces a computation-sparse framework for multimodal large language models that dynamically adjusts recursive depth based on token complexity, improving efficiency while maintaining performance.
Reality Card
MoR-MLLM significantly reduces training memory and computation complexity compared to existing tiny MLLMs while achieving high performance on vision-language tasks.
Achieved a reduction in training memory and computation complexity while retaining high performance on various vision-language tasks.
The paper does not specify the authors, which may hinder reproducibility and validation of results.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.