🧪 Test?View on arXiv
Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking
Not provided in the abstract
Mixture-of-Expertsmodel compressionsensitivity analysisexpert masking
2608.13565
Builder Relevance
1h ago80%
Abstract
This paper presents a systematic layer-wise sensitivity analysis of Mixture-of-Experts models, revealing depth-dependent sensitivity and providing insights for model compression.
Reality Card
Core Claim
Layer sensitivity in Mixture-of-Experts models is strongly depth-dependent, with late layers tolerating aggressive expert masking while maintaining output quality.
Method / Result
The narrow very-late policy (layers 35-39 @ 50%) retains 419/500 Good+Similar outputs while masking only 640 of 10,240 total experts.
Limitations
The findings may not yet compose cleanly with aggressive expert masking, indicating potential challenges in practical implementation.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.