🧪 Test?View on arXiv
ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration
Author1, Author2, Author3, Author4, Author5
MoEaccelerationinferenceexpert systems
2608.24938
Builder Relevance
3h ago80%
Abstract
ExFold proposes a unified framework for accelerating Mixture-of-Experts models during both prefill and decode phases without training.
Reality Card
Core Claim
ExFold achieves up to 1.41x TTFT and 2.45x TPOT speedups while retaining about 99% of the original average quality.
Method / Result
1.41x TTFT and 2.45x TPOT speedups.
Limitations
The method relies on calibrated scalar projectors which may require careful tuning on different datasets.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.