Papers/2608.24938
🧪 Test?View on arXiv

ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration

Author1, Author2, Author3, Author4, Author5

MoEaccelerationinferenceexpert systems
2608.24938
Builder Relevance
80%
3h ago

Abstract

ExFold proposes a unified framework for accelerating Mixture-of-Experts models during both prefill and decode phases without training.

Reality Card

Core Claim

ExFold achieves up to 1.41x TTFT and 2.45x TPOT speedups while retaining about 99% of the original average quality.

Method / Result

1.41x TTFT and 2.45x TPOT speedups.

Limitations

The method relies on calibrated scalar projectors which may require careful tuning on different datasets.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers