Papers/2608.27513
🧪 Test?View on arXiv

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

Not provided in the abstract

quantizationrecurrent neural networkslanguage modelsmemory optimization
2608.27513
Builder Relevance
80%
2h ago

Abstract

DAMP introduces a method for quantizing recurrent states in language models to reduce memory usage and improve decoding latency while maintaining accuracy.

Reality Card

Core Claim

DAMP reduces recurrent-state storage by 69.1% and accelerates the recurrent-state update kernel by up to 2.01x while maintaining average accuracy close to the FP32 baseline.

Method / Result

DAMP achieves 9.9 bits per state value with a 69.1% reduction in storage.

Limitations

The study focuses on specific language models (GDN and KDA), which may limit generalizability to other architectures.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers