🧪 Test?View on arXiv
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
workload analysisLLM servingcachingload-balancing
2608.13573
Builder Relevance
1h ago80%
Abstract
This paper provides a comprehensive analysis of real-world LLM serving workloads over a one-year period, revealing insights into user interactions and workload evolution.
Reality Card
Core Claim
The study presents a longitudinal analysis of LLM serving workloads, capturing detailed user-model interactions and workload evolution across various models and users.
Method / Result
The paper includes a full one-year production trace from Chutes, which allows for in-depth analysis of LLM serving workloads.
Limitations
The study's findings may be limited by the specific context of the Chutes production environment, which may not generalize to all LLM serving scenarios.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.