Papers/2608.13573
🧪 Test?View on arXiv

A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing

workload analysisLLM servingcachingload-balancing
2608.13573
Builder Relevance
80%
1h ago

Abstract

This paper provides a comprehensive analysis of real-world LLM serving workloads over a one-year period, revealing insights into user interactions and workload evolution.

Reality Card

Core Claim

The study presents a longitudinal analysis of LLM serving workloads, capturing detailed user-model interactions and workload evolution across various models and users.

Method / Result

The paper includes a full one-year production trace from Chutes, which allows for in-depth analysis of LLM serving workloads.

Limitations

The study's findings may be limited by the specific context of the Chutes production environment, which may not generalize to all LLM serving scenarios.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers