HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference
Author1, Author2, Author3, Author4, Author5
Abstract
This paper presents HybridInfer, a thermal-aware reinforcement-learning router that optimizes the selection of inference tiers for language models based on thermal constraints and query complexity.
Reality Card
HybridInfer significantly improves inference quality while managing thermal constraints by using a thermal-aware routing policy, outperforming hand-tuned heuristics.
The learned router achieves significantly higher quality than two hand-tuned heuristics on a benchmark of 210 prompts (p < 0.02).
The method's performance may vary based on specific hardware configurations and thermal conditions, which could limit reproducibility across different devices.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.