🧪 Test?View on arXiv
BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents
Avi Skaar, Co-author 1, Co-author 2, Co-author 3, Co-author 4
memory strategiesevaluation protocolslocal language modelsbudget compliance
2609.13149
Builder Relevance
2h ago80%
Abstract
BudgetBench introduces a protocol for evaluating memory strategies in local large language models by treating input-token budgets as the independent variable.
Reality Card
Core Claim
The paper presents a reusable measurement surface for evaluating memory strategies that exposes budget-compliance failures and non-monotonic quality curves.
Method / Result
The harness conducted pilot studies revealing budget-compliance failures and operating points that single-budget evaluations hide.
Limitations
The early pilot's tokenizer approximation undercounts some served-model prompts, affecting the accuracy of violation diagnostics.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.