Papers/2609.13149
🧪 Test?View on arXiv

BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents

Avi Skaar, Co-author 1, Co-author 2, Co-author 3, Co-author 4

memory strategiesevaluation protocolslocal language modelsbudget compliance
2609.13149
Builder Relevance
80%
2h ago

Abstract

BudgetBench introduces a protocol for evaluating memory strategies in local large language models by treating input-token budgets as the independent variable.

Reality Card

Core Claim

The paper presents a reusable measurement surface for evaluating memory strategies that exposes budget-compliance failures and non-monotonic quality curves.

Method / Result

The harness conducted pilot studies revealing budget-compliance failures and operating points that single-budget evaluations hide.

Limitations

The early pilot's tokenizer approximation undercounts some served-model prompts, affecting the accuracy of violation diagnostics.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers