🧪 Test?View on arXiv
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
Not specified in the provided content
memorybenchmarkingtool-using agentscost accounting
2609.05441
Builder Relevance
1d ago80%
Abstract
This paper evaluates the utility of long-term memory in tool-using LLM agents through a new benchmark called MERIT, focusing on cost-aware task execution.
Reality Card
Core Claim
The introduction of the MERIT benchmark demonstrates that memory significantly improves dependent-task success rates for LLM agents, achieving scores from 0.55 to 1.00.
Method / Result
Memory lifts dependent-task success from a leak-verified floor of 0.00 to 0.55-1.00 across 23,440 scored episodes.
Limitations
The unpredictability of embedding retrieval performance and the variability in task success based on memory implementation may hinder reproducibility.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.