MERIT benchmark measures the marginal utility of memory for tool-using LLM agents with explicit cost accounting
Read the original at arxiv.org→arXiv:2609.05441v1 Announce Type: new Abstract: Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history,...
Original headline: "When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents"
Coverage timeline
- Sep 9, 04:00 UTC arXiv cs.AI lead source When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
- Sep 9, 04:00 UTC arXiv cs.LG AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning