MemCon: Dynamic Memory Boosts Agents
New framework MemCon uses reinforcement learning to dynamically manage LLM agent memory, showing improved task success and reduced token costs.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
The MemCon framework models memory operations for LLM agents as a Markov Decision Process, using an online learning policy to adaptively decide when and how much to retrieve.
Most existing agents rely on fixed heuristics for external memory access, which can be inefficient during early task stages or long-running sessions. MemCon employs a lightweight contextual bandit algorithm that converges without pretraining or additional LLM calls.
Experiments across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones show the method improves task success by up to 15.2 percentage points over baselines while reducing token consumption by 5–20%.
These results are self-reported in a preprint (arXiv:2607.13591v2) and have not yet been independently reproduced.