Preprint: Residual Memory Network Helps Long-Horizon Agents Approach Full-Context Performance with 5.2% Input Positions
New neural memory network supplements summaries with soft tokens, improving source attribution and reducing tool errors while maintaining insight coverage.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Key Finding: The REMORY preprint introduces a neural memory network that appends a bounded sequence of soft memory tokens after text summaries, enabling frozen LLMs to approach full-context joint scores using only 5.2% of input positions.
Background: Long-horizon agents typically compress history via textual summaries to fit finite context windows, but pure text summaries often fail to support all subsequent decisions, leading to information loss or hallucinations.
Result: On the SummHay benchmark, the method improved source attribution with nearly unchanged insight coverage. Across long-horizon agent benchmarks like BrowseComp and Terminal-Bench 2.1, Qwen3.8-27B and GLM-5.3-Flash showed consistent gains and substantially fewer repeated tool outputs and errors.
Limitation: Results are self-reported by the authors and have not yet been independently reproduced; validation is primarily limited to specific open-source model architectures.