Preprint proposes delayed-compression memory, self-tested gains over 10 F1 in long-context reasoning
CoEM postpones compression until later context confirms an excerpt's relevance; authors' own tests show 10+ F1 gains, pending independent replication.
A preprint introduces CoEM, a memory-management method for long-context reasoning: potentially useful source excerpts are kept verbatim in a pending set, and only after later context clarifies their relevance does a learned policy commit them to compressed memory or discard them, avoiding the information loss of premature compression.
The policy is trained with reinforcement learning, and a frozen verifier accepts memory facts only when supported by retained excerpts. On 6,400-document inputs, the authors report CoEM beats the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B.
The experiments cover a single model, Qwen3.5-9B, and a single input length of 6,400 documents; performance on other models and shorter or longer contexts is unknown. The code is open on GitHub for verification.
Sources:https://arxiv.org/abs/2609.36935