Delayed-compression memory CoEM lifts long-context reasoning F1 by over 10 points in self-tests
CoEM postpones compression until later context confirms an excerpt's relevance; authors' own tests show 10.4-11.4 F1 gains over the strongest baseline on Qwen3.5-9B, pending independent replication.
BedeutungLokalBeweisE2 nicht repliziertAufbereitungSchnell
Readers can now learn of CoEM, a memory-management method for long-context reasoning: potentially useful source excerpts are kept verbatim in a pending set, and only after later context clarifies their relevance does the method commit them to compressed memory or discard them, avoiding the loss of key details from premature compression; the authors' own tests show F1 gains of more than 10 points.
The prior practice was to compress memory early, which could discard key details before an excerpt's relevance became clear, causing information loss in long-context reasoning.
The decision policy is trained with reinforcement learning, and a frozen verifier accepts memory facts only when supported by retained excerpts. On 6,400-document inputs, the authors' self-tests show CoEM beats the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B.
The experiments cover a single model, Qwen3.5-9B, and a single input length of 6,400 documents; performance on other models and shorter or longer contexts is unknown, and no independent replication exists yet. The code is open on GitHub for verification.