Preprint claims dual-track memory fixes multi-party attribution, unverified
Authors self-report 62.33% on EverMemBench as leaderboard best, but all results are first-party; track as a method signal.
Original event 2026-09-22
The preprint SpeakerMem-R1 targets two bottlenecks in multi-party dialogue memory: who said what, and reconstructing states across members and time. The method stores speaker-labeled verbatim messages and derived states in dual tracks, merging evidence by entity, event and time at query time.
The authors self-report binary accuracies of 47.9%, 69.2% and 61.9% on GroupMemBench, SocialMemBench and EverMemBench; their 62.33% submission to EverMind-AI's public EverMemBench leaderboard is the highest reported. In a 305-question controlled test, RL lifted the SFT writer's mean accuracy from 57.38% to 68.20%.
All numbers are first-party evaluations, and the leaderboard best is the authors' own submission, not independently reproduced; the sub-48% score on GroupMemBench shows the problem is far from solved. Treat as a method signal to track.