Dual-track memory self-reports top EverMemBench score, unreplicated
SpeakerMem-R1 self-reports 62.33% on EverMemBench as leaderboard best, but all results are first-party; track as a method signal.
重要度局所的証拠E2 未複製執筆簡易
Multi-party dialogue memory has a new method signal: SpeakerMem-R1 self-reports a 62.33% submission to EverMind-AI's public EverMemBench leaderboard, the highest reported, claiming to tackle both bottlenecks of who said what and reconstructing states across members and time. The method stores speaker-labeled verbatim messages and derived states in dual tracks, merging evidence by entity, event and time at query time.
Before this, no approach handled speaker attribution and state reconstruction together; the authors self-report binary accuracies of only 47.9%, 69.2% and 61.9% on GroupMemBench, SocialMemBench and EverMemBench.
In the authors' self-reported 305-question controlled test, RL lifted the SFT writer's mean accuracy from 57.38% to 68.20%; the leaderboard-best 62.33% is also the authors' own submission.
All numbers are first-party evaluations, not independently reproduced, and the sub-48% GroupMemBench score shows the problem is far from solved; the preprint material is dated 2026-09-22. Treat as a method signal to track.