Agent memory evaluation can now quantify accuracy, cost and latency together
DolphinBench requires memory systems to report accuracy, total cost and latency together; it is author-built and worth tracking as a method signal.
ImportanceLocalPreuvesE2 non réplicableTraitementRapide
Readers building agent memory systems now have an evaluation option that puts accuracy, total cost and latency on equal footing: DolphinBench requires all three, with none optional.
Before this, evaluations typically looked only at accuracy, leaving cost and latency unreported and making the real price of a memory method hard to judge.
The benchmark has three knowledge-work personas with roughly 500k tokens of user messages each, and 200 tasks per persona that depend on that history; the authors verified each task by requiring an agent to succeed with the relevant history and fail without it. This is first-party design and verification.
There is no independent reproduction or third-party adoption yet; the preprint material is dated 2026-09-21 and can be tracked as an evaluation option.