New benchmark adds cost and latency to agent memory evaluation
DolphinBench requires memory systems to report accuracy, cost and latency together; it is author-built and worth tracking as a method signal.
Original event 2026-09-21
A new preprint, DolphinBench, argues agent memory evaluation should report accuracy, total cost and latency together, with none of the three optional.
The benchmark has three knowledge-work personas with roughly 500k tokens of user messages each, and 200 tasks per persona that depend on that history. The authors verified each task by requiring an agent to succeed with the relevant history and fail without it.
This is first-party design and verification with no independent reproduction or third-party adoption yet; readers building memory systems can track it as an evaluation option.