Memory evaluations must now report cost and latency together
DolphinBench moves long-term memory evaluation from QA to agent task completion and mandates reporting cost, latency and accuracy together.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
Long-term memory evaluation can now show cost, latency and accuracy at once: DolphinBench evaluates long-term memory through agent task completion rather than question-answer retrieval.
Previous memory benchmarks tested only question-answer retrieval, and the authors state no existing memory benchmark combines cost, latency and accuracy.
The benchmark has three knowledge-work personas, each with roughly 500k tokens of user messages and 200 tasks; every task was verified by the authors: the agent must succeed with the relevant history and fail without it. This is a first-party, author-run benchmark; the dataset and evaluation code are public at dolphinbench.ai.
Three authors released the DolphinBench preprint on arXiv (submitted September 21), but there is no third-party adoption or independent reproduction yet.