Surgical video QA benchmark self-reports 14,256 pairs and 14.61% gain
SurgRAW authors self-report SurgCoTBench with 14,256 QA pairs and retrieval-augmented multi-agent reasoning beating supervised models by 14.61%, with no independent verification yet.
ImportanceLocalEvidenceE2 unreplicated
Surgical video AI now has a question-answering benchmark covering five surgical task types, SurgCoTBench, which the authors self-report contains 14,256 QA pairs, with retrieval-augmented multi-agent reasoning accuracy exceeding supervised models by 14.61%.
Previously the field lacked a unified surgical video question-answering benchmark, making it hard to compare methods across five surgical tasks.
Chang Han Low et al. self-reported the above benchmark and the 14.61% accuracy gain, updating arXiv v3 on September 16, with the entry marked as citing IEEE RA-L 2026; code is open on GitHub and the dataset has been released.
No independent verification yet.
Sources: