Humans Start Speaking a Median 151 ms Early; Current Voice Systems Struggle to Match
TurnBench benchmark tests 14 voice turn-taking systems: human listeners start a median 151 ms early in smooth turn transitions, and current systems struggle to match that timing without excessive false interruptions.
ImportanceLocalEvidenceE2 unreplicated
In smooth turn transitions, human listeners begin speaking a median 151 ms early, while current voice turn-taking systems cannot yet match that timing without producing too many false interruptions, with false positives concentrated in backchannel-dense conversational styles.
Previously there was no turn-taking evaluation benchmark covering multiple conversational styles with human annotation and a public leaderboard, making it hard to compare systems against human timing.
Researchers released the TurnBench benchmark: 30 hours of human-annotated two-person conversation data, a 104-hour training set and a public leaderboard, covering six conversational styles with triple annotation; the paper tested 14 systems, and the authors report the above results.
The paper was accepted to IEEE SLT 2026, with the camera-ready updated on September 16; it does not mention whether the benchmark has been reproduced by third parties.
Source: TurnBench paper page (arXiv) ↗ | TurnBench leaderboard and dataset ↗