Humans Start Speaking a Median 151 ms Early; Current Voice Systems Struggle to Match
TurnBench benchmark tests 14 voice turn-taking systems: human listeners start a median 151 ms early in smooth turn transitions, and current systems struggle to match that timing without excessive false interruptions.
ImportânciaLocalEvidênciaE2 não replicada
In smooth turn transitions, human listeners begin speaking a median 151 ms early, while current voice turn-taking systems cannot yet match that timing without producing too many false interruptions, with false positives concentrated in backchannel-dense conversational styles.
Previously there was no turn-taking evaluation benchmark covering multiple conversational styles with human annotation and a public leaderboard, making it hard to compare systems against human timing.
Researchers released the TurnBench benchmark: 30 hours of human-annotated two-person conversation data, a 104-hour training set and a public leaderboard, covering six conversational styles with triple annotation; the paper tested 14 systems, and the authors report the above results.
The paper was accepted to IEEE SLT 2026, with the camera-ready updated on September 16; it does not mention whether the benchmark has been reproduced by third parties.
Source: TurnBench paper page (arXiv) ↗ | TurnBench leaderboard and dataset ↗