Full-duplex speech can cut interruptions to 8%, but task accuracy trails
Tencent's self-tested Gander posts the lowest interruption rate at 8% yet slightly trails the weakest rival on task accuracy; weights are unreleased.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
A full-duplex speech model can hold user interruptions to 8 percent, starting to speak at the right moment in all 100 scenarios, the lowest of any system tested. Previously, comparable systems struggled to time their turns: GPT-Realtime interrupted at 13.5 percent and the weakest competitor at nearly 48 percent. Tencent's Gander technical report, self-run, shows its 8 percent interruption rate below GPT-Realtime's 13.5 percent and the weakest competitor's nearly 48 percent, at the cost of task accuracy slightly below even that weakest system, which the team attributes to whole-system scoring that counts speech recognition errors. There is no independent replication and no standard evaluation for this model class; weights and training data will only be released after the open-source process completes, with a code repository and demo page available now.