XTurnix Unifies Dialogue Turn Control
New model uses 5.5M samples for unified listen/speak state detection, hitting 89% accuracy on a custom benchmark.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
The XTurnix model simplifies dialogue turn control into two binary decisions, achieving 89.06% accuracy on an author-curated benchmark.
Traditional voice AI systems struggle with separate logic for "when to start speaking" and "when to stop," leading to complexity and errors. XTurnix pretrains on 5.5 million causal action examples automatically derived from timestamped two-speaker transcripts, enabling a single text-based model to predict a unified control token based on whether the AI is currently listening or speaking.
Across four public benchmarks, XTurnix achieved the best results on all SemanticVAD and LiveKit splits and the highest incomplete-turn accuracy on Easy-Turn. On the authors' balanced self-curated benchmark, its accuracy exceeded the strongest third-party baseline by more than 20 percentage points.
These results come from a preprint submitted on October 3. While code is open-source, the headline performance relies heavily on a non-public custom test set and awaits independent community reproduction.