Suno opens speech-plus-music joint generation beta, admits accents drift
Voice and score can now be generated in one pass, skipping the splicing step, but accent drift means human review still matters in beta.
ImportanceLocalEvidenceE3 inspectableWrite-upQuick
On October 1, Suno announced Speech, opening it to all users in public beta and calling it the first audio model to generate speech and music as a single coherent track.
The traditional workflow converts text to speech first, then splices in background music separately. Speech generates end to end: enter a text and describe the voice and music style, and it returns spoken audio with original scoring. The feature is built into the AI music platform Suno.
The company itself listed current limitations: accents can drift mid-track, such as a British accent sliding into Australian and back, and intonation and emotional pauses can be overly dramatic, slowing the pacing of sentences.