releasesuno

Suno Speech, and what it does with the words you write

Silas Moser4 min read
Open Studio

You paste a script into Suno's new Speech mode, and the music under it is better than you expected. Then a word comes out wrong, or the voice stops for breath in the middle of a sentence. Users heard both in Suno's own tutorial, a little over three minutes in.

Speech opened to everyone on 1 October 2026, in beta. Before you blame the voice, check how much of what you hear was decided by the text you typed.

Suno splits the words from the delivery

The launch post calls Speech "the first audio model that generates voice and music together as one cohesive track." You "type an idea, a poem or something you've written, then describe the voice and musical style you have in mind."

The practical rules are in Suno's tutorial video. The help center had no Speech article as of 3 October. The script box holds what is said. A separate style box is, in the presenter's words, "where I can direct things like tone, pacing, mood, and setting." And the formatting advice runs opposite to song mode:

"Plain paragraphs work best. You don't need lyric formatting or bracketed section tags like you might need when you're creating a song."

Generations run "up to around 8 minutes," and the music can be switched off. Suno's web app caps the script at 5,000 characters and the style box at 1,000, though neither number appears on any Suno page.

Suno's demo reads at about 80 words a minute

The tutorial gives one finished spoken-word piece, so we timed it against the captions Suno published with the video. It delivers 53 words in about 40 seconds, roughly 80 words a minute. The presenter introducing it speaks 99 words in 33 seconds, about 178 a minute.

Much of that time is space between phrases. The style asked for a voice "restrained at first" over a score that grows, and Speech left the score room to do it.

That changes how long a script you can write. Our toast below averages 4.7 characters a word, spaces included, so a full 5,000-character box holds about 1,050 words.

At the demo's pace that is 13 minutes of audio, well past the 8-minute ceiling, and nobody has published what Speech does with the excess. For an unhurried read, about 600 words is what fits.

One demo, timed from caption cues, with a style that asked for slow. Treat 80 as the pace of that one request.

What the first users hear go wrong

Replies in the launch thread point to "particular" mispronounced at 3:13 of the tutorial, in the train announcement, and to a pause mid-sentence. Others report a voice that drifts after one to two minutes, and no way to use the voices saved for songs.

On pauses there is one structured test so far, a 48-minute French walkthrough of about 40 generations. Its finding: several ellipses between sentences gave better pauses than a bracketed pause tag with a duration, which needed several generations before it responded. These are first impressions from a few days of beta, and a beta changes week to week.

Two of those problems start on the page

On Suno's singing models, a misread word is decided by the text before the model makes a sound. The model performs the spelling and cannot see which meaning you intended, so a word with two readings gets one of them by chance. The mispronunciation guide shows the respelling that fixes it in songs.

Here is a toast written for this piece, 63 words with four traps in it, as one plain paragraph:

Before we raise a glass, I want to read the card Sam wrote after their
first date. Sam read it to me once, years ago, and made me promise never
to repeat it. So here it is. It says: I will live closer to you, even if
it makes me late for everything. They live a minute apart now. Sam is
still late.

The traps are "read" twice, once as reed and once as red, "live" as a verb both times, and "minute", which a model can read as my-NOOT. The second version keeps Suno's plain paragraph and changes only the spelling and the pauses:

Script:
Before we raise a glass, I want to reed the card Sam wrote after their
first date. Sam red it to me once, years ago, and made me promise never
to repeat it... So here it is... It says: I will liv closer to you, even
if it makes me late for everything. They liv a minnit apart now... Sam
is still late.

Style:
Warm best-man toast, unhurried, soft piano, a beat before the last line

We have not run this on Speech yet. Nobody has published whether a respelling survives into speech, or whether it comes out sounding like a typo.

What nobody has measured yet

  • Respelling. Whether "reed", "red" and "liv" fix the reading in Speech as they do in song mode.
  • Line breaks. Suno says plain paragraphs work best, and nobody has timed what a new line does instead.
  • Brackets. Suno says you do not need them, and testers disagree about whether a bracketed cue helps or confuses it.
  • Overflow. What happens to a script longer than the 8 minutes it can perform.

Paste the toast both ways, with the same style line and Variety at zero, and listen for the four words before you write anything longer.

Read next