One change per generation is right for the wrong reason
Every prompt guide arrives at the same discipline, and it is good advice. Change one thing between generations. One of them tells you to open the lyrics, change only the problem line and regenerate, and it publishes hit rates for its method: about eight songs in ten against a quarter for the casual version, finishing at take three rather than take twelve.
Precise numbers, for a process nobody can hold still. The reason always given for the discipline is that it isolates the variable, the way a controlled experiment does.
It does not, and the gap matters, because it changes how many generations you need before you believe anything.
There is no seed to hold still
A controlled experiment needs everything except the one variable to stay put. In Suno, nothing stays put. There is no exposed seed, no way to say run that again exactly, and the same sheet returns different takes on Tuesday than it did on Monday. People have been asking for a reusable seed for years, and one of them is in the thread about songs cutting off, guessing correctly that the day's randomness was the culprit and wishing they could pin it.
So when you change a word and generate once, you are not comparing your old prompt against your new one. You are comparing one draw from the old distribution against one draw from the new one.
What a single comparison is actually worth
Say a generation gives you a keeper about three times in ten, which is generous on a bad week.
Now change nothing at all, generate before and after, and ask what you would conclude:
| What you see | How often, when the change does nothing |
|---|---|
| A difference in either direction | 42 per cent |
| It looks like the change helped | 21 per cent |
| It looks like the change hurt | 21 per cent |
| Both takes agree, so you learn nothing either way | 58 per cent |
Two takes in five will show you a difference that is not there. That is not a flaw in your listening. It is two coin flips.
This is also why the second generation feels so decisive and so often is not: the pattern where a change appears to work once and never again is exactly what randomness produces.
Change one thing, and wait for the third time
The discipline survives. The reason changes, and so does the stopping rule.
Here is a change worth logging, one line split in two and nothing else touched:
[Verse 1]
I drive back to the place on Aldridge Road
and each time I get there I forget the thing I came to say
[Verse 1]
I drive back to the place on Aldridge Road
I get as far as the porch
and forget what I came to say
Log it and generate. If the rush goes away, you have one observation, worth about as much as one coin landing heads.
The number that matters is what happens on repetition. If a change genuinely does nothing, the odds of it appearing to help three separate times, on three different songs, are about one in a hundred. So the rule is not "change one thing and trust the result". It is change one thing, write down what you saw, and act on it the third time you see it.
That also explains why the log beats memory. The evidence for any single edit is thin enough that it only becomes evidence in aggregate, and nobody remembers take 12 by take 14.
Log your next change and leave the conclusion until it has happened three times.