What you'll be able to do
- Pick settings that match the kind of audio you are making
- Diagnose why a voice sounds robotic or unstable
- Write scripts that survive being read by a machine
- Make a defensible call on voice cloning consent
Inside the path
A focused set of five-minute lessons. Each one ends with a hands-on exercise, not a quiz you can guess.
What the sliders do 5 min
Stability, similarity, and style, in plain terms.
Settings by use case 5 min
Narration, conversation, character work: different targets.
Writing for the ear 5 min
Punctuation, sentence length, and the words that trip TTS up.
Cloning and consent 5 min
Where the technology outruns what you should do with it.
Try a sample exercise
This is the kind of card you'd practice inside Iro: you do the thinking, then get feedback.
◆ Sample exercise · Tool judgment
You are narrating a 20-minute technical tutorial. Your first render sounds flat and slightly robotic, and listeners said it was hard to stay with.
Your task: Choose the change most likely to fix it.
- Raise stability to 100% so the delivery is perfectly consistent.
- Lower stability to roughly 65 to 75% and leave similarity around 75%, then re-render a short section to compare.
- Raise similarity to 95% to make the voice sound more like the original speaker.
- Raise style exaggeration to add emotion across the whole tutorial.
See why the second setting wins
Flatness is usually too much stability, not too little. At 100% the model removes the small pitch and timing variations listeners subconsciously read as a human speaking, which is exactly the robotic quality being described. Dropping into the 65 to 75% band keeps a technical read clear while restoring some natural cadence. Pushing similarity to 95% tends to introduce artifacts rather than realism, since the setting governs how hard the model matches the target voice, not how good it sounds. Cranking style adds theatrical delivery that fights a technical tutorial. And re-rendering a short section first is the habit worth keeping: settings interact with the specific voice, so compare rather than assume.
Stability is the setting that matters
Of the three sliders, stability does the most work. It governs how much the voice varies between sentences.
High stability produces consistent, predictable delivery. Push it to 100% and you strip the micro-variations in pitch and timing that listeners subconsciously associate with a real person, which is why maximum consistency sounds least human. Low stability lets cadence shift and small imperfections through, which reads as spontaneous but can wander on long passages.
Rough targets: 40 to 55% for storytelling, podcasts, and character dialogue, where emotional range is the point. 65 to 75% for technical tutorials and corporate narration, where clarity matters more but you still want it alive.
Similarity and style
Similarity controls how closely the output tracks the target speaker and sharpens clarity. It is tempting to max it, and that is usually a mistake: above roughly 75 to 80% it starts producing artifacts. If a clone sounds slightly metallic or crunchy, similarity is the first thing to bring down.
Style exaggerates delivery. Leave it at 0 unless you have a specific reason. It is a strong effect and it fights most straight narration.
A reliable starting point for a new voice is stability 50, similarity 75, style 0. Change one at a time and re-render the same 20 seconds so you can actually hear what each change did.
Write for the ear
Settings only carry you so far if the script fights the reader. Text that works on a page often does not work aloud.
- Shorten sentences. A clause that works visually can leave a synthetic voice running out of breath.
- Punctuate for pacing. Commas and full stops are the main lever you have over rhythm.
- Spell out anything ambiguous. Numbers, acronyms, and units get read in ways you did not intend. Write "twenty twenty six" if that is what you want to hear.
- Read it out loud yourself first. If you stumble, the model will too.
Cloning, consent, and where the line is
Voice cloning works well enough that the technical question stopped being interesting and the ethical one took over. The line is straightforward and worth stating plainly: clone your own voice freely, clone someone else's only with their explicit permission, and never use a cloned voice to make someone appear to say something they did not say.
That is not only an ethical position. Impersonation is increasingly a legal exposure, and platforms are getting faster at removing it. The reputational cost of getting this wrong is far larger than the time saved.