Give the whole work one principal voice and baseline rate, pitch and volume. Let AI classify the scene and each paragraph’s purpose, then make small adjustments within the styles the selected voice actually supports. Do not rebuild a separate voice context for every sentence.
What each SSML control is for
| Control | Purpose | Risk |
|---|---|---|
| break / punctuation | Section, transition and emphasis pauses | Too many pauses fragment the delivery |
| prosody | Small rate, pitch and volume movement | Wide ranges change perceived identity |
| express-as | Supported speaking styles and emotions | Unsupported styles may be ignored or rejected |
| sub / phoneme | Acronyms, brands and specialist pronunciation | Must match language and phoneme set |
Why the same voice can sound like a different person
A matching voice identifier does not guarantee matching acoustic context. Moving directly from poetic delivery to strong empathy, while also changing rate, pitch and style strength, creates a discontinuity. Keep one continuous voice range and change one dominant dimension only when the meaning really turns.
Lock the voice, baseline rate and narrator identity before adding paragraph variation. Emotion is a treatment, not a character switch. Narration, training and branded content should prioritise “one person speaking” over theatrical range.
Rate can change—when meaning calls for it
Dense information, numbers, disclaimers and key conclusions benefit from a slower rate. Transitions and light descriptive passages can move slightly faster. Add pauses before a new section and around the main conclusion. Work at paragraph or phrase level instead of assigning a random percentage to every sentence.
How intelligent direction should classify emotion
First determine the content type and narrator role, then classify each paragraph’s job: opening attention, neutral explanation, narrative tension, reflective warmth or energetic call to action. Generated direction must be constrained to the selected voice’s verified style list, with unknown values removed or safely returned to a natural style.
Validate SSML before synthesis
- Escape text nodes so ampersands and angle brackets cannot break XML.
- Allow only known elements, attributes and values; never execute invented model output.
- Verify closing tags, matching voice locale and document language.
- Fall back gracefully when one paragraph requests an unsupported style.
- Chunk long works at safe boundaries while preserving voice, parameters and order.
Why every preview needs a version
Regeneration may not reproduce identical prosody. Store generation time, text snapshot, SSML snapshot, voice, region, scene and billing context with each preview. Users can compare, download or mark an approved result without overwriting the last usable take.
Pre-publication listening checklist
- Do adjacent paragraphs still sound like one narrator?
- Are brands, names, acronyms, numbers and units pronounced correctly?
- Does each emotional shift have a semantic reason?
- Do rate and pauses suit the destination and subtitle timing?
- Do the downloaded audio, text and SSML belong to the same version?
Keep the script, direction and every preview in one project
Use the Windows client to cast voices, generate and edit SSML, compare previews and retain downloadable versions.