Voice¶
Voice lives on the timeline: the bottom track holds one or more voice clips, and you build the dialogue by creating, auditioning, and arranging them.
Creating a voice clip¶
Click the + button on the voice track (the clip is created at the playhead position). In the dialog:
- Type the line. There's a character limit — clips are short by design, up to about ten seconds of speech.
- Pick a voice and a tone preset from the dropdowns.
- Click Generate. The audio plays back so you can hear it right away.
Happy with the read? Click Submit. The studio then does two things: it enhances the audio (the raw generated voice is quiet and a little rough — this cleans and normalises it) and it analyzes the lip sync so the avatar actually speaks the line. This takes a few moments, and then the clip lands on the timeline.
If you have an existing WAV file you want the avatar to speak, use Load instead of generating.
Writing lines that generate well¶
A few practical rules that make a big difference:
- Write numbers as words. Type "zero", not "0" — digits confuse the voice model.
- Punctuation shapes the pacing. Commas give short pauses; an ellipsis ("…") forces a slightly longer break.
- For a real pause, use separate clips. Rather than fighting punctuation, split the text into two clips and position them on the timeline with the gap you want.
Every generation is different¶
The voice is generated fresh each time — the same line generated twice will sound different. This is inherent to the model, not a fault, and the workflow is built around it: generate, listen, and regenerate until you get a read you like.
- Occasional mispronunciations or odd phrasing happen; regenerating usually clears them.
- The tone presets are flavours of the avatar's speech drawn from his reference recordings — they steer the read but don't guarantee it. Some tones are more consistent than others. Always judge with your ears; if a read sounds off, regenerate or try another tone.
Building longer speech from short clips¶
This is the recommended way to work. Instead of putting a whole paragraph into one clip, split the speech into short phrases and give each its own clip on the timeline:
- A long line has to land perfectly end-to-end — one slip and you regenerate the whole thing. Short segments can each be regenerated individually until they're right.
- Real speech varies its tone between phrases; separate clips let you give each phrase its own tone preset.
- You control the pacing precisely by positioning the clips.
This is exactly how the campaign videos were produced.
Working with clips on the timeline¶
- Drag a clip to reposition it (it needs free room to move into).
- Trim a clip by dragging its edges — useful to tighten a slightly long take.
- Double-click a clip to reopen it, pre-filled with its text and tone. Change what you need, regenerate, and submit — the clip is replaced in place.
- Delete a clip you don't want.
The lip sync you see is a preview
In-studio lip sync is a fast approximation whose job is to help you direct and time the performance. The final, delivery-quality lip sync is produced later by the AI render.
Screenshot placeholder
Voice dialog and a timeline with multiple voice clips (to be added).