What to expect¶
NK-Studio is built around generative AI, and generative systems behave differently from traditional software. None of the points below are surprises or faults — they follow directly from how the system works, and the studio is designed to work comfortably within them. Knowing them up front makes the difference between fighting the tool and flowing with it.
The voice varies between attempts¶
Each generation is a fresh take — the same line will sound different every time, and occasionally a take has an artifact, a mispronunciation, or a drift in character. Regenerating and discarding is the normal workflow, not a workaround. Long sentences are riskier than short ones — build longer speech from several short voice clips instead of one long generation.
Renders vary between attempts too¶
The AI render carries the same kind of randomness. An unsatisfying result just means: re-render. Each attempt is independent, and seeds make a good result reproducible.
Subtle expressions render softer than bold ones¶
Smiles and strong expressions come through the AI render well; subtler ones — eyebrow raises and the like — can read weaker in the finished image than in the studio preview. If an expression matters to the shot, make it a bold one, and judge the result from the render rather than the preview.
In-studio lip sync is a preview¶
What you see while directing is a fast approximation to support timing decisions. The final mouth movement is produced by the AI render's dedicated lip-sync pass — that's what appears in the finished video.
Lip sync is tuned for English¶
Other languages are not supported.
The look is stylized by design¶
Expressive and slightly digital, not photo-real. The finished image comes from the AI render, not from the 3D preview you work with.
There are no eyes or blinking¶
The avatar wears sunglasses, so eye and blink behavior is not part of the performance.
Rendering takes real time and runs one job at a time¶
The two-pass AI render is the heavy stage and needs the full graphics card — expect around 15–20 minutes for a ten-second clip, proportionally less for shorter ones. That's why renders go through a queue you can stack up and leave running, why the studio switches into Render Mode while it works, and why editing waits until the queue is done. Clip length drives render time, so keep clips as long as they need to be and no longer.
One avatar, operator-driven¶
The studio is built around a single avatar (with Choupette as an optional companion), and nothing is live or autonomous — the avatar performs exactly what you compose and record, nothing more.