Lore

Voice Cloning and Per-Character Voice Assignment Workflow

Overview

Dialogue audio is produced by splitting a scene's dialogue into a separate track per character, then assigning each character a voice: the creator's own cloned voice for characters they perform, and preset voices for secondary characters.

Method

Why it matters

Cloning by timbre-swap (rather than full resynthesis) is what supposedly keeps the original performance's prosody intact — the acting happens once, at record time, and voice cloning only relabels who it sounds like.

Pixel Audio: Timbre-Only Voice Swap and Preset Voices

Higgsfield's Pixel Audio 'change voice' feature swaps a recording's vocal timbre while claiming to preserve the original emotion, pacing, and delivery: "The emotions, the pacing, the delivery, all of that stays exactly the same." The workflow: record (or generate) a natural line reading first, capturing the real performance choices, then run that recording through 'change voice' to swap in a different voice identity. The implication is that this style of voice cloning acts as a timbre filter over a source performance rather than a performance generator — the acting comes from the original recording, not the model.

For multi-character scenes, split dialogue onto one audio track per character before cloning, so each character's track can be assigned its own voice independently. Not every character needs a personal clone: secondary characters can instead draw a fitting preset voice from Pixel Audio's built-in library, reserving full voice cloning for characters whose vocal identity matters most.