animation
Dialogue and lip-sync planning in Dragonframe's X-Sheet
Dragonframe
Definition
The workflow, built into Dragonframe's Audio Workspace and X-Sheet, for breaking a pre-recorded dialogue track into per-frame phonemes before a puppet is animated, so the animator knows exactly which mouth shape a replacement face set or hand-sculpted mouth needs to hit on which exposure -- the planning stage that precedes, and is distinct from, the Foley and sound-design pass performed in post.
Overview
A stop-motion puppet's mouth cannot improvise the way a live actor's can: the mouth shapes it will hit are either a fixed set of replacement parts or a small number of hand-sculpted poses, decided before the animator ever touches the puppet. That forces the order of operations in the opposite direction from live action -- the dialogue has to exist, frame-accurately mapped, before the performance is built, not recorded to match a performance that already happened.
Dragonframe's Audio Workspace is the tool most productions use to do this mapping. An editor imports the dialogue track, transcribes the line into a text box running alongside the waveform, then scrubs through the audio slowly and enters the phoneme heard on each individual frame -- not the whole word, just the vowel or consonant sound, with particular attention to vowels and the visually distinct B, P and M shapes, entering a sound a frame or two early rather than late. The X-Sheet can then carry a Phonetics column per character alongside Mouth and Eyes columns, so the assigned sound and the replacement part it calls for sit next to each other on the same exposure row.
That phoneme track stays visible while the animator works -- in the X-Sheet, the Timeline, or a configurable on-screen Audio HUD -- so matching the puppet's mouth to the sound is a lookup, not a guess. The replacement mouths themselves are typically built as a Custom Face Set: a multi-layered file with separate groups for mouth, eyes, brows and other swappable parts, 3D printed or fabricated once the phoneme count for the line is known.
This is the planning half of a two-stage process; the performance half comes later. On Guillermo del Toro's Pinocchio (2022), sound designer Scott Martin Gershin built the film's Foley to an animated scratch track the animators had already shot -- meaning the X-Sheet's phoneme pass happens first, driving the puppet's performance, and Foley or dialogue mixing is layered onto the finished animation afterward, not the other way around (see stop-motion-foley).
Workflow
- Record and lock a reference dialogue track before animation begins, typically a scratch or guide track from the voice performance.
- Import the audio into Dragonframe's Audio Workspace and transcribe the line into the dialogue text box that runs alongside the waveform.
- Scrub through the clip and mark the phoneme heard on each frame -- vowels and the B, P and M consonants matter most, since those are the shapes a mouth replacement or sculpt has to physically hit.
- Add Phonetics, Mouth and Eyes columns for the character to the X-Sheet so the marked sound and its matching replacement part are visible on the same row.
- Reference the phoneme and waveform data live while shooting, via the X-Sheet, Timeline or Audio HUD, swapping in the Custom Face Set part that matches the sound marked for that exposure.