sound
Recording the voice track before the puppet moves
Definition
In a scripted stop-motion feature, actors record their dialogue in a booth against the finished script before a single frame of the finished shot is animated -- not against picture, the way live-action ADR loops a performance to footage that already exists, and not to an unscripted interview, the way vox pop technique inverts the order of production entirely. The fixed recording becomes the animator's brief: a director walks a prototype puppet through the selected take as a rehearsal, and only then is that take broken into per-frame phonemes for the animator who will build the finished physical performance around it.
Overview
On Early Man (2018), Aardman animation supervisor Will Becher described the order of operations plainly: voice actors record dialogue first. "Nick works closely with the actors to find the character's voice," Becher said, and once a session is recorded, animators "animate the prototype puppet to recorded dialogue from Nick's session with the actor" -- reading the fixed audio like a director's brief before a finished puppet or set exists (VFX Voice, 2018-04-03).
That order can flatten what the recording session itself actually contributes. Hugh Grant, who voiced the lead in Aardman's The Pirates! Band of Misfits, recorded his lines to a script with no picture to react to -- and by his own account, that session was only the starting point, not the finished performance: "these animated geniuses are doing 80% of the work," he said, describing how the puppet animators built the comedic and emotional performance on top of a voice recording that only had to "hit the right spots and get the right tone" (ScreenCrush, 2012-04-25). The audio is fixed the moment the session wraps; almost everything a viewer reads as "performance" -- timing, expression, physical comedy -- is still ahead of it, built frame by frame against a soundtrack that cannot change.
The same dependency runs even further at LAIKA, where the fixed audio does not just brief a body animator -- it locks a fully rendered facial performance before that animator ever picks up the puppet. On Missing Link (2019), Brian McLean, LAIKA's director of rapid prototyping, described how the studio's 3D-printed replacement-face pipeline reverses the usual animation order: "Normally in animation...the last thing that they do is they add the facial animation on top," but at LAIKA "we're doing facial animation months before an animator's even on set with their puppet," because each face has to be animated in CG, printed, tested and delivered before the shoot reaches that scene -- so "when an animator is out on set...the facial animation is already pre-determined and already locked down" (3dprint.com, 2019-03-20). Director Chris Butler traced that chain back to the recording booth itself: voice actors work "with just a script and headphones," and a performance's inflections -- "for humor or sorrow" -- become "a road map for the rest of the years-long adventure in the making" (MovieMaker Magazine, 2019-05-10). Where Aardman's fixed track briefs one animator's physical performance, LAIKA's turns that same track into a second, pre-built performance layer -- the printed face -- that is already finished before body animation begins.
This is the mirror image of two other stop-motion sound practices. It is not vox pop dialogue, where unscripted, real speech is recorded first and a character is invented to match it (see vox pop dialogue) -- here the script exists first, and a professional actor performs it to a booth microphone with no scene to react to. And it is not ADR or live-action looping, where an actor matches new lines to picture that has already been shot -- in stop motion there is no picture yet; the puppet does not move until the audio already has. Once a take is selected, it moves from the recording stage into the shooting stage through Dragonframe's Audio Workspace and X-Sheet, where the fixed recording is broken into per-frame phonemes so the line animator knows exactly which mouth shape a hand-sculpted or replacement mouth needs to hit on each exposure (see dialogue and lip-sync planning in Dragonframe's X-Sheet).
LAIKA's Wildwood (2026) adds a further variant: filming the voice session itself as physical-performance reference, on top of the fixed audio. Cast members recorded lines years ahead of animation -- Peyton Elizabeth Lee has said "when we were in the booth, they filmed us the whole time recording" -- and animators drew on that footage, not just the phonemes, for "certain facial expressions or hand gestures" the actors carried into the booth (The Nerds of Color, 2026-09-18). Jacob Tremblay described receiving little more than "a little bit of concept art...just some early sketches" before recording, so the performance itself, not a pre-built visual target, set the direction for the character. Director Chris Butler has framed this as a deliberate escalation of the practice's live-action-reference habit: "we were really trying to sell this as real people," using "more live-action reference [than] any typical animation" so that, as he put it, viewers forget the result "has been manipulated frame by frame" (The Nerds of Color, 2026-09-18). Where Aardman's fixed track briefs an animator's physical performance and LAIKA's own replacement-face pipeline pre-locks a facial layer from that track, Wildwood's filmed booth sessions add a third input: a recorded physical performance the puppet animator can study and echo, on top of the audio that still cannot change once the session wraps.
Workflow
- The script is locked before the recording session -- unlike live-action ADR, there is no picture yet for the actor to match.
- The actor performs the scripted lines to a booth microphone; the session is recorded and edited down to a selected take, the same editorial pass any dialogue recording gets.
- The director and puppet department screen the selected take against a prototype puppet as a rehearsal, reading the physical performance the fixed audio calls for before a finished puppet or set exists.
- The take passes to Dragonframe's Audio Workspace, where it is broken into per-frame phonemes on the X-Sheet so the animator knows which mouth shape to hit on which exposure.
- The animator builds the finished physical performance frame by frame against that fixed, unchangeable audio -- the reverse of live-action ADR, where the performance is already finished on screen and the actor instead matches newly recorded lines to it.