Lip sync and audio-led editing
Lip syncing AI-generated characters to an existing vocal performance. The Odyssey as a test case. A genuinely humbling challenge.
Updated July 2026
I’m a huge fan of Epic the Musical, Jorge Rivera-Herrans’ serial adaptation of Homer’s Odyssey. My Goodbye is my favourite song in the entire soundtrack. It’s Athena delivering a cold, devastating farewell to Odysseus on the shores of burning Troy. With Greek mythology having a bit of a moment right now thanks to Christopher Nolan’s The Odyssey, I wanted to create a visual interpretation of that scene. Specifically the moment she says goodbye for the last time.
The goal was simple in theory: lip sync an AI-generated character to an existing vocal performance. In practice it was one of the most technically challenging things I’ve attempted in this space.
My Goodbye. A fan interpretation of the song by Jorge Rivera-Herrans. Non-commercial fan project.
Audio first is a different problem entirely
Typically how I work is to generate the voice to match the video. Doing it the other way around meant every tool in my pipeline was fighting me. The audio existed, it could not be changed, and the visuals had to serve it completely. The timing, the mouth shapes, the emotional delivery all had to match a performance I had no control over.
On top of that, lip sync results are very mixed the moment your character is anything other than a tight close-up. The further away the character is in the shot, or the moment you introduce a second character into the scene, the less reliable the results become. That creates a real tension between wanting cinematic variety in your shots and actually getting the lip sync to land.
What happens when you don’t lock the style first
The biggest mistake was one I’d made before and should have known better. I started generating character images before I’d locked either the characters or the visual world they existed in. Both drifted significantly as a result.
The early versions pulled toward photorealism. Wrong direction for the world I was building.
By the time I had a consistent background, the two characters had already evolved away from each other stylistically. I had to stop, rebuild reference sheets for both from scratch and re-establish the visual language before I could move forward.
The restyled versions landed on an illustrated mosaic aesthetic: gold and amber tones, graphic linework, glowing eyes. Much more fitting for a goddess.
Some of that inconsistency from the early drift is still visible in the final film. That’s entirely down to the sequence of decisions, not the tools.
Getting the characters locked
Having these locked before animation was the thing that held the later sequences together. Should have been step one, not an afterthought.
Building the environment
The environment was built in Midjourney to match the illustrated style of the characters once they were locked. Burning Troy, Greek ships in the harbour, that amber and smoke palette. Getting the background and characters to feel like they existed in the same visual universe took longer than it should have. Because the characters came first.
Four things I would tell myself at the start
01
Audio-first is a completely different workflow
When the audio already exists and can’t change, everything else has to bend to it. Your shot choices, your character distances, your pacing. It removes creative freedom and demands much tighter planning upfront.
02
Lip sync degrades fast with distance
Close-up, single character, facing camera: workable. Mid-shot, two characters, any movement: unreliable. Plan your lip sync shots as close-ups and cut away for everything else.
03
Lock the style before you lock the characters
The visual language of the world has to come first. Once characters start generating in the wrong style, you’re rebuilding from scratch anyway. Do it in the right order.
04
Reference sheets are not optional
Without a full character sheet locked early, every new generation pulls the character slightly somewhere else. The drift is subtle per shot and obvious across a sequence.
What I used and what each one did
ChatGPT
Character and prompt development. Used for working through the visual language, building character descriptions, and iterating on Midjourney prompt structure.
Midjourney
Character and environment assets. All still images: Athena, Odysseus, burning Troy, reference sheets. The illustrated mosaic aesthetic was developed and locked here.
Kling Omni
Video generation. Used for shots where lip sync was not required: wide shots, environment sequences, non-speaking character moments.
Higgsfield Seedance
Lip sync generation. The primary tool for all mouth-synced character close-ups. Results were workable on tight single-character shots and significantly less reliable on anything wider.
Nano Banana Pro
Audio isolation. Used to extract the vocal performance cleanly from the source track before feeding it into the lip sync pipeline.
Lalal.ai
Audio stem separation. Separated vocals, instrumentation and backing elements so each could be handled independently in the edit.
DaVinci Resolve
Editing, sound design and colour grade. Where everything came together: cutting to the audio, managing the sync, laying in the separated stems, grading toward the amber and smoke palette.