Forcing Gemini to Keep Time: Inside midnight-memory

How a local-first workflow uses multimodal LLMs to solve the tedious math of lyric alignment and AI lip-syncing.

6 min read • View on GitHub • More from Sunwood-ai-labs

A human hand carefully inserts tiny, precise gears into a massive, fast-moving industrial loom, illustrating the manual effort required to time audio for AI video generation.
AI lip-sync generation requires perfectly sliced 7 to 15-second audio segments, a process that historically demands hours of manual timeline scrubbing.
Key Takeaways

The Lip-Sync Bottleneck

AI video generators require audio chopped into highly specific segments. For ecosystems like LTX, these chunks must be strictly between 7 and 15 seconds long. Doing this manually for a three-minute song involves hours of scrubbing a timeline, slicing audio, and adjusting SRT files. The friction is immense, turning automated video generation into a deeply manual chore. Midnight-memory was built to eliminate this exact bottleneck.

Reading Instead of Guessing

Whisper and traditional ASR models transcribe by guessing what they hear. When AI-generated vocals are muddy or buried in a complex mix, these models hallucinate. Midnight-memory takes a different path. It uses Gemini's massive multimodal context window to feed the model the exact lyrics alongside the audio file. It turns a transcription task into a pure alignment task. The AI is no longer guessing the words; it is simply acting as a highly precise timing engine.

A split composition showing a blindfolded person listening to a tin can string on the left, and a person tracking a printed script with a silver pointer while listening to a phonograph on the right.
Traditional ASR guesses blindly at muddy audio, while multimodal alignment anchors the model with a perfect reference text.
FeatureASR (e.g., Whisper)Multimodal Alignment
Core MechanismAudio-to-Text inferenceAudio+Text synchronous mapping
Muddy Audio AccuracyDegrades rapidly (hallucinates)Highly resilient (anchored by text)
Output StructureRaw text blockEnforced JSON cue array
Gap HandlingIgnores silenceInjects [melody] segments

Defensive Programming for Timelines

LLMs are notoriously bad at adhering to strict numerical constraints. If Gemini outputs an overlapping timestamp or a cue that exceeds the audio duration, the entire subtitle file corrupts. To solve this, midnight-memory implements defensive programming in its Python backend. The script utilizes a clamping function that forces the AI's output into mathematical compliance. It ensures minimum durations and prevents any two timeline segments from overlapping.

def clamp_cues(cues: list[Cue], max_duration: float) -> list[Cue]:
    for i in range(len(cues)):
        # Enforce minimum duration
        if cues[i].end - cues[i].start < 0.05:
            cues[i].end = cues[i].start + 0.05
        
        # Prevent overlap with next cue
        if i < len(cues) - 1 and cues[i].end > cues[i+1].start:
            cues[i].end = cues[i+1].start - 0.01
            
        # Clamp to bounds
        cues[i].end = min(cues[i].end, max_duration)
    return cues

The Lane-Based Alignment Engine maps fine-grained lyrics to coarse video segments, filling gaps with [melody] cues to maintain a continuous timeline.

The Local-First Safety Net

You cannot blindly trust AI-generated timelines for production rendering. The project includes a vanilla JavaScript viewer that runs entirely locally in the browser. It parses a generated manifest to render parallel tracks of coarse video segments and fine-grained lyrics. This allows creators to visually verify the alignment without uploading their original assets to a cloud provider, bridging the gap between automated generation and manual quality control.

A close-up of a magnifying glass focused tightly on a strip of physical celluloid film perfectly sliced into exact segments, measured by a steel machinist's caliper.
Local-first verification tools allow creators to measure and validate AI outputs before committing to expensive rendering pipelines.