Forcing Gemini to Keep Time: Inside midnight-memory
How a local-first workflow uses multimodal LLMs to solve the tedious math of lyric alignment and AI lip-syncing.
- Midnight-memory abandons traditional Automatic Speech Recognition by feeding both raw audio and exact lyric text into Gemini's multimodal context window.
- The workflow enforces strict timeline integrity using a Python backend that clamps AI-generated timestamps to prevent overlapping intervals.
- A zero-dependency, local-first vanilla JavaScript viewer allows creators to visually verify the alignment between fine-grained lyrics and coarse video segments without cloud uploads.
The Lip-Sync Bottleneck
AI video generators require audio chopped into highly specific segments. For ecosystems like LTX, these chunks must be strictly between 7 and 15 seconds long. Doing this manually for a three-minute song involves hours of scrubbing a timeline, slicing audio, and adjusting SRT files. The friction is immense, turning automated video generation into a deeply manual chore. Midnight-memory was built to eliminate this exact bottleneck.
Reading Instead of Guessing
Whisper and traditional ASR models transcribe by guessing what they hear. When AI-generated vocals are muddy or buried in a complex mix, these models hallucinate. Midnight-memory takes a different path. It uses Gemini's massive multimodal context window to feed the model the exact lyrics alongside the audio file. It turns a transcription task into a pure alignment task. The AI is no longer guessing the words; it is simply acting as a highly precise timing engine.
| Feature | ASR (e.g., Whisper) | Multimodal Alignment |
|---|---|---|
| Core Mechanism | Audio-to-Text inference | Audio+Text synchronous mapping |
| Muddy Audio Accuracy | Degrades rapidly (hallucinates) | Highly resilient (anchored by text) |
| Output Structure | Raw text block | Enforced JSON cue array |
| Gap Handling | Ignores silence | Injects [melody] segments |
Defensive Programming for Timelines
LLMs are notoriously bad at adhering to strict numerical constraints. If Gemini outputs an overlapping timestamp or a cue that exceeds the audio duration, the entire subtitle file corrupts. To solve this, midnight-memory implements defensive programming in its Python backend. The script utilizes a clamping function that forces the AI's output into mathematical compliance. It ensures minimum durations and prevents any two timeline segments from overlapping.
def clamp_cues(cues: list[Cue], max_duration: float) -> list[Cue]:
for i in range(len(cues)):
# Enforce minimum duration
if cues[i].end - cues[i].start < 0.05:
cues[i].end = cues[i].start + 0.05
# Prevent overlap with next cue
if i < len(cues) - 1 and cues[i].end > cues[i+1].start:
cues[i].end = cues[i+1].start - 0.01
# Clamp to bounds
cues[i].end = min(cues[i].end, max_duration)
return cues
The Local-First Safety Net
You cannot blindly trust AI-generated timelines for production rendering. The project includes a vanilla JavaScript viewer that runs entirely locally in the browser. It parses a generated manifest to render parallel tracks of coarse video segments and fine-grained lyrics. This allows creators to visually verify the alignment without uploading their original assets to a cloud provider, bridging the gap between automated generation and manual quality control.