ai-interview-platform: The cage around the AI interviewer

A real-time voice interview stack that treats Gemini like a component, not a decision-maker.

11 min read • View on GitHub • More from rakamindev

A recruiter’s desk reimagined as a controlled audio pipeline. Browser speech enters as a waveform through a narrow conduit, passes a valve-like gate, and reaches an AI interviewer behind a glass partition with status lights for reconnect, silence, and session hold. It explains the article’s core idea: the system surrounds the model with policy and transport controls.
The interview is only the surface. The real product is the control plane that keeps a live voice session coherent, recoverable, and governable.
Key Takeaways

This project is easy to misread. On the surface, it is an AI interviewer. Underneath, it is a stateful control system for a live voice session, built to keep a non-deterministic model inside deterministic rails.

That matters because interviews are not just conversations. They are workflow, policy, continuity, and evidence. The repo treats Gemini as one moving part in a larger assessment machine, not as the thing that gets to decide what happens next.

The interview is just the surface

The repo frames hiring as infrastructure. It records a candidate, routes audio through a backend-owned session, evaluates against skills and levels, then emits a fit-gap report and proctoring trail. That is a very different ambition from a chatbot with a microphone.

The engineering choice is the story. Instead of letting the model drive the experience, the system wraps it in transport, prompt governance, turn-taking control, and recovery logic. The interview feels conversational, but the platform underneath is procedural.

The browser talks audio. The backend talks policy.

The backend is not just relaying audio. It is preserving session state, regulating speech boundaries, and keeping the interview alive across browser churn.

The most important middleware lives in the backend. The repo uses a custom Rack WebSocket proxy instead of treating browser audio as a simple request-response problem, because voice sessions need low latency, state retention, and controlled reconnects.

That design is what lets the platform survive a refresh without collapsing the interview. The backend owns the session lifecycle, keeps the Gemini connection warm through a grace period, and routes audio in a way that the frontend alone could not safely manage.

A close mechanical scene of speech timing. A candidate waveform approaches a threshold, a thin silent shim is inserted between two gears, and a model response gate unlocks only after that pause lands. A side path shows reconnect continuity bypassing a broken browser link. It explains how turn-taking is actively controlled instead of guessed.
Turn-taking is the hardest part of voice AI. This repo solves it by forcing clear speech boundaries and keeping the session state separate from the browser.

Turn-taking is the hard part

Voice interfaces fail in the seams. If the model starts speaking too early, it echoes the user. If it waits too long, the conversation feels broken. This repo attacks that problem directly with a silence-pumping pattern and a model-speaking gate.

That is the subtle move. The system does not trust speech activity alone to mark the end of a turn. It actively nudges the stream with synthetic silence, then lets the model answer only when the backend decides the coast is clear.

candidate_speaks -> silence_pump injects frame -> backend detects turn end -> model_speaking gate opens -> Gemini responds -> transcript + scoring pipeline continue

The benefit is not just cleaner audio. It is fewer echo loops, fewer awkward overlaps, and a session that behaves more like a controlled service than a brittle demo.

The prompt is not a prompt. It is a policy compiler.

The prompt system is doing governance work. It compiles skills, levels, and interview rules into a large system instruction so the model behaves like an assessor, not a tutor.

The backend also injects a system signal token that the model cannot override. That is the key point: the interview ends because the platform says it ends, not because the model gets creative about wrapping up.

Rakamin helps Peruri to increase their hiring efficiency & effectivity with 6000+ applicants shortlisted in less than 5 days for their massive recruitment programs and 900+ applicants assessed without any hassle using our integrated recruitment & assessment platform.

Rakamin Academy, Platform Description · Recruitment & Development Redefined

That quote is product marketing, but it still reveals the intended operating mode. The platform is built to process candidates at volume, with enough control to make the output usable for real hiring decisions.

Scoring happens after the conversation, but not by vibes

The fit-gap engine splits the evaluation problem in two. Quantitative skill gaps are handled deterministically. Qualitative narrative is synthesized afterward, which keeps the scoring legible instead of asking the model to invent the whole rubric.

DimensionSimple wrapperai-interview-platformWhy it matters
Audio handlingBrowser sends audio directly to an APICustom WebSocket proxy, PCM handling, silence injectionThe system can shape latency and speech boundaries instead of hoping the model behaves
Session recoveryRefresh usually means restartBackend-owned session lifecycle with grace-period reconnectCandidates do not lose context when the browser drops
Turn-taking controlModel decides when to answerServer-side gate and silence pump manage the boundaryThis reduces overlaps, echo, and awkward interruptions
Prompt governanceOne prompt sets the tonePrompt compiler plus system signal tokenThe backend keeps authority over the interview rules
ScoringFree-form summary onlyDeterministic skill deltas plus generated narrativeThe output is easier to trust and compare
Multi-tenancyOften absentTenant-scoped data modelSeparate organizations can share the platform safely
ProctoringUsually minimalFace, window, and integrity signals are part of the stackAssessment data is more defensible
Deployment maturitySingle-page demoDocker, Redis, Sidekiq, Kubernetes manifestsThe repo is shaped like infrastructure, not a prototype

The table makes the distinction plain. This is not a thin UI on top of an LLM. It is a workflow system that uses generative models where they help and hard rules where they must.

Why the architecture feels unusually mature

There are signs of a system built for real deployment pressure: tenant scoping, background jobs, proctoring, and k8s manifests. Those pieces do not make the product glamorous, but they do make it believable.

That maturity shows up in the architecture choices too. A custom audio proxy, explicit reconnect handling, and downstream analysis are all signs that the team has thought past the first successful demo.

The result is a platform that sits between candidate-facing prep tools and recruiter-facing assessment systems. It is closer to assessment infrastructure than to interview theater.

What this project says about AI hiring tools

The bigger lesson is simple. If you want AI interviews to be useful, the question is not only what the model can say. It is what the platform can prevent, preserve, and prove.

That is why this repo stands out. It shows how to make a voice LLM act like a dependable assessment component without letting it control the session. The cage is the product.