ai-interview-platform: The cage around the AI interviewer
A real-time voice interview stack that treats Gemini like a component, not a decision-maker.
- The repo’s real invention is the control plane around the model, not the model itself.
- Its audio path, reconnect logic, and silence pumping turn a fragile voice demo into a stateful system.
- The backend owns the interview lifecycle, so Gemini behaves like a constrained component instead of the source of truth.
- Scoring stays credible because deterministic skill deltas and generated narrative are separated on purpose.
This project is easy to misread. On the surface, it is an AI interviewer. Underneath, it is a stateful control system for a live voice session, built to keep a non-deterministic model inside deterministic rails.
That matters because interviews are not just conversations. They are workflow, policy, continuity, and evidence. The repo treats Gemini as one moving part in a larger assessment machine, not as the thing that gets to decide what happens next.
The interview is just the surface
The repo frames hiring as infrastructure. It records a candidate, routes audio through a backend-owned session, evaluates against skills and levels, then emits a fit-gap report and proctoring trail. That is a very different ambition from a chatbot with a microphone.
The engineering choice is the story. Instead of letting the model drive the experience, the system wraps it in transport, prompt governance, turn-taking control, and recovery logic. The interview feels conversational, but the platform underneath is procedural.
The browser talks audio. The backend talks policy.
The most important middleware lives in the backend. The repo uses a custom Rack WebSocket proxy instead of treating browser audio as a simple request-response problem, because voice sessions need low latency, state retention, and controlled reconnects.
That design is what lets the platform survive a refresh without collapsing the interview. The backend owns the session lifecycle, keeps the Gemini connection warm through a grace period, and routes audio in a way that the frontend alone could not safely manage.
Turn-taking is the hard part
Voice interfaces fail in the seams. If the model starts speaking too early, it echoes the user. If it waits too long, the conversation feels broken. This repo attacks that problem directly with a silence-pumping pattern and a model-speaking gate.
That is the subtle move. The system does not trust speech activity alone to mark the end of a turn. It actively nudges the stream with synthetic silence, then lets the model answer only when the backend decides the coast is clear.
candidate_speaks -> silence_pump injects frame -> backend detects turn end -> model_speaking gate opens -> Gemini responds -> transcript + scoring pipeline continue
The benefit is not just cleaner audio. It is fewer echo loops, fewer awkward overlaps, and a session that behaves more like a controlled service than a brittle demo.
The prompt is not a prompt. It is a policy compiler.
The prompt system is doing governance work. It compiles skills, levels, and interview rules into a large system instruction so the model behaves like an assessor, not a tutor.
The backend also injects a system signal token that the model cannot override. That is the key point: the interview ends because the platform says it ends, not because the model gets creative about wrapping up.
Rakamin helps Peruri to increase their hiring efficiency & effectivity with 6000+ applicants shortlisted in less than 5 days for their massive recruitment programs and 900+ applicants assessed without any hassle using our integrated recruitment & assessment platform.
That quote is product marketing, but it still reveals the intended operating mode. The platform is built to process candidates at volume, with enough control to make the output usable for real hiring decisions.
Scoring happens after the conversation, but not by vibes
The fit-gap engine splits the evaluation problem in two. Quantitative skill gaps are handled deterministically. Qualitative narrative is synthesized afterward, which keeps the scoring legible instead of asking the model to invent the whole rubric.
| Dimension | Simple wrapper | ai-interview-platform | Why it matters |
|---|---|---|---|
| Audio handling | Browser sends audio directly to an API | Custom WebSocket proxy, PCM handling, silence injection | The system can shape latency and speech boundaries instead of hoping the model behaves |
| Session recovery | Refresh usually means restart | Backend-owned session lifecycle with grace-period reconnect | Candidates do not lose context when the browser drops |
| Turn-taking control | Model decides when to answer | Server-side gate and silence pump manage the boundary | This reduces overlaps, echo, and awkward interruptions |
| Prompt governance | One prompt sets the tone | Prompt compiler plus system signal token | The backend keeps authority over the interview rules |
| Scoring | Free-form summary only | Deterministic skill deltas plus generated narrative | The output is easier to trust and compare |
| Multi-tenancy | Often absent | Tenant-scoped data model | Separate organizations can share the platform safely |
| Proctoring | Usually minimal | Face, window, and integrity signals are part of the stack | Assessment data is more defensible |
| Deployment maturity | Single-page demo | Docker, Redis, Sidekiq, Kubernetes manifests | The repo is shaped like infrastructure, not a prototype |
The table makes the distinction plain. This is not a thin UI on top of an LLM. It is a workflow system that uses generative models where they help and hard rules where they must.
Why the architecture feels unusually mature
There are signs of a system built for real deployment pressure: tenant scoping, background jobs, proctoring, and k8s manifests. Those pieces do not make the product glamorous, but they do make it believable.
That maturity shows up in the architecture choices too. A custom audio proxy, explicit reconnect handling, and downstream analysis are all signs that the team has thought past the first successful demo.
The result is a platform that sits between candidate-facing prep tools and recruiter-facing assessment systems. It is closer to assessment infrastructure than to interview theater.
What this project says about AI hiring tools
The bigger lesson is simple. If you want AI interviews to be useful, the question is not only what the model can say. It is what the platform can prevent, preserve, and prove.
That is why this repo stands out. It shows how to make a voice LLM act like a dependable assessment component without letting it control the session. The cage is the product.