Interview-Agent: How a Plain Web Stack Becomes a Convincing AI Interviewer

A close look at the resume-to-question pipeline, the browser-native voice loop, and the small state machine that turns mock interviews into a product.

8 min read • View on GitHub • More from Ankur909

A job candidate sits at a laptop while a resume stack feeds a question engine and a speaking interviewer appears on screen. The illustration explains how the project turns uploaded background material, browser speech, and timed video playback into a believable mock interview loop.
The trick is not a big avatar system. It is an orchestration layer that makes ordinary browser pieces feel continuous.
Key Takeaways

The most interesting thing in Interview-Agent is not that it runs mock interviews. It is that it makes a mock interview feel present tense with a surprisingly small stack. The avatar is not a special effects layer. It is a browser video that plays only when the AI is speaking.

The Interviewer Is Mostly Browser APIs

That sounds almost too simple, but the repo leans hard on native primitives. The question is spoken with speechSynthesis, the response is captured with webkitSpeechRecognition, and the face on screen is just a video element synced to the speech lifecycle. When the utterance starts, the video starts. When the utterance ends, the video stops.

A close-up of a small machine with three linked levers and a moving belt. One lever starts speech, one captures a spoken answer, and one toggles a video frame, showing how event hooks create the illusion of a live interviewer.
A tiny event loop does the work of a much heavier avatar stack.

The app works because data, speech, and playback are stitched into one loop rather than bolted together as separate features.

From Resume PDF to Personalized Questions

The personalization starts before the first question. In the backend, the app extracts text from an uploaded resume with pdfjs-dist, then sends that text into a structured LLM prompt that returns roles, skills, and projects. That turns a generic interview script into one that can actually reference the candidate’s background.

// High-level flow in the interview controller
const resumeText = await extractPdfText(file);
const profile = await askAi({
  system: 'Return structured JSON with roles, skills, and projects.',
  user: resumeText,
});

const questions = await generateQuestion(profile);
return res.json({ profile, questions });

AI powered interview agent to conduct initial rounds of technical interviews.

A Three-Step State Machine Keeps the Experience Tight

The product spine is a simple step flow. Step 1 sets up the interview, Step 2 runs the live Q and A loop, and Step 3 turns the session into a report. That structure matters because it keeps the UI focused and prevents the app from feeling like a chat window that wandered into video.

StepUser seesSystem does
1. SetupUpload resume and prepare the sessionParse PDF, build profile, generate questions
2. InterviewHear questions and answer out loudSpeak question, record response, sync avatar video
3. ReportReview scores and export resultsSummarize performance, render charts, generate PDF

Why the Interview Gets Harder Over Time

The difficulty curve is not an accident. The repo pushes candidates through a progression that starts easy and climbs toward harder questions, which makes the experience feel like an interview instead of a quiz. It is a small choice, but it changes the tone from generic testing to controlled pressure.

Design choiceWhat it optimizesTradeoff
Difficulty progressionA more realistic interview rhythmLess freedom for arbitrary question order
Credit gateA simple monetization boundaryAdds friction before the user even starts
Pre-generated question setPredictable control flowLess room for spontaneous follow-ups

What This Stack Buys You, and What It Costs

This is where the repo feels disciplined. Browser speech keeps costs low and the app easy to ship. Video playback tied to speech events gives the avatar a convincing rhythm. But the same choices also cap how natural the voice can sound, because quality depends on the browser and operating system the user has in front of them.

Interview-AgentHeavier interview platforms
Browser-native speech and local video triggersDedicated avatar, telephony, or streaming infrastructure
Narrow mock-interview scopeBroader enterprise assessment suite
Low setup complexityHigher deployment and integration burden
Fast iteration for a solo builderMore operational overhead and vendor lock-in

Where It Fits in the Interview-Tech Landscape

Against commercial tools like HackerRank, Glider AI, and TestGorilla, Interview-Agent makes a different bet. It does not try to be the system of record for hiring. It aims to be a focused, customizable interview simulator that feels polished without carrying enterprise weight.

That niche matters. A lot of interview software ends up either too barebones to feel useful or too enterprise-heavy to adopt casually. This repo sits in the middle: lightweight enough to build and change quickly, but opinionated enough to deliver a coherent experience.


Why This Little Project Matters

Interview-Agent is a case study in product minimalism. The interface feels richer than the architecture because the orchestration is doing the work: state transitions, speech events, resume parsing, and report generation all reinforce one another. The lesson is not that every interview tool should be this small. It is that a small stack can still create a convincing product when the timing is right.