InboundAIVoice: The Open-Source Voice Stack Built for Indian Calls
A production-minded AI phone framework that mixes LiveKit, Sarvam AI, SIP telephony, and a browser control panel to automate booking, support, and lead qualification in Hinglish and regional languages.

Most people building AI voice agents never leave Level 1. They stay on VAPI, bleed 5 cents a minute, and never find out their own data might be getting sold. I've spent 1000+ hours building voice agents and deployed them for million-dollar companies. This is the exact 3-level path — VAPI to your own pipeline to a productized SaaS.
- InboundAIVoice turns AI calling into an operational stack for Indian businesses, not just a conversational demo.
- Its strongest differentiator is localization, with Hinglish and regional-language handling built into the call experience.
- The repository is production-minded because it treats booking, logging, alerts, and failure recovery as first-class parts of the system.
- The project wins on control and margin economics, but it trades away the convenience of a fully managed platform.
Why this project exists
Most voice agent frameworks are built to prove a point. InboundAIVoice is built to run a business. It focuses on Indian phone workflows, where code-switching, regional accents, local telephony, and fast booking logic matter more than a polished demo.
That is the real wedge here. The repository is not trying to be the most general voice platform. It is trying to be the most useful one for agencies and operators who need calls answered, leads qualified, and appointments booked without handing the whole margin to a managed SaaS.
The unusual part is not the model. It is the localization.
The stack leans on Sarvam AI, Vobiz SIP trunks, and LiveKit Agents, but the differentiator is cultural. The agent is tuned for Hinglish and Indian regional language behavior, which changes the whole product. It is not just about transcription quality. It is about how people actually speak on the phone in India.
That shows up in the configuration layer too. The project uses per-client JSON configs and language presets, so one call can sound formal, another can feel conversational, and another can be tuned for a specific language mix without rewriting the agent.
Inside the call path
The core flow starts in make_call.py. That script creates the call, hands metadata into a LiveKit session, and dispatches the outbound agent. From there, the agent handles the real-time conversation while SIP trunks connect the phone network to the browser-based orchestration layer.
# High-level call flow
room = create_livekit_room(phone_number)
metadata = {
'phone_number': phone_number,
'client_config': load_config(phone_number),
'language': 'hinglish'
}
start_agent(room=room, metadata=metadata)
bridge_sip_to_livekit(room)
The call is not treated as a dead-end text exchange. It becomes a live session with language selection, voice settings, and operational context attached from the start. That is what lets the project behave like infrastructure instead of a chatbot wrapper.
Why the booking flow is the quietest impressive part
The booking layer looks simple until you inspect it. In calendar_tools.py, the repo abstracts both Cal.com and Google Calendar, then computes availability with time-zone-aware logic for IST. That matters because phone scheduling fails in small ways before it fails in big ones.
The system does not trust a calendar API to do the entire job. It fetches busy ranges, calculates open windows, and keeps the booking logic deterministic. That is exactly the kind of boring engineering that makes a voice agent usable in production.
| Calendar path | How it behaves | Why it matters |
|---|---|---|
| Cal.com | Unified interface for scheduling operations | Keeps one booking surface across clients |
| Google Calendar | Manual slot calculation from busy ranges | Avoids timezone mistakes and overbooking |
| IST handling | Explicit Asia/Kolkata context | Matches the business reality of Indian callers |
The code is defensive in ways that matter
This is where InboundAIVoice stops looking like a prototype. The database layer in db.py includes schema fallback behavior and retry logic so transcripts are not lost if migrations lag behind the application. The system would rather save partial data than fail silently.
The same bias shows up in rate limiting and notifications. The agent guards against repeated calls, while notify.py pushes alerts through Telegram and WhatsApp so humans do not have to babysit the stack. That combination is modest, practical, and exactly right for a small production deployment.
| Toy voice agent | Production voice agent |
|---|---|
| Assumes perfect schema | Falls back when columns are missing |
| Logs only if everything works | Retries inserts and preserves transcripts |
| Conversation is the product | Conversation is one step in an operations flow |
| No safety against loops | Rate limits repeated calls |
| One notification path | Owner and customer alerts through multiple channels |
What it competes with, and where it wins
Against Vapi and Retell AI, InboundAIVoice loses on convenience. Those platforms are faster to start. Against raw LiveKit Agents or Pipecat, it loses on generality. But that is not the right comparison. This repo is optimized for ownership, Indian telephony, and agency economics.
| Project | Deployment model | Language localization | Control panel | Self-hosting | Best fit |
|---|---|---|---|---|---|
| InboundAIVoice | Self-hosted Python stack | Hinglish and regional presets | FastAPI browser UI | Yes | Indian voice agencies |
| Vapi | Managed SaaS | General-purpose | Hosted UI | No | Fast prototyping |
| Retell AI | Managed SaaS | General-purpose | Hosted UI | No | Teams wanting low ops |
| Pipecat | Open-source framework | Depends on integration | DIY | Yes | Builders who want flexibility |
| Raw LiveKit Agents | Framework primitives | Depends on your code | DIY | Yes | Teams building from scratch |
The distinction is simple. InboundAIVoice is not trying to win on breadth. It is trying to win on fit. If you need a stack that understands Indian callers, owns the phone path, and lets you keep the margin, the project makes a strong case.

The UI Server allows you to easily configure your agent's prompt, select voices/models, and securely paste your API keys without touching code.
The trade-off
The architecture is strong, but it is still opinionated. Local JSON config, in-memory rate limiting, and a deployment pattern that favors a single instance all point to the same reality. This is production-ready beta infrastructure, not a globally distributed multi-tenant platform.
That is not a weakness so much as a boundary. The repo solves one important class of problem very well: localized, operationally useful voice automation for Indian businesses. For the audience it targets, that is enough to matter.