InboundAIVoice: The Open-Source Voice Stack Built for Indian Calls

A production-minded AI phone framework that mixes LiveKit, Sarvam AI, SIP telephony, and a browser control panel to automate booking, support, and lead qualification in Hinglish and regional languages.

9 min read • View on GitHub • More from toprmrproducer

A wide black-ink editorial scene of an Indian office reception desk at night. A ringing phone connects into a visible telephony line, while a control surface behind it combines a calendar grid and a browser dashboard with voice settings. The image explains that this project is not a demo chatbot, but an operating stack for phone workflows.
InboundAIVoice bundles telephony, scheduling, and operator control into one self-hosted stack for Indian businesses.

Most people building AI voice agents never leave Level 1. They stay on VAPI, bleed 5 cents a minute, and never find out their own data might be getting sold. I've spent 1000+ hours building voice agents and deployed them for million-dollar companies. This is the exact 3-level path — VAPI to your own pipeline to a productized SaaS.

Shreyas Raj, Founder, RapidXAI · AI Voice Agency: The 3 Levels of Scaling
Key Takeaways

Why this project exists

Most voice agent frameworks are built to prove a point. InboundAIVoice is built to run a business. It focuses on Indian phone workflows, where code-switching, regional accents, local telephony, and fast booking logic matter more than a polished demo.

That is the real wedge here. The repository is not trying to be the most general voice platform. It is trying to be the most useful one for agencies and operators who need calls answered, leads qualified, and appointments booked without handing the whole margin to a managed SaaS.

A WSJ hedcut-style portrait of Shreyas Raj based on his verified GitHub avatar at https://avatars.githubusercontent.com/u/167063974?v=4. The portrait is meant to identify the maintainer behind the project and ground the article in a real public contributor.

The unusual part is not the model. It is the localization.

The stack leans on Sarvam AI, Vobiz SIP trunks, and LiveKit Agents, but the differentiator is cultural. The agent is tuned for Hinglish and Indian regional language behavior, which changes the whole product. It is not just about transcription quality. It is about how people actually speak on the phone in India.

That shows up in the configuration layer too. The project uses per-client JSON configs and language presets, so one call can sound formal, another can feel conversational, and another can be tuned for a specific language mix without rewriting the agent.

A close-up black-ink diagram of a call-routing switchboard. One phone line enters from the left and branches into three guarded paths: language preset, calendar check, and logging plus notifications. A small firewall-like gate sits in front of the fallback path. The image explains that the system is built to adapt the call path, not just respond to speech.
Localization is not a skin on top of the stack. It shapes routing, prompts, and the operator experience.

This flow shows how one call becomes a localized business event, not just a conversation.

Inside the call path

The core flow starts in make_call.py. That script creates the call, hands metadata into a LiveKit session, and dispatches the outbound agent. From there, the agent handles the real-time conversation while SIP trunks connect the phone network to the browser-based orchestration layer.

# High-level call flow
room = create_livekit_room(phone_number)
metadata = {
    'phone_number': phone_number,
    'client_config': load_config(phone_number),
    'language': 'hinglish'
}
start_agent(room=room, metadata=metadata)
bridge_sip_to_livekit(room)

The call is not treated as a dead-end text exchange. It becomes a live session with language selection, voice settings, and operational context attached from the start. That is what lets the project behave like infrastructure instead of a chatbot wrapper.

Why the booking flow is the quietest impressive part

The booking layer looks simple until you inspect it. In calendar_tools.py, the repo abstracts both Cal.com and Google Calendar, then computes availability with time-zone-aware logic for IST. That matters because phone scheduling fails in small ways before it fails in big ones.

The system does not trust a calendar API to do the entire job. It fetches busy ranges, calculates open windows, and keeps the booking logic deterministic. That is exactly the kind of boring engineering that makes a voice agent usable in production.

Calendar pathHow it behavesWhy it matters
Cal.comUnified interface for scheduling operationsKeeps one booking surface across clients
Google CalendarManual slot calculation from busy rangesAvoids timezone mistakes and overbooking
IST handlingExplicit Asia/Kolkata contextMatches the business reality of Indian callers

The code is defensive in ways that matter

This is where InboundAIVoice stops looking like a prototype. The database layer in db.py includes schema fallback behavior and retry logic so transcripts are not lost if migrations lag behind the application. The system would rather save partial data than fail silently.

The same bias shows up in rate limiting and notifications. The agent guards against repeated calls, while notify.py pushes alerts through Telegram and WhatsApp so humans do not have to babysit the stack. That combination is modest, practical, and exactly right for a small production deployment.

Toy voice agentProduction voice agent
Assumes perfect schemaFalls back when columns are missing
Logs only if everything worksRetries inserts and preserves transcripts
Conversation is the productConversation is one step in an operations flow
No safety against loopsRate limits repeated calls
One notification pathOwner and customer alerts through multiple channels

What it competes with, and where it wins

Against Vapi and Retell AI, InboundAIVoice loses on convenience. Those platforms are faster to start. Against raw LiveKit Agents or Pipecat, it loses on generality. But that is not the right comparison. This repo is optimized for ownership, Indian telephony, and agency economics.

ProjectDeployment modelLanguage localizationControl panelSelf-hostingBest fit
InboundAIVoiceSelf-hosted Python stackHinglish and regional presetsFastAPI browser UIYesIndian voice agencies
VapiManaged SaaSGeneral-purposeHosted UINoFast prototyping
Retell AIManaged SaaSGeneral-purposeHosted UINoTeams wanting low ops
PipecatOpen-source frameworkDepends on integrationDIYYesBuilders who want flexibility
Raw LiveKit AgentsFramework primitivesDepends on your codeDIYYesTeams building from scratch

The distinction is simple. InboundAIVoice is not trying to win on breadth. It is trying to win on fit. If you need a stack that understands Indian callers, owns the phone path, and lets you keep the margin, the project makes a strong case.

The UI Server allows you to easily configure your agent's prompt, select voices/models, and securely paste your API keys without touching code.

Shreyas Raj, Maintainer · InboundAIVoice GitHub README

The trade-off

The architecture is strong, but it is still opinionated. Local JSON config, in-memory rate limiting, and a deployment pattern that favors a single instance all point to the same reality. This is production-ready beta infrastructure, not a globally distributed multi-tenant platform.

That is not a weakness so much as a boundary. The repo solves one important class of problem very well: localized, operationally useful voice automation for Indian businesses. For the audience it targets, that is enough to matter.