Voice & Multimodal AI

Real-time voice companions, multimodal agents, speech-to-speech, and audio-first interfaces

145 explainers
kids-storybook-generator: Kids Storybook Generator: The AI Story App That Keeps Working When the Models Fail
Voice & Multimodal AI
kids-storybook-generator: Kids Storybook Generator: The AI Story App That Keeps Working When the Models Fail
A surprisingly mature full-stack project that combines fallback templates, multi-provider image generation, and model orchestration into one resilient storybook pipeline.
8 min read
AI-Based-Oral-Health-Analysis-System-YOLOv8-: AI-Based Oral Health Analysis System: The Split-Brain Architecture That Turns YOLOv8 Into a Dental Workflow
Voice & Multimodal AI
AI-Based-Oral-Health-Analysis-System-YOLOv8-: AI-Based Oral Health Analysis System: The Split-Brain Architecture That Turns YOLOv8 Into a Dental Workflow
A Node.js orchestrator, a Flask inference engine, and structured prediction records combine into a surprisingly disciplined blueprint for accessible medical imaging.
8 min read
snap-bite: SnapBite: The AI Food Scanner That Treats Uncertainty as a First-Class Feature
Voice & Multimodal AI
snap-bite: SnapBite: The AI Food Scanner That Treats Uncertainty as a First-Class Feature
A close look at the React Native, Expo, and Supabase stack behind a nutrition app that turns camera input into structured macro data, then gives users the final word.
8 min read
Nodos Plugins: The Real-Time Graph That Refuses to Sleep
Voice & Multimodal AI
Nodos Plugins: The Real-Time Graph That Refuses to Sleep
How a schema-driven plugin system, GPU-first execution, and always-dirty nodes turn Nodos into a live media engine instead of a static dependency graph.
8 min read
clawtotalk: the shortest path from a self-hosted agent to your phone
Voice & Multimodal AI
clawtotalk: the shortest path from a self-hosted agent to your phone
A push-to-talk companion that makes OpenClaw feel portable without opening a public front door.
7 min read
openclaw-textclaw: The iMessage bridge that lets OpenClaw run without a Mac
Voice & Multimodal AI
openclaw-textclaw: The iMessage bridge that lets OpenClaw run without a Mac
openclaw-textclaw strips Apple messaging down to a relay problem, then solves it with a lean Node.js adapter, a persistent WebSocket, and just enough state to keep replies natural.
9 min read
sendblue-openclaw-relay: How a tiny Python bridge makes OpenClaw feel native in iMessage
Voice & Multimodal AI
sendblue-openclaw-relay: How a tiny Python bridge makes OpenClaw feel native in iMessage
A public webhook, a private WebSocket, and ffmpeg in the middle. This repo shows that chatbot UX is really a delivery problem dressed up as AI.
7 min read
The AI Hangover: Unpacking Avik-Jain.github.io
Voice & Multimodal AI
The AI Hangover: Unpacking Avik-Jain.github.io
How an ambitious voice-controlled portfolio hit the limits of the generative era and retreated to static code.
6 min read
last_year_project: The Fitness Coach That Thinks in Angles, States, and Timers
Voice & Multimodal AI
last_year_project: The Fitness Coach That Thinks in Angles, States, and Timers
A deep dive into an AI exercise system that keeps pose tracking local, turns motion into finite-state logic, and uses simple math to do surprisingly disciplined coaching.
10 min read
Sign-Language-Translator: When a Webcam Learns to Ignore Everything Except the Hand
Voice & Multimodal AI
Sign-Language-Translator: When a Webcam Learns to Ignore Everything Except the Hand
A Python vision pipeline that turns finger-spelling into a clean skeletal signal, then pushes it through a recognition stack built for real-world clutter, not lab-perfect backgrounds.
7 min read
Online-Tech-Hiring-Platform: PeerChat: The Technical Hiring Platform That Treats Interviews Like Server-Captured Media
Voice & Multimodal AI
Online-Tech-Hiring-Platform: PeerChat: The Technical Hiring Platform That Treats Interviews Like Server-Captured Media
A close look at the Polyglot stack, the Mediasoup recording pipeline, and the workflow logic that turns a video call into a durable hiring system.
9 min read
MediConnect: The Hospital App That Treats Scheduling Like a Clinical System
Voice & Multimodal AI
MediConnect: The Hospital App That Treats Scheduling Like a Clinical System
A deep dive into a React and Node platform where appointments, records, and AI revolve around one idea: healthcare only works when state changes are handled carefully.
8 min read
ai-mock-interview-platform: AI Mock Interview Platform: The App That Grades Candidates by Surviving LLM Failure
Voice & Multimodal AI
ai-mock-interview-platform: AI Mock Interview Platform: The App That Grades Candidates by Surviving LLM Failure
A full-stack interview coach that chains fallback models, normalizes messy AI output, and scores both technical answers and STAR-style behavioral responses.
8 min read
Ink-Well Turns SVG Paths Into InkML, One Stroke at a Time
Voice & Multimodal AI
Ink-Well Turns SVG Paths Into InkML, One Stroke at a Time
A tiny TypeScript bridge that translates browser drawing commands into portable ink traces for handwriting, signatures, and sketching workflows.
5 min read
Interview-Agent: How a Plain Web Stack Becomes a Convincing AI Interviewer
Voice & Multimodal AI
Interview-Agent: How a Plain Web Stack Becomes a Convincing AI Interviewer
A close look at the resume-to-question pipeline, the browser-native voice loop, and the small state machine that turns mock interviews into a product.
8 min read
PrepBuddy-frontend: The Mock Interview App That Treats Voice Like a First-Class Feature
Voice & Multimodal AI
PrepBuddy-frontend: The Mock Interview App That Treats Voice Like a First-Class Feature
A lean React SPA turns interview prep into a live simulation, using Whisper-powered transcription, sequential evaluation, and a deliberately simple architecture to make practice feel immediate.
8 min read
EduLinguaBridge: The Classroom Translator Built Entirely in the Browser
Voice & Multimodal AI
EduLinguaBridge: The Classroom Translator Built Entirely in the Browser
A no-backend edtech prototype that uses speech recognition, local storage, PDF parsing, and phrase-first translation to help multilingual students keep up without sending data to the cloud.
8 min read
Real-Time-Malpractice-Detection-in-classroom: Real-Time Malpractice Detection in Classroom: How a Webcam Becomes an Evidence Machine
Voice & Multimodal AI
Real-Time-Malpractice-Detection-in-classroom: Real-Time Malpractice Detection in Classroom: How a Webcam Becomes an Evidence Machine
A lightweight Flask and OpenCV system does more than spot inattentiveness. It logs behavior, captures proof, and turns a simple face detector into a classroom audit trail.
8 min read
InterviewIQ: The Mock Interview App That Makes a Browser Feel Like a Human Interviewer
Voice & Multimodal AI
InterviewIQ: The Mock Interview App That Makes a Browser Feel Like a Human Interviewer
A React, Node, and MongoDB stack that turns resumes into structured interviews, simulates a talking interviewer with native browser APIs, and packages the result as a scored PDF report.
10 min read
ai-travel-companion-platform, the AI trip planner that refuses to trust its own output
Voice & Multimodal AI
ai-travel-companion-platform, the AI trip planner that refuses to trust its own output
Gemini drafts the itinerary, OpenStreetMap grounds it in reality, and the app lets you regenerate one day or one activity without blowing up the whole trip.
9 min read
skyview: SkyTrack: The Flight Tracker Where the Backend Takes the Blame
Voice & Multimodal AI
skyview: SkyTrack: The Flight Tracker Where the Backend Takes the Blame
A real-time aviation dashboard that turns noisy ADS-B state vectors into a calm, map-first experience by caching, normalizing, and shielding the UI from upstream failure.
9 min read
NexChat: The Spring Boot Chat Stack That Treats Messaging, Calls, and AI Summaries as One System
Voice & Multimodal AI
NexChat: The Spring Boot Chat Stack That Treats Messaging, Calls, and AI Summaries as One System
A real-time Java backend that uses WebSockets for chat, WebRTC signaling for calls, and Gemini for conversation summaries, all wrapped in JWT security and rate limits.
8 min read
Mind-Mend-App: Mind Mend AI: When a Mood Journal Becomes a Triage Engine
Voice & Multimodal AI
Mind-Mend-App: Mind Mend AI: When a Mood Journal Becomes a Triage Engine
A privacy-first wellness app that encrypts journal entries, routes AI through Supabase Edge Functions, and turns emotional check-ins into targeted therapeutic actions.
8 min read
Crop-Cure: The AI crop doctor that knows when to think and when to rule
Voice & Multimodal AI
Crop-Cure: The AI crop doctor that knows when to think and when to rule
A full-stack farm management system that pairs image-based disease diagnosis with plain-spoken irrigation advice, built for farmers who may have a phone number, a leaf photo, and not much else.
9 min read
HireSense-AI: The Mock Interview Room That Syncs Code, Video, and Judgment in Real Time
Voice & Multimodal AI
HireSense-AI: The Mock Interview Room That Syncs Code, Video, and Judgment in Real Time
A deep dive into a full-stack collaboration app that wires together Socket.io, PeerJS, Judge0, and Gemini to turn interview prep into a live, shared workflow.
8 min read
KhataAI-Project: KhataAI: Turning WhatsApp Voice Notes Into a Working Credit Ledger
Voice & Multimodal AI
KhataAI-Project: KhataAI: Turning WhatsApp Voice Notes Into a Working Credit Ledger
An AI-powered Khata system for Kirana shops that combines Hinglish transcription, fuzzy customer matching, and rule-based risk scoring so informal credit can be tracked without forcing shopkeepers into software-shaped behavior.
8 min read
promptHire-ai-interview-mocker: PromptHire: The AI Interview Mocker That Turns Your Browser Into the Hiring Panel
Voice & Multimodal AI
promptHire-ai-interview-mocker: PromptHire: The AI Interview Mocker That Turns Your Browser Into the Hiring Panel
A close look at the browser-native loop behind PromptHire, where Gemini generates the questions, speech APIs run the interview, and a feedback table turns every answer into a reviewable lesson.
8 min read
IntervYou-MERN-Stack-project: IntervYou: The MERN Interview App That Grades Your Answers and Watches Your Face
Voice & Multimodal AI
IntervYou-MERN-Stack-project: IntervYou: The MERN Interview App That Grades Your Answers and Watches Your Face
A technical look at a prototype that splits evaluation across browser-side emotion tracking, cloud LLM scoring, and a hard three-strikes proctoring system.
8 min read
`leaf-diseases-detect`: When a Plant Doctor Is Just a Prompt
Voice & Multimodal AI
`leaf-diseases-detect`: When a Plant Doctor Is Just a Prompt
How a small Python app turns a vision LLM into a leaf triage pipeline, using a strict JSON contract, an invalid-image guardrail, and a fast API wrapper around a surprisingly practical workflow.
7 min read
EasyBudgetAI: SmartKhata: The AI Ledger That Understands Hinglish, Khata, and Human Sloppiness
Voice & Multimodal AI
EasyBudgetAI: SmartKhata: The AI Ledger That Understands Hinglish, Khata, and Human Sloppiness
A deep look at how EasyBudgetAI turns chat and voice into structured expenses, matches fuzzy Indian names, and uses Redis to keep AI costs, abuse, and security under control.
10 min read
PDF_to_Podcast_Generator Turns PDFs Into a Two-Voice Audio Pipeline
Voice & Multimodal AI
PDF_to_Podcast_Generator Turns PDFs Into a Two-Voice Audio Pipeline
A Streamlit app that preserves document structure, scripts a conversational episode, and pushes the result through fast LLM summarization and neural text-to-speech.
7 min read
aiecho-react-chat: AI Echo React Chat: A Debugger for Gemini’s Hidden Conversation
Voice & Multimodal AI
aiecho-react-chat: AI Echo React Chat: A Debugger for Gemini’s Hidden Conversation
A local React viewer that reconstructs turns, reveals reasoning on demand, and turns Gemini JSON exports into a navigable, citation-aware transcript.
7 min read
AI_surveillance: The Tiny Geometry Trick That Turns Detections Into Security Events
Voice & Multimodal AI
AI_surveillance: The Tiny Geometry Trick That Turns Detections Into Security Events
A multi-service entrance monitor that goes beyond bounding boxes, using line-crossing state, snapshot evidence, and delayed notifications to convert vision into actionable alerts.
8 min read
yukkuri-video-generator: The Repo That Builds Videos by Listening First
Voice & Multimodal AI
yukkuri-video-generator: The Repo That Builds Videos by Listening First
A Python pipeline for Yukkuri commentary that turns script text into rendered MP4s only after the voice exists, the duration is known, and the timeline can be locked to reality.
8 min read
exitZER0: The P2P Bridge That Turns a Phone Into an AI Agent’s Remote Sense-Act Loop
Voice & Multimodal AI
exitZER0: The P2P Bridge That Turns a Phone Into an AI Agent’s Remote Sense-Act Loop
A Rust-and-Swift system that pairs iPhone, Mac, and LLM backends directly over Iroh, keeps the connection alive through flaky networks, and treats the phone as more than a chat window.
11 min read
SPOTTO: The parking app that turns streets into a probability map
Voice & Multimodal AI
SPOTTO: The parking app that turns streets into a probability map
A Flutter prototype that predicts where you can stop, redraws the map around your position, and hands off to navigation only when the decision is made.
7 min read
celo-personality-miniapp Is Almost Empty, and That Is the Thesis
Voice & Multimodal AI
celo-personality-miniapp Is Almost Empty, and That Is the Thesis
A tiny Celo MiniApp repo that hints at a bigger shift: mobile-native distribution, onchain identity, and personality as a product surface.
6 min read
cafeg: The Android app that latches onto live audio sessions
Voice & Multimodal AI
cafeg: The Android app that latches onto live audio sessions
A foreground service, a broadcast receiver, and Android’s audio effects API combine to add binaural virtualization and reverb directly onto playback sessions.
7 min read
SkipBackRecorder: The Recorder That Starts Before You Hit Record
Voice & Multimodal AI
SkipBackRecorder: The Recorder That Starts Before You Hit Record
A Windows desktop app in Python that keeps a live audio buffer, then stitches the past onto the present the moment you decide to save.
7 min read
cafem: CaféTone makes Android feel like a premium audio device
Voice & Multimodal AI
cafem: CaféTone makes Android feel like a premium audio device
Instead of owning playback, this Kotlin app latches onto audio sessions, keeps a foreground service alive, and slips Virtualizer and Reverb into the signal path.
9 min read
playai-gradio: The Glue Code for the Multimodal Era
Voice & Multimodal AI
playai-gradio: The Glue Code for the Multimodal Era
How a minimalist Python bridge uses the Gradio registry pattern to turn complex streaming text-to-speech APIs into a single composable UI block.
6 min read
moltys-run: The Architecture of the Ubiquitous Agent: Inside OpenClaw
Voice & Multimodal AI
moltys-run: The Architecture of the Ubiquitous Agent: Inside OpenClaw
How a local-first daemon turned everyday messaging apps into an action-oriented AI control plane.
9 min read
RapRecord: The Zero-Latency Bridge from YouTube to the Booth
Voice & Multimodal AI
RapRecord: The Zero-Latency Bridge from YouTube to the Booth
How a minimalist TypeScript monorepo turns the world's largest video library into a high-speed scratchpad for rappers.
era2: Joxy and the Art of the Elegant Exit
Voice & Multimodal AI
era2: Joxy and the Art of the Elegant Exit
How evinjohnn/era2 uses sub-second RAG and state-driven handoffs to bridge the gap between AI assistance and luxury retail.
galando/alexa-skill: The Architecture of a Silent Conversation
Voice & Multimodal AI
galando/alexa-skill: The Architecture of a Silent Conversation
How a minimalist Kotlin implementation decoupled the "Voice-First" monolith and brought backend discipline to the Alexa Skills Kit.
Prompt-Engineering the Human Voice: Inside louislva/read
Voice & Multimodal AI
Prompt-Engineering the Human Voice: Inside louislva/read
How a minimalist Chrome extension uses multimodal LLMs to turn static web text into a performance with a thick Kiwi accent.
Wayfarer: An AI Travel Companion That Actually Picks Up the Phone
Voice & Multimodal AI
Wayfarer: An AI Travel Companion That Actually Picks Up the Phone
By bridging the gap between LLM reasoning and the global telephone network, Wayfarer transforms the travel assistant from a search engine into a functional concierge.
8 min read
Decoding the Shaky Hand: How Assistant-for-stage-fear Computes Confidence
Voice & Multimodal AI
Decoding the Shaky Hand: How Assistant-for-stage-fear Computes Confidence
A look inside the multimodal pipeline turning raw video frames into a real-time anxiety heatmap.
6 min read