Decoding the Shaky Hand: How Assistant-for-stage-fear Computes Confidence
A look inside the multimodal pipeline turning raw video frames into a real-time anxiety heatmap.
- The system quantifies stage fright by translating behavioral cues like gaze deviation and hand-to-face proximity into geometric ratios.
- A multimodal pipeline uses MediaPipe for spatial tracking, Whisper for transcription, and Gemini for contextual coaching.
- Multi-threading prevents audio-visual synchronization lag by separating heavy I/O processing from the real-time video capture loop.
- The tool moves beyond simple word counters by correlating vocal fillers with physical tells to provide holistic behavioral analysis.
The Geometry of a Nervous Glance
Public speaking anxiety is notoriously subjective. A human coach reads the room, senses tension, and offers intuitive advice. The Assistant-for-stage-fear project takes a different approach: it mechanizes empathy. By translating "stage fright" into hard geometric ratios, it attempts to quantify the invisible physical symptoms of anxiety using standard webcams and off-the-shelf AI models.
The core logic lies in its eye-tracking and fidget-detection algorithms. Instead of relying on vague behavioral cues, the system uses facial landmarks to calculate "audience engagement" as a mathematical ratio. The is_looking_away function evaluates the distance of the iris relative to the eye's bounding box. If the gaze vector deviates beyond a specific threshold, a boolean flag is tripped. The speaker is losing the room.
Similarly, "fidgeting" is defined not by a psychological state, but by physical proximity. By tracking hand landmarks via MediaPipe, the script monitors how close the speaker's hands get to their face. It is a clinical, geometric interpretation of a deeply human feeling.
The Three-Layered Brain
Processing a human being in real-time requires a multimodal stack. The project achieves this by dividing the workload across three distinct AI models acting as the system's "brain."
MediaPipe acts as the eyes, mapping the spatial coordinates of the speaker's body and face. OpenAI's Whisper acts as the ears, transcribing the audio stream into text and surfacing filler words. Finally, Google's Gemini acts as the judgment layer, ingesting the transcript and the physical telemetry to generate contextual coaching.
The technical challenge here is synchronization. Running synchronous capture for audio and video in Python often leads to blocked threads and stuttering frames. To prevent the heavy I/O of audio processing from halting the video feed, the system leverages multi-threading, keeping the visual capture loop running smoothly while Whisper chews through the audio buffer.
Beyond the "Um" Counter
Basic speech analytics tools have existed for years, primarily functioning as glorified metronomes and word counters. They can tell a speaker their Words Per Minute (WPM) or how many times they said "um," but they lack contextual awareness.
This project attempts to correlate vocal crutches with physical tells. A fast speaking rate combined with strong eye contact might indicate passion; a fast speaking rate combined with downward gaze and facial touching indicates panic. By fusing these data streams, the feedback evolves from simple heuristics into holistic behavioral coaching.
| Legacy Speech Tools | Multimodal Assistants |
|---|---|
| Counts filler words ("um", "ah") via simple regex. | Correlates filler words with physical distress signals. |
| Measures raw Words Per Minute (WPM). | Evaluates pacing alongside emotional sentiment analysis. |
| Audio-only analysis; oblivious to body language. | Tracks gaze deviation and hand-to-face proximity. |
| Provides static, rule-based reports. | Uses LLMs for contextual, conversational feedback. |
From Notebook to Logic
The architectural evolution of the repository offers a raw look at how AI prototypes are built today. The project began as an experimental Jupyter notebook, effectively a sandbox for isolating and testing the heavy-lifting libraries like OpenCV and MediaPipe.
Once the baseline heuristics were established—such as using regex to catch filler words and calculating speech rates—the logic migrated into a standalone Python script. This transition from interactive cells to a threaded application marks the shift from passive analysis to real-time, low-latency feedback.
While still in its early stages, the repository illustrates a broader shift in EdTech and mental wellness tools. It moves away from post-game analytics and toward live exposure therapy, providing a judgment-free, geometric mirror for human anxiety.
Sources: Siya-05/Assistant-for-stage-fear.