Decoding the Shaky Hand: How Assistant-for-stage-fear Computes Confidence

A look inside the multimodal pipeline turning raw video frames into a real-time anxiety heatmap.

6 min read • View on GitHub • More from Siya-05

A wide shot of a glass-walled stage. Inside, a speaker made of glowing green geometric wireframes stands before a dark, looming cloud of data points.
The geometry of performance anxiety: mapping the human body as a series of trackable coordinates.
Key Takeaways

The Geometry of a Nervous Glance

Public speaking anxiety is notoriously subjective. A human coach reads the room, senses tension, and offers intuitive advice. The Assistant-for-stage-fear project takes a different approach: it mechanizes empathy. By translating "stage fright" into hard geometric ratios, it attempts to quantify the invisible physical symptoms of anxiety using standard webcams and off-the-shelf AI models.

The core logic lies in its eye-tracking and fidget-detection algorithms. Instead of relying on vague behavioral cues, the system uses facial landmarks to calculate "audience engagement" as a mathematical ratio. The is_looking_away function evaluates the distance of the iris relative to the eye's bounding box. If the gaze vector deviates beyond a specific threshold, a boolean flag is tripped. The speaker is losing the room.

A medium shot of a person’s face, with their hand approaching their chin. A translucent 'Warning' perimeter is drawn around the face in ink.
Fidgets are quantified by measuring the proximity of hand landmarks to the facial bounding box.

Similarly, "fidgeting" is defined not by a psychological state, but by physical proximity. By tracking hand landmarks via MediaPipe, the script monitors how close the speaker's hands get to their face. It is a clinical, geometric interpretation of a deeply human feeling.

The Three-Layered Brain

Processing a human being in real-time requires a multimodal stack. The project achieves this by dividing the workload across three distinct AI models acting as the system's "brain."

MediaPipe acts as the eyes, mapping the spatial coordinates of the speaker's body and face. OpenAI's Whisper acts as the ears, transcribing the audio stream into text and surfacing filler words. Finally, Google's Gemini acts as the judgment layer, ingesting the transcript and the physical telemetry to generate contextual coaching.

A flow diagram showing 'The Anatomy of a Feedback Loop'. It starts with Camera/Mic Input on the left. The flow splits: Video goes to MediaPipe (Pose/Face)

The technical challenge here is synchronization. Running synchronous capture for audio and video in Python often leads to blocked threads and stuttering frames. To prevent the heavy I/O of audio processing from halting the video feed, the system leverages multi-threading, keeping the visual capture loop running smoothly while Whisper chews through the audio buffer.

Beyond the "Um" Counter

Basic speech analytics tools have existed for years, primarily functioning as glorified metronomes and word counters. They can tell a speaker their Words Per Minute (WPM) or how many times they said "um," but they lack contextual awareness.

This project attempts to correlate vocal crutches with physical tells. A fast speaking rate combined with strong eye contact might indicate passion; a fast speaking rate combined with downward gaze and facial touching indicates panic. By fusing these data streams, the feedback evolves from simple heuristics into holistic behavioral coaching.

Legacy Speech Tools Multimodal Assistants
Counts filler words ("um", "ah") via simple regex. Correlates filler words with physical distress signals.
Measures raw Words Per Minute (WPM). Evaluates pacing alongside emotional sentiment analysis.
Audio-only analysis; oblivious to body language. Tracks gaze deviation and hand-to-face proximity.
Provides static, rule-based reports. Uses LLMs for contextual, conversational feedback.

From Notebook to Logic

The architectural evolution of the repository offers a raw look at how AI prototypes are built today. The project began as an experimental Jupyter notebook, effectively a sandbox for isolating and testing the heavy-lifting libraries like OpenCV and MediaPipe.

A close-up of a typewriter where the paper is a film strip. As the keys hit, they stamp both words and tiny icons of 'eyes' or 'hands'.
The system weaves discrete audio transcripts and continuous video frames into a single diagnostic thread.

Once the baseline heuristics were established—such as using regex to catch filler words and calculating speech rates—the logic migrated into a standalone Python script. This transition from interactive cells to a threaded application marks the shift from passive analysis to real-time, low-latency feedback.

While still in its early stages, the repository illustrates a broader shift in EdTech and mental wellness tools. It moves away from post-game analytics and toward live exposure therapy, providing a judgment-free, geometric mirror for human anxiety.


Sources: Siya-05/Assistant-for-stage-fear.