gaze-coder: GazeCoder: Programming at the Speed of Sight

By syncing eye-tracking with LLM function calling, this experimental IDE replaces the mouse and keyboard with intent and focus.

• View on GitHub • More from louislva

A wide shot of a developer's silhouette with crosshatched beams of light extending from their eyes to a specific floating block of code.
GazeCoder envisions a hands-free development environment driven entirely by eye tracking and voice commands.

Key Takeaways

The Eye as a Pointer

You are looking at a bug. You press the spacebar, say "fix this," and release the key. The AI knows exactly which "this" you meant. No clicking, no highlighting, no scrolling.

The keyboard-less future of coding has long been a science fiction trope. GazeCoder turns it into a technical reality by treating the human eye as a high-bandwidth cursor. Built as a Next.js application, it combines WebGazer.js for eye tracking, Whisper for voice transcription, and GPT-4 for code generation.

How GazeCoder translates raw webcam coordinates into actionable semantic context in the DOM.

Solving the Saccade Struggle

Browser-based eye tracking is notoriously difficult to calibrate. Eyes are inherently jittery. They dart around in rapid movements called saccades. If an IDE simply passed your live gaze coordinates to an LLM while you spoke a command, the context window would be a chaotic mess of every line you glanced at.

GazeCoder solves this with a clever state management trick inside its React component tree. The interface relies on a custom hook that freezes the gaze focus the moment recording starts.

An extreme close-up of a human eye with a pause icon reflected in the pupil, staring at a single sharp line of code among blurred surroundings.
The Freeze-on-Record state prevents natural eye movements from shifting context during voice dictation.

When the user presses the spacebar, a hook locks the current gaze state. You can look away or read other functions while speaking your command. The system has already captured your initial intent.

The Multimodal Orchestrator

Most AI coding assistants operate on a selection-centric loop. You highlight text, right-click, and type a prompt. GazeCoder operates on a context-augmented generation loop.

The orchestration happens in a dedicated backend route that proxies requests to OpenAI. It bundles the visual context (the gazed-at lines), the audio context (the transcribed voice command), and environmental context (clipboard contents) into a single structured payload. It leverages OpenAI function calling to force the model to return a structured JSON response containing the exact code diff.

Probabilistic Quality Control

Traditional unit tests check for exact string matches or deterministic outputs. When your input is a voice command and your output is generated by an LLM, exact string matching becomes a brittle liability.

To safely develop this non-deterministic system, the developer built an AI-in-the-loop testing framework. GazeCoder extends the standard Jest testing library with a custom matcher.

"Please respond with <reasoning></reasoning> tags... and then finally a <grade></grade> tag"

This matcher sends the test output and a reference answer to an OpenAI model. It forces the model to output a step-by-step reasoning chain before issuing a final grade. It is a probabilistic safety net that ensures the semantic intent of the code remains intact even if the exact syntax varies.

DimensionTraditional AI IDEsGazeCoder
Input MethodKeyboard and MouseEye Tracking and Voice
Context SelectionManual Text HighlightingAutomatic Gaze Intersection
Cognitive LoadHigh (Context Switching)Low (Continuous Flow)
Testing ParadigmExact String MatchingSemantic Evaluation (toBeAsGoodAs)