gaze-coder: GazeCoder: Programming at the Speed of Sight
By syncing eye-tracking with LLM function calling, this experimental IDE replaces the mouse and keyboard with intent and focus.
- GazeCoder eliminates manual text selection by using eye-tracking coordinates to define the context for LLM prompts.
- A state management hook freezes the gaze focus at the start of a voice command to prevent natural eye jitters from breaking the context.
- The system uses OpenAI function calling to transform multimodal inputs into structured JSON code diffs.
- An AI-in-the-loop testing framework replaces traditional string matching with semantic evaluation to validate non-deterministic code outputs.
The Eye as a Pointer
You are looking at a bug. You press the spacebar, say "fix this," and release the key. The AI knows exactly which "this" you meant. No clicking, no highlighting, no scrolling.
The keyboard-less future of coding has long been a science fiction trope. GazeCoder turns it into a technical reality by treating the human eye as a high-bandwidth cursor. Built as a Next.js application, it combines WebGazer.js for eye tracking, Whisper for voice transcription, and GPT-4 for code generation.
Solving the Saccade Struggle
Browser-based eye tracking is notoriously difficult to calibrate. Eyes are inherently jittery. They dart around in rapid movements called saccades. If an IDE simply passed your live gaze coordinates to an LLM while you spoke a command, the context window would be a chaotic mess of every line you glanced at.
GazeCoder solves this with a clever state management trick inside its React component tree. The interface relies on a custom hook that freezes the gaze focus the moment recording starts.
When the user presses the spacebar, a hook locks the current gaze state. You can look away or read other functions while speaking your command. The system has already captured your initial intent.
The Multimodal Orchestrator
Most AI coding assistants operate on a selection-centric loop. You highlight text, right-click, and type a prompt. GazeCoder operates on a context-augmented generation loop.
The orchestration happens in a dedicated backend route that proxies requests to OpenAI. It bundles the visual context (the gazed-at lines), the audio context (the transcribed voice command), and environmental context (clipboard contents) into a single structured payload. It leverages OpenAI function calling to force the model to return a structured JSON response containing the exact code diff.
Probabilistic Quality Control
Traditional unit tests check for exact string matches or deterministic outputs. When your input is a voice command and your output is generated by an LLM, exact string matching becomes a brittle liability.
To safely develop this non-deterministic system, the developer built an AI-in-the-loop testing framework. GazeCoder extends the standard Jest testing library with a custom matcher.
"Please respond with <reasoning></reasoning> tags... and then finally a <grade></grade> tag"
This matcher sends the test output and a reference answer to an OpenAI model. It forces the model to output a step-by-step reasoning chain before issuing a final grade. It is a probabilistic safety net that ensures the semantic intent of the code remains intact even if the exact syntax varies.
| Dimension | Traditional AI IDEs | GazeCoder |
|---|---|---|
| Input Method | Keyboard and Mouse | Eye Tracking and Voice |
| Context Selection | Manual Text Highlighting | Automatic Gaze Intersection |
| Cognitive Load | High (Context Switching) | Low (Continuous Flow) |
| Testing Paradigm | Exact String Matching | Semantic Evaluation (toBeAsGoodAs) |