opencode-voice: When the Tool Description Becomes the Voice

A small OpenCode plugin uses ElevenLabs v3, expressive audio tags, and macOS playback to make an AI agent sound directed, not robotic.

8 min read • View on GitHub • More from anomalyco

A conductor's score, a terminal window, and a speaker cone are fused into one scene. The image shows that the plugin's voice comes from instructions in the tool description before speech is synthesized and played back.
The plugin treats expressive delivery as a written score, not a hidden mixer.

An OpenCode plugin that adds text-to-speech capabilities using ElevenLabs v3 with expressive audio tags.

R44VC0RP, Contributor · anomalyco/opencode-voice README
Key Takeaways

Most TTS plugins start at the audio layer. opencode-voice starts one layer earlier, inside the speak tool's description. That is the surprise: the model is taught how to perform, not just what to say.

The control surface is the description

The repo's core move is simple and unusual. It uses ElevenLabs v3, but the real product decision is to bake examples like [whispers], [laughs], and [sarcastic] into the tool metadata itself. In other words, the prompt is the interface.

That keeps the project small. The code lives mainly in src/plugin.ts, with index.ts exporting the plugin and setup.sh handling installation, secrets, and OpenCode config wiring. The setup script matters because it removes the annoying part: finding the key, setting permissions, and registering the plugin.

Portrait of R44VC0RP rendered as a WSJ hedcut from a verified GitHub avatar. The face is shown with crisp stippling and hatching so the contributor feels like a source, not an icon.

How a few lines become performance

A tight close-up of a cue card, a magnifying glass, and a thread that leads into a speaker cone. The image explains how the tool description acts like margin notes that guide the generated voice.
The tag system works because the model sees examples where it needs them, at tool-call time.

At runtime, the plugin asks ElevenLabs v3 for an MP3, writes the result to a temporary file, and launches playback with afplay on macOS. Playback is detached, so the agent can keep going while the audio speaks in the background. Cleanup still happens after exit, so the convenience does not leak temp files.

A tiny pipeline, but the interesting control point is upstream, where the tool description shapes the output.

Why the background playback matters

That detached process is not an implementation footnote. It defines the user experience. The agent can answer in text and voice at the same time, which makes the feature feel like a live companion instead of a blocking media player.

Non-blocking - Audio plays in background, control returns immediately

R44VC0RP, Contributor · anomalyco/opencode-voice README

That choice also clarifies where the project sits in the OpenCode ecosystem. It is an output-first plugin, not a full voice stack. The broader voice work around OpenCode points toward integrated modes that combine listening and speaking, while this repo proves the smaller claim first.

How it compares

Aspectopencode-voiceIntegrated voice mode in OpenCode
Primary jobSpeak with expressive outputHandle both listening and speaking
Control modelThe tool description teaches expressive tagsProduct-level voice controls manage the session
Runtime feelLightweight and non-blockingDeeper integration with continuous voice workflows
TradeoffFast to graft onto the TUIBroader surface area and more moving parts

That is why opencode-voice is interesting even if you never ship it. It proves a bigger thesis: some of the most important UX can be created by changing the instructions around a tool, not by rebuilding the tool itself.

The setup script is the product

One more detail makes the repo feel complete. setup.sh is not a throwaway installer. It checks dependencies, captures the secret, locks down file permissions with chmod 600, and uses jq to patch the OpenCode config automatically. That is a good sign in a small tool: the path from curiosity to working voice is short.

The stack is intentionally thin: TypeScript on Bun, an OpenCode plugin SDK, ElevenLabs over HTTP, and native afplay for playback. That restraint is the point. The repo spends its complexity budget on expression, not infrastructure.