The C-Suite in Your Terminal: Unpacking gstack
How Garry Tan’s open-source framework turns Claude Code from a simple autocomplete tool into a multi-role software factory with persistent visual QA.
I don’t need modafinil with this revolution
- A persistent Chromium daemon reduces visual feedback latency from seconds to milliseconds for real-time UI verification.
- Markdown-based role templates force the AI to simulate a full management hierarchy to prevent low-quality code generation.
- Diff-based evaluations use structured LLM scoring to audit code changes without the cost of running full-suite tests on every commit.
- Security is maintained through strict localhost binding and UUID bearer tokens to prevent unauthorized process access.
The Browser That Never Sleeps
Most AI coding tools operate completely blind to the visual output of their code. They can write a React component, but they cannot tell you if the button is rendering off-center. To solve this, developers often bolt on web-scraping scripts, which introduce a massive performance penalty. Spinning up a fresh headless browser for every command takes seconds. That latency destroys the flow state required for rapid iteration.
gstack attacks this problem with a persistent browser daemon. Instead of launching a new instance per request, it runs a long-lived Chromium process managed by a Bun-based HTTP server. The AI agent communicates with this daemon via a local API. Because the browser is always running and holding state, interactions drop from three seconds to roughly a hundred milliseconds.
This speed changes the fundamental behavior of the AI workflow. The agent can take a semantic ARIA snapshot of the DOM, click a button, and verify the resulting state instantly. It is a continuous visual feedback loop. The architecture relies on a "crash-only" philosophy. If the headless browser enters a bad state, gstack simply exits the process and allows the CLI to auto-restart it, prioritizing recovery over complex error handling.
Hiring a Terminal C-Suite
Providing an AI with a browser is useless if it lacks the discipline to use it effectively. This is where gstack introduces its core orchestration layer. Rather than interacting with a generic assistant, developers invoke specific organizational roles using slash commands.
These roles are defined by strict Markdown templates. The /plan-ceo-review command forces the AI to challenge the product-market fit of the proposed changes before any code is written. It acts as a gatekeeper against unnecessary features. Once approved, the /plan-eng-review command dictates the architecture. Finally, the /review command acts as a paranoid staff engineer, aggressively auditing the generated code for performance regressions and security flaws.
This forced friction solves the common problem of AI slop. By making the AI argue with itself across different personas, gstack ensures that the final output meets a rigorous standard. The system tracks user preferences and will proactively suggest adjacent skills, simulating a high-functioning engineering team anticipating the needs of its technical lead.
The Engine and Diff-Based Evals
Under the hood, gstack is engineered for maximum throughput. It leverages Bun for native SQLite support, which is critical for decrypting Chrome cookies without requiring heavy external dependencies. Security is handled via strict localhost binding and UUID bearer tokens. Even if another process discovers the random port used by the browser daemon, it cannot issue commands without the token stored in an owner-only state file.
Because running complex LLM evaluations on every commit is prohibitively expensive, gstack employs a tiered testing strategy. The system uses a diff-based approach to determine which tests to run.
// Example of smart JS wrapping in gstack's read-commands.ts
export function wrapForEvaluate(code: string): string {
if (hasAwait(code) || isMultiLine(code)) {
return `(async () => {
${code}
})()`;
}
return code;
}
By comparing the current Git diff against the base branch, the system maps changed files to specific LLM-as-judge evaluations. This ensures that a minor documentation update does not trigger a costly, full-suite end-to-end test. The LLM judges use structured scoring rubrics, turning subjective code quality into hard metrics based on clarity, completeness, and actionability.
The Post-Junior Developer Era
The landscape of AI coding tools is diverging. Standard CLI agents and AI-integrated IDEs focus primarily on writing lines of code faster. gstack represents a different philosophy entirely: managing the organizational structures that produce code.
| Tool | Primary Interface | Visual QA | Persona Management |
|---|---|---|---|
| gstack | Meta-framework / CLI | Persistent Playwright Daemon | Strict multi-role templates |
| Aider | Direct CLI | None (Text/File only) | Single developer agent |
| Cursor | GUI / IDE | Manual user verification | Context-aware autocomplete |
The core argument of gstack is that the marginal cost of completeness has fallen to zero. When boilerplate, testing, and deployment scripts can be generated reliably in seconds, there is no longer a valid excuse for skipping them. By wrapping raw generative power in strict organizational protocols, gstack provides technical founders the leverage to operate at the scale of a full engineering department.
Sources: gstack GitHub Repository; "Garry Tan's Claude Code Setup," Bitcoin World; OpenClaw Repository.