Engine-Core: The Framework That Puts AI Engineers in a Sandbox

While others let LLMs hallucinate directly into production, Engine-Labs built a "shadow filesystem" to verify every line of code before it is committed.

8 min read • View on GitHub • More from Engine-Labs

A giant mechanical hand placing a delicate glass gear into a small, contained glass dome clockwork mechanism while a human figure observes.
The "Shadow Filesystem" acts as a quarantine zone. The AI tests its code in a parallel environment before touching the production repository.
Sudhir Balaji

Engine is a remote AI coding agent built for professional software engineers. Here we detail some technical challenges we faced building robust LLM terminal access.

— Sudhir Balaji, Author/Maintainer (Terminals for LLMs: a Halting Problem - by Sudhir)
Key Takeaways

The artificial intelligence software engineer bubble is saturated with thin wrappers. Most tools simply send a prompt to an LLM and blindly execute the response against a live codebase. They treat the model like an infallible senior developer. Engine-Core takes a different stance. It treats the LLM as a highly capable but fundamentally unreliable operator of complex machinery.

Built by Engine Labs, this TypeScript framework operates as a shadow-filesystem orchestrator. It does not just chat with code. It builds a temporary parallel universe for the AI to play in, testing every assumption before making a permanent change.

The Shadow Filesystem

The most compelling technical contribution of Engine-Core lives inside a file called editBackendFileFunctionAction.ts. This is where the framework implements its validation loop. When the AI suggests a code change, the engine does not write it to disk immediately.

Instead, it initiates a pre-flight check. The engine uses mkdtempSync to create a temporary directory and copies the entire project file tree into it. It applies the LLM's requested line-based edits to this shadow copy. Then, it runs a full TypeScript compilation (npx tsc). If the code compiles, the engine goes one step further. It spawns a temporary Fastify server to smoke-test the application. Only if the server boots successfully does the engine commit the changes to the real repository.

A horizontal process flow showing the "Validation Gauntlet". Start with a node labeled "AI Suggests Edit". Draw an arrow to a secure box labeled "Temp Directory Created (mkdtempSync)". Inside this box

This sandbox prevents the primary failure mode of AI-driven development. Hallucinated variables, syntax errors, and broken imports are caught in the quarantine zone. The AI acts as a junior developer, and the framework acts as an automated CI/CD pipeline for every single thought the model has.

Strategies vs. Adapters

Engine-Core achieves modularity through a strict separation of concerns. The architecture divides the system into Adapters and Strategies. Adapters normalize the communication layer. They handle the idiosyncrasies of OpenAI, Anthropic, or local Ollama models, translating generic messages into provider-specific tool schemas.

A robot engineer unplugging a glowing cube head and plugging in a glowing sphere head.
The Adapter pattern allows developers to swap the underlying LLM (the brain) without changing the tools or the workflow (the hands).

Strategies define the job description. A Strategy dictates the system prompt, the available toolset, and the specific execution environment. For example, the BackendStrategy knows how to run Prisma migrations and boot a Fastify server. The ShellStrategy knows how to manage raw bash processes. By decoupling the brain from the hands, developers can hot-swap Claude 3.5 for GPT-4o without rewriting a single tool definition.

The "Think-Then-Act" Loop

At the heart of the orchestration layer is src/chat.ts. This file manages the agentic loop. Unlike a simple request-response cycle, this function uses a recursive while loop. It continually processes tool calls and model responses until the AI decides the task is complete.

When the AI makes a mistake, the engine employs a technique called "Augmented Errors". If the TypeScript compiler throws an error in the shadow filesystem, the engine does not just return a raw stack trace. It slices the source code around the failing line and sends a highly contextualized error message back to the LLM. The AI sees exactly where it failed and is prompted to try again.

A typewriter with a red mechanical stamp highlighting a specific typo with a magnifying glass.
Augmented errors provide the LLM with surgical context, showing the exact source code lines surrounding a compilation failure.

Engineering the Terminal

Giving an LLM raw terminal access is notoriously difficult. The system must know when a command has finished executing, manage standard input streams, and handle interactive prompts. Engine-Core tackles the "halting problem" of LLM terminals by carefully managing process execution and standard output streams.

Simon Spurrier

Anthropic’s example creates a persistent Bash session for the LLM to interact with, whereas OpenAI’s example runs an LLM-generated Bash command directly as a subprocess.

— Simon Spurrier, Author/Maintainer (Terminals for LLMs: a Halting Problem - cto.new)

The framework uses Node's spawn to execute shell commands. It sources the user profile before running to ensure environment variables and paths are correctly loaded. This low-level control is what allows the engine to perform complex, multi-step operations reliably.

Engine Labs has posted a state of the art score on Terminal Bench- the top score with a Claude Sonnet 4 class model and #2 overall. This improves on Claude Code's equivalent score by ~25%, or around 10 percentage points and is a huge achievement for Engine's small but experienced technical team.

The "Goldilocks" Framework

Engine-Core sits in a specific niche within the LLM orchestration landscape. It is designed for developers who find enterprise frameworks too bloated but consider raw API calls too primitive.

Feature Engine-Core LangChain Microsoft Guidance
Abstraction Level Medium (Strategy & Tooling) High (Chains & Agents) Low (Token & Syntax)
Primary Unit of Code Chat Strategies Chains Handlebars Templates
Safety Guardrails Shadow Filesystem Compilation Custom Callbacks Constrained Generation
Terminal Support Native Subprocess Orchestration Third-party Tools N/A

Frameworks like LangChain attempt to be everything to everyone, often resulting in complex abstractions that obscure the underlying prompt logic. Guidance focuses heavily on constrained syntax generation. Engine-Core provides a structured way to define conversation logic and safely execute side effects without the overhead of a massive ecosystem.


Sources: