SmallCode: The Coding Agent That Treats Small LLMs Like a System to Be Engineered

A deep dive into the local-first agent that narrows tools, shrinks context, compiles prompt logic, and escalates only when the model actually needs help.

8 min read • View on GitHub • More from Doorman11991

A workshop-like control system with a compact machine at the center, narrow pipes feeding in from a few labeled tool drawers, and a tight context panel that limits what the machine can see. It explains the project’s core idea: small models become useful when the software around them absorbs complexity.
SmallCode is not trying to make a small model smarter. It is trying to make the surrounding system stricter, narrower, and more forgiving.
Key Takeaways

SmallCode’s wager is simple: a small local model should not be treated like a weaker version of Claude or GPT. It should be treated like a capable but brittle coworker. That changes everything about the agent around it.

The Bet: Small Models Need a Different Agent

The repo is built for the 8B to 35B class of models, the ones developers actually run on consumer hardware. That matters because the failure modes are different. Tool calls break format. Context drifts. Multi-step plans unravel.

SmallCode responds by reducing the model’s freedom before it starts doing work. It compresses context, narrows tool choice, and forces the system to recover from errors instead of pretending they will not happen.

SmallCode is a terminal-native coding agent designed from the ground up to extract useful work from local models (8B-35B) running on consumer hardware.

Doorman11991, Project Maintainer · Doorman11991/smallcode - GitHub

MarrowScript Turns Prompts Into Compiled Behavior

The most distinctive move in the repo is MarrowScript, a custom DSL for agent behavior. Prompts are not just strings. They are declared with constraints, retries, and caching, then compiled into executable code under src/compiled/.

A close-up of a mechanical press stamping a prompt specification into a row of executable blocks. One side shows typed declarations with retries and cache constraints. The other side shows the resulting workflow feeding an agent loop. It explains how the repo turns prompt logic into compiled software.
MarrowScript is the project’s clearest philosophical statement: agent behavior should be declared, compiled, and inspected like code.

That is why the repo feels different from a wrapper around an API. The intelligence layer is not hidden in prompt spaghetti. It is formalized, compiled, and shipped as part of the system.

I was frustrated that every coding agent (OpenCode, Cursor, Claude Code) assumes you're running GPT-5.4 or Claude Opus. If you try them with a local model like Gemma or Qwen they fall apart.

Glittering_Focus1538, Project Creator (Implied) · r/ollama benchmark post

Why the Agent Cuts Choice Down Before It Cuts Code

SmallCode’s tool strategy is built around the same premise. The system does not throw a huge tool surface at the model and hope for the best. It stages decisions.

This diagram shows the real design pattern. SmallCode does not maximize the model’s options. It meters them out in stages.

The first move is routing. The model classifies the task, then receives only the relevant tools. The second move is shaping output. Token-aware rewriting trims noisy command output before it hits the model. The third move is editing discipline. Patch-based writes are safer than full-file rewrites because exact string replacement is easier to verify.

PatternWhat it assumesWhat usually breaksSmallCode’s answer
Frontier-model-first agentThe model can hold lots of context and call tools cleanlyLocal models drift, overreach, or format tool calls badlyStage decisions and expose fewer tools at once
Full-file rewrite workflowThe model can regenerate a file accurately from scratchSmall models miss details and corrupt unrelated codeUse exact patches and read-before-write guards
Dump-context-and-prayMore context means better resultsNoise swamps the model and wastes tokensSummarize, rewrite, and budget aggressively
Single-shot tool choiceOne pass should be enoughThe wrong tool choice compounds failureUse a two-stage router and narrow categories first

The effect is not cosmetic. It is the difference between an agent that looks powerful in a demo and one that survives on weaker hardware in real sessions.

The Governor Watches for Failure, Then Changes Strategy

The governor is the control layer that makes the whole system feel disciplined instead of naive. It tracks whether tools are succeeding, notices repeated failures, and changes strategy when the local path stops working.

That can mean trying a different approach. It can also mean escalation. In other words, SmallCode does not insist that local execution solve every problem. It knows when to hand off.

That matters because failure handling is where many agents quietly collapse. They keep calling the same tool in the same way and call it autonomy. SmallCode treats repeated failure as a signal, not a personality trait.

What This Rewrites About Coding Agents

The cleanest comparison is not “SmallCode versus another feature list.” It is a clash of assumptions. Frontier-first agents assume abundant capability. SmallCode assumes scarcity and builds around it.

DimensionFrontier-model-first agentsSmallCode
Primary assumptionThe model is the center of intelligenceThe system must absorb model weakness
Tool surfaceMany tools, broad freedomFewer tools, staged access
Context strategyMore context, more dataBudgeted context, aggressive trimming
Editing styleOften full-file regenerationPatch-first, exact replacement
Failure policyRetry and hopeScore, adapt, escalate
Target userCloud-first teamsLocal-first, privacy-conscious, budget-limited developers

That is why the repo matters. It is not claiming that small models beat frontier models at everything. It is showing that a different market exists: people who want local control, predictable cost, and an agent shaped for imperfect hardware.

Where the Project Is Going

If prompt logic can be compiled, tools can be tiered, and failures can be routed, the agent starts to look less like a chat interface and more like an adaptive operating layer. That is the bigger idea hiding inside SmallCode.

The project’s contribution is architectural. It reframes coding agents as systems of constraints, not just bigger prompt windows. For small models, that is not a compromise. It is the product.