Temper: The AI Coding Framework That Treats Context Like a Hazard

A deep dive into the wrapper that replaces open-ended prompting with staged intent, source checks, and feedback loops that make AI code feel less like improvisation and more like engineering.

11 min read • View on GitHub • More from galando

A wide assembly line inside a stark machine room. Loose prompts and paper scraps enter on the left, then pass through four chambers before exiting as a stamped artifact. The image explains Temper's core idea: AI coding becomes a controlled pipeline instead of a freeform chat.
Temper turns loose context into staged work. The point is not more prompting, but tighter control over what survives into the next phase.
Key Takeaways

Temper is not interesting because it adds another chat surface to AI coding. It is interesting because it distrusts the thing most AI tools sell as a strength: open-ended context. The repo’s premise is blunt. If code is going to pass through an AI, the work needs a shape, a boundary, and a way to prove it was judged against something real.

That makes Temper feel less like a coding assistant and more like a process wrapper. It asks a harder question than “what can the model write?” It asks, “how do we make the model’s output survive review without turning review into theater?”

The uncomfortable thesis: more context can make AI coding worse

Temper’s answer is counterintuitive. Give the model less, but better. Not a sprawling thread of accumulated prompts. Not a pile of stale decisions. Instead, break the job into bounded stages and let each stage carry only the context it needs. In the repo’s own framing, that is how you reduce drift, hallucinated confidence, and the tendency to overbuild.

But "fast" without "right" creates bugs, technical debt, and features that miss the point.

Project README, Repository documentation · galando/temper README

From chat to pipeline

The central orchestrator lives in `.claude/commands/temper.md`, and its job is structural. It does not ask one agent to remember everything. It splits work into Plan, Design, Build, and Review. Each stage is its own checkpoint, with its own prompt shape and its own output artifact.

Temper treats each stage as a separate context window. That is the mechanism that keeps one phase from contaminating the next.

That separation matters because AI systems are good at sliding across boundaries. A loose planning note becomes a hidden requirement. A design tradeoff becomes a speculative implementation detail. Temper forces those things apart before they can blur together.

/temper:plan  -> write intent.md
/temper:design -> narrow the implementation path
/temper:build  -> implement against the staged intent
/temper:review -> run the validation contract

if review fails:
  check-to-build
  create fix task
  re-enter build stage

`intent.md` is the project’s most novel artifact

This is the part that changes the feel of the whole system. `intent.md` is not just a note. It is a contract that says how success will be judged. Temper’s validate types, including `scenario`, `code`, `metric`, and `manual`, turn review into something closer to a predeclared test plan than a human opinion.

A close-up desk scene centered on a single intent.md file under a mechanical validation stamp. Around it are four small tokens for scenario, code, metric, and manual. The image explains that Temper predeclares how work will be judged before the implementation begins.
`intent.md` is the key move. It binds the build to the review criteria before code is written.
DimensionChat-first assistantsTemper
Context modelEverything stays in one conversation.Each stage gets a narrow, staged context.
ValidationThe model or reviewer makes a vibe-based judgment.Review runs against the validation type declared in `intent.md`.
Failure handlingA bad answer often means a retry with more prompting.A failed check can loop back into build through `check-to-build`.
MemoryPreferences are usually implicit and ephemeral.Repeated dismissals can be stored in review memory.
API freshnessStale examples can slip through easily.Source-driven development fetches current docs before code is written.
Human roleThe human steers the chat.The human sets constraints, validates outcomes, and shapes policy.

The contrast is subtle but important. Chat-first tools are optimized for momentum. Temper is optimized for legibility. It would rather slow the machine down than let a blurry requirement leak forward as if it were settled.

Source-driven development as an anti-hallucination system

Temper’s source-driven workflow is a guardrail against the most familiar AI coding mistake: confidently using the wrong API. The repo’s pattern is simple. Detect the actual framework version, fetch the current docs, then write code against the source of truth instead of memory.

1. Read package.json, pom.xml, or equivalent
2. Detect the exact framework version
3. Fetch current documentation through MCP
4. Write implementation with fresh source references
5. Avoid stale patterns and deprecated APIs

That matters because the failure mode is not just broken code. It is code that looks plausible enough to pass a casual review. Temper tries to prevent the illusion by forcing the model to re-anchor itself in current documentation before it starts composing output.

Context engineering as token triage

The repo’s context-engineering rules make the discipline explicit. Keep the working set under roughly 2,000 lines. Prioritize rules over architecture, architecture over source, source over errors. Stop at two hops when traversing dependencies. This is not generosity. It is triage.

The policy is doing something subtle. It assumes that every extra token has a cost, but not every token has the same value. So the system reserves room for the things most likely to change the outcome, and cuts off the rest before it becomes noise.

Priority order:
1. Rules
2. Architecture
3. Source
4. Errors

Traversal rule:
- Stop at 2 hops
- Do not expand the graph indefinitely
- Keep the context window readable and local

The feedback loop that teaches the tool your preferences

Temper’s `.temper/review-memory.json` is where the product gets personal. If a developer keeps dismissing the same suggestion, the system can remember that preference and suppress the repeat. Then the check-to-build loop turns failed review into a concrete fix path instead of a dead end.

That changes the feel of the tool. It is no longer just a static policy engine. It starts to behave like a local operating agreement, one that learns what your team treats as signal and what it treats as churn.

BehaviorWhat happensWhy it matters
Repeated suggestionsDismissed rules can be suppressed after enough repetition.The system stops wasting attention on known false positives.
Check failureA failed validation can route work back into build.Review becomes an actionable loop, not a terminal verdict.
Local memoryPreference data persists across sessions.The assistant adapts to the team, not just the repository.

Where Temper fits in the AI coding stack

Temper is not trying to replace a coding assistant, a linter, or an evaluation suite. It sits above them. Chat tools help generate options. Linters and tests provide deterministic signals. Temper’s role is to orchestrate the behavior between those layers so the AI does not get to improvise its own process.

LayerJobTemper’s relationship to it
Chat-first coding assistantProduces code and suggestions quickly.Temper constrains how that output is staged and reviewed.
Linting and static analysisCatch deterministic defects.Temper can route validation through those signals, but does not replace them.
Evaluation frameworksMeasure code quality or task success.Temper borrows the idea of validation, then makes it part of the build contract.

That positioning is the real bet. The repo is betting that teams will eventually care less about a model’s raw fluency and more about whether its work can be inspected, replayed, and corrected without drama.

What Temper is really betting on

Temper’s philosophy is simple enough to fit on one line. AI coding gets better when you stop treating context as a conversational convenience and start treating it as a controlled asset. The project is a reminder that the hard part of AI software is often not generation. It is governance.

If it works, Temper will not be remembered as a wrapper. It will be remembered as a design pattern for making AI work feel native to engineering teams. That is a much bigger ambition than a better prompt template.