Temper: The AI Coding Framework That Treats Context Like a Hazard
A deep dive into the wrapper that replaces open-ended prompting with staged intent, source checks, and feedback loops that make AI code feel less like improvisation and more like engineering.
- Temper’s real argument is that AI coding fails when context stays conversational instead of becoming an audited work product.
- Its most important invention is `intent.md`, which turns review into a predeclared validation contract rather than a subjective vibe check.
- Source-driven development and context engineering make the system less ambitious and more disciplined, which is exactly the point.
- Temper is trying to sit above coding assistants and below the human team, shaping process instead of generating code directly.
Temper is not interesting because it adds another chat surface to AI coding. It is interesting because it distrusts the thing most AI tools sell as a strength: open-ended context. The repo’s premise is blunt. If code is going to pass through an AI, the work needs a shape, a boundary, and a way to prove it was judged against something real.
That makes Temper feel less like a coding assistant and more like a process wrapper. It asks a harder question than “what can the model write?” It asks, “how do we make the model’s output survive review without turning review into theater?”
The uncomfortable thesis: more context can make AI coding worse
Temper’s answer is counterintuitive. Give the model less, but better. Not a sprawling thread of accumulated prompts. Not a pile of stale decisions. Instead, break the job into bounded stages and let each stage carry only the context it needs. In the repo’s own framing, that is how you reduce drift, hallucinated confidence, and the tendency to overbuild.
But "fast" without "right" creates bugs, technical debt, and features that miss the point.
From chat to pipeline
The central orchestrator lives in `.claude/commands/temper.md`, and its job is structural. It does not ask one agent to remember everything. It splits work into Plan, Design, Build, and Review. Each stage is its own checkpoint, with its own prompt shape and its own output artifact.
That separation matters because AI systems are good at sliding across boundaries. A loose planning note becomes a hidden requirement. A design tradeoff becomes a speculative implementation detail. Temper forces those things apart before they can blur together.
/temper:plan -> write intent.md
/temper:design -> narrow the implementation path
/temper:build -> implement against the staged intent
/temper:review -> run the validation contract
if review fails:
check-to-build
create fix task
re-enter build stage
`intent.md` is the project’s most novel artifact
This is the part that changes the feel of the whole system. `intent.md` is not just a note. It is a contract that says how success will be judged. Temper’s validate types, including `scenario`, `code`, `metric`, and `manual`, turn review into something closer to a predeclared test plan than a human opinion.
| Dimension | Chat-first assistants | Temper |
|---|---|---|
| Context model | Everything stays in one conversation. | Each stage gets a narrow, staged context. |
| Validation | The model or reviewer makes a vibe-based judgment. | Review runs against the validation type declared in `intent.md`. |
| Failure handling | A bad answer often means a retry with more prompting. | A failed check can loop back into build through `check-to-build`. |
| Memory | Preferences are usually implicit and ephemeral. | Repeated dismissals can be stored in review memory. |
| API freshness | Stale examples can slip through easily. | Source-driven development fetches current docs before code is written. |
| Human role | The human steers the chat. | The human sets constraints, validates outcomes, and shapes policy. |
The contrast is subtle but important. Chat-first tools are optimized for momentum. Temper is optimized for legibility. It would rather slow the machine down than let a blurry requirement leak forward as if it were settled.
Source-driven development as an anti-hallucination system
Temper’s source-driven workflow is a guardrail against the most familiar AI coding mistake: confidently using the wrong API. The repo’s pattern is simple. Detect the actual framework version, fetch the current docs, then write code against the source of truth instead of memory.
1. Read package.json, pom.xml, or equivalent
2. Detect the exact framework version
3. Fetch current documentation through MCP
4. Write implementation with fresh source references
5. Avoid stale patterns and deprecated APIs
That matters because the failure mode is not just broken code. It is code that looks plausible enough to pass a casual review. Temper tries to prevent the illusion by forcing the model to re-anchor itself in current documentation before it starts composing output.
Context engineering as token triage
The repo’s context-engineering rules make the discipline explicit. Keep the working set under roughly 2,000 lines. Prioritize rules over architecture, architecture over source, source over errors. Stop at two hops when traversing dependencies. This is not generosity. It is triage.
The policy is doing something subtle. It assumes that every extra token has a cost, but not every token has the same value. So the system reserves room for the things most likely to change the outcome, and cuts off the rest before it becomes noise.
Priority order:
1. Rules
2. Architecture
3. Source
4. Errors
Traversal rule:
- Stop at 2 hops
- Do not expand the graph indefinitely
- Keep the context window readable and local
The feedback loop that teaches the tool your preferences
Temper’s `.temper/review-memory.json` is where the product gets personal. If a developer keeps dismissing the same suggestion, the system can remember that preference and suppress the repeat. Then the check-to-build loop turns failed review into a concrete fix path instead of a dead end.
That changes the feel of the tool. It is no longer just a static policy engine. It starts to behave like a local operating agreement, one that learns what your team treats as signal and what it treats as churn.
| Behavior | What happens | Why it matters |
|---|---|---|
| Repeated suggestions | Dismissed rules can be suppressed after enough repetition. | The system stops wasting attention on known false positives. |
| Check failure | A failed validation can route work back into build. | Review becomes an actionable loop, not a terminal verdict. |
| Local memory | Preference data persists across sessions. | The assistant adapts to the team, not just the repository. |
Where Temper fits in the AI coding stack
Temper is not trying to replace a coding assistant, a linter, or an evaluation suite. It sits above them. Chat tools help generate options. Linters and tests provide deterministic signals. Temper’s role is to orchestrate the behavior between those layers so the AI does not get to improvise its own process.
| Layer | Job | Temper’s relationship to it |
|---|---|---|
| Chat-first coding assistant | Produces code and suggestions quickly. | Temper constrains how that output is staged and reviewed. |
| Linting and static analysis | Catch deterministic defects. | Temper can route validation through those signals, but does not replace them. |
| Evaluation frameworks | Measure code quality or task success. | Temper borrows the idea of validation, then makes it part of the build contract. |
That positioning is the real bet. The repo is betting that teams will eventually care less about a model’s raw fluency and more about whether its work can be inspected, replayed, and corrected without drama.
What Temper is really betting on
Temper’s philosophy is simple enough to fit on one line. AI coding gets better when you stop treating context as a conversational convenience and start treating it as a controlled asset. The project is a reminder that the hard part of AI software is often not generation. It is governance.
If it works, Temper will not be remembered as a wrapper. It will be remembered as a design pattern for making AI work feel native to engineering teams. That is a much bigger ambition than a better prompt template.