`agents-best-practices`: The Repo That Teaches Agents How to Build Their Own Harness

A provider-neutral skill for Claude Code and Codex that turns agent design into a runtime problem, not a prompt-writing exercise.

8 min read • View on GitHub • More from DenisSergeevitch

A wide control room with a central harness console routing model proposals through gates, validators, and a state ledger before anything reaches the outside world. The image explains that the runtime around the model owns authority, not the model itself.
The thesis in one scene: the model proposes, but the harness decides what survives the trip to execution.
Key Takeaways

The Prompt Is Not the System

Most agent tutorials start with prompts. This repo starts one layer deeper: the harness, the runtime shell that sits between the model and the world. That is the argument in one line. If an agent is unreliable, the missing piece is often not better wording, but better control over permissions, state, validation, and execution.

That is why agents-best-practices is interesting. It is not a framework you run. It is a skill that teaches other systems how to build the layer around the model. The repo treats the model like a proposal engine and the harness like the authority that decides what becomes real.

A Skill That Teaches Skills

The repository is designed for agentic IDEs and CLI tools, with SKILL.md acting as the activation layer and references/ holding the architectural playbook. In other words, this is documentation that is meant to be loaded into an agent, not just read by a human.

# SKILL.md

## Activation Triggers
- build an agent
- design tools
- audit harness

## Default behavior
When activated, stop giving vague advice.
Generate a concrete harness blueprint.

That structure matters because it changes the unit of value. Instead of a static best-practices document, the repo behaves like a portable operating manual. It is also provider-neutral, which makes the ideas easier to move between Claude Code, Codex, and other agentic environments.

The Manifest Decides When the Skill Wakes Up

The repo’s activation layer is the clever part. SKILL.md does not just describe the project. It tells the agent when to switch modes, and what to do once it does. The idea of “MVP Builder Mode” is a useful tell. It pushes the model away from generic guidance and toward a concrete harness blueprint.

The control loop is the product. The model proposes, but every other stage exists to make sure proposals are validated, authorized, executed once, and remembered safely.

Inside the Harness: The Model Proposes, the Harness Disposes

The architecture file centers on a 15-component control plane. That list is not decoration. It is a map of where authority lives: instruction management, permissions, execution control, memory, observation, and compaction all sit between the LLM and the outside world.

The best way to read that design is as a separation-of-duties system. The model can suggest. The harness can validate, approve, execute, persist, and compact. That split is what turns a chatty assistant into something closer to an operating system for work.

Traditional wrapperHarness-first system
Model touches tools directlyModel proposes, harness mediates
Permissions handled informallyPermissions enforced outside the model
Memory mostly lives in chat historyDurable state lives outside the prompt
Tool calls can be ambiguousTool results are exactly once and recorded
Safety depends on prompting careSafety depends on runtime contracts
Hard to move across providersProvider-neutral by design

Why the Loop Matters More Than the Model

The runtime logic is where the repo becomes more than a philosophy deck. references/agentic-loop.md defines an invariant: validate before permission, permission before execution, and one tool call equals one result. That sounds simple, but it is the difference between a loosely supervised chatbot and a predictable agent loop.

while true:
  request = read_input()
  plan = model.propose(request)
  validate(plan)
  check_permissions(plan)
  result = execute(plan)
  assert exactly_one_result(result)
  durable_state.update(result)
  context = compact(durable_state)
  respond(context)

The repo also distinguishes manual loops from hosted loops. That distinction matters because the control contract has to survive platform changes. If the provider runs part of the loop, the harness still has to own the rules.

ConcernManual loopHosted loop
Who drives the cycleClient or local runtimeProvider service
Where policy livesIn the harnessStill in the harness
How results are recordedExplicitly in local stateExplicitly through the contract
PortabilityHighHigh if the contract stays intact
Main riskDrift in implementationDrift in provider behavior

Security Is a Product Decision, Not a Prompt Trick

The security layer makes the repo practical. Instead of asking the model to be careful, it separates tools by risk and uses draft-versus-commit patterns for sensitive actions. That is a much stronger move than a polite warning in a prompt.

A close-up split workbench shows a risky action on one side with a direct send lever, and on the other side the same task broken into draft and commit stages with a human hand holding the commit key. The image explains why high-risk agent actions need separate approval boundaries.
Risky actions become safer when drafting and committing are different tools, not different moods.

This is the repo’s most enterprise-friendly idea. It replaces trust in the model’s judgment with trust in the system’s structure. That matters anywhere agents can move money, send messages, change infrastructure, or commit to external actions.

High-risk patternSafer harness pattern
One tool that drafts and sendsSeparate draft and commit tools
Model decides inside the same stepHuman checkpoint before commit
Risk hidden inside a promptRisk encoded in the API surface
Ambiguity about side effectsClear boundary before execution

The Quiet Revolution Is Documentation That Executes

The broader shift here is subtle but important. This repo treats documentation as runtime policy. It is written for machines first, humans second, which means the real audience is the agent that has to act inside these rules.

That is the category change. The best open-source repos are no longer just libraries or apps. Some are instruction sets for systems that will read, interpret, and act. agents-best-practices is one of those. It does not just explain the harness. It is a harness for thinking about harnesses.

Old documentation modelNew documentation model
Reference for humansPolicy for machines
Explains what to doConstrains what agents may do
Read once, consult laterLoaded into runtime behavior
Static guidanceExecutable operating assumptions