deepseek-harness: DeepSeek Harness: The Agent Framework That Makes Itself Keep the Receipts

dsh turns documentation, translation, and architectural decisions into enforced runtime discipline, then wraps it in a plugin system built for modular agent infrastructure.

8 to 10 min read • View on GitHub • More from deepseek-ai

A wide bureaucratic machine where an agent sits at a drafting table, stamping cards into compartments labeled proposed, implemented, rejected, and archived. Thin wires run from the filing system into a central control panel marked Context, showing that the record-keeping itself is part of the runtime.
DeepSeek Harness treats decision records as infrastructure, not as after-the-fact paperwork. The agent is expected to move its own ideas through a governed lifecycle.

Hi I'm one of the authors of DeepSeek Harness. It's just an early developer preview version we're presenting in MIT license currently. Expect lots of rough edges and compatibility-breaking changes.

Tianyi Cui, Author / Maintainer · DeepSeek Harness developer preview | Hacker News
Key Takeaways

Most agent frameworks optimize for the output. DeepSeek Harness optimizes for control. Its strangest move is the Agent Note system, where the agent is expected to write down, translate, version, and move its own decisions through a strict lifecycle.

That changes the job of the repository. It is not just a runtime for tools and prompts. It is a system for keeping the agent accountable to its own history.

The strangest thing about dsh: the agent has to document itself

The note system is structured like a decision ledger. Notes move through proposed, implemented, rejected, and archived, and the archived side keeps a frozen manifest so the rationale does not drift after the fact.

That is not normal repo hygiene. It is process enforced at the file level. The project treats explanation as something the agent must produce before it gets to change the system.

A hedcut-style portrait of Tianyi Cui, rendered as black ink stippling and hatching on a pure white background. The image presents him as the maintainer associated with the repository’s governance-heavy design.
.agents/notes/
  proposed/
  implemented/
  rejected/
  archived/

.agents/notes/README.md
.agents/notes/README.i18n.yaml
.agents/notes/archived/manifest.json

The sidecar translation files matter for the same reason. English and Chinese are tied together with hash checks, so the repo does not just ask for bilingual documentation. It verifies that the two versions stay in sync.

Everything is a plugin is not a slogan here

The core architecture runs on Cordis, a service-based context model. Plugins provide and consume capabilities through shared context instead of hard-coding dependencies into one central loop.

Cordis is the key to the repo’s modularity. The shared context lets runtime pieces cooperate without turning the harness into a monolith.

That makes the system feel less like an app and more like a chassis. Sessions, extensions, command adapters, and interfaces all sit on the same conceptual footing because the context orchestrates the relationships.

A modular machine built from detachable black metal blocks, with a shared core in the center and surrounding components that can be plugged in or removed. The scene shows the same assembly mid-swap, making the modular runtime model easy to grasp at a glance.
The architecture is swappable by design. The point is not a single fixed chain, but a context-bound mesh of replaceable parts.

The harness has a runtime, not just a prompt loop

The repo is not only about prompts and tool calls. It includes a Web UI, a terminal UI, ACP support, session replay, and trajectory-style workflows for inspecting how an agent arrived at an outcome.

That matters because it changes what the product is. A chat wrapper shows text. A runtime shows state, history, and control surfaces.

Web UI 3080
  -> session log
  -> trajectory view
  -> tool execution
  -> replay / fork

TUI
  -> interactive session control
  -> reloadable commands
  -> terminal-first agent work

Why the comparison is the point

The easiest way to misread dsh is to compare it to polished agent products and expect the same goal. It is not trying to be the friendliest surface. It is trying to be the most composable and auditable runtime.

ProjectPrimary goalWhat is swappableModel dependenceRuntime styleGovernance / traceabilityBest fit
DeepSeek HarnessComposable, auditable agent runtimeModels, tools, sessions, UI, adapters, storageLow to mediumContext-bound plugin meshStrongTeams that want control and provenance
Claude CodeRefined coding assistantLimitedHighProductized CLIModerateUsers who want an opinionated tool
OpenClawBroad open-source agent surfaceVaries by integrationMediumMulti-channel applicationMixedGeneral-purpose deployment
Hermes AgentMemory-centric agent designMemory and workflow componentsMediumPipeline-orientedModerateExperiments centered on long-term memory
OpenCode / Codex-style harnessesDeveloper coding workflowSome integrationsMedium to highCLI-first harnessBasic to moderateStraightforward coding assistance

The contrast is philosophical as much as technical. Claude Code is a finished product. dsh looks more like infrastructure you can rewire.

Bilingual governance as a scaling strategy

The translation system is not cosmetic. The sidecar .i18n.yaml files and hash checks turn bilingual documentation into a verifiable part of the repo, while the archived manifest freezes decisions after they have been settled.

That is the broader pattern in the project. DeepSeek is encoding consistency into the repository itself, so the harness can grow without losing its own rules.

Everything is a plugin, literally everything. The core idea driving the whole project: every component of the agent — model adapter, tool registry, session log, agent loop, storage, scheduling, sandboxing, even the UI itself — is a swappable plugin conforming to the same interface.

DeepSeek AI, Project Maintainer · DeepSeek Harness: Everything is a Plugin

What dsh is really optimized for

The project is optimized for sovereignty, traceability, and change over time. It assumes the agent is not a black box to be trusted blindly, but a participant whose decisions must be explainable and replayable.

That is why the note system matters as much as the runtime. dsh does not just orchestrate agents. It orchestrates the rules by which agents explain themselves.