deepseek-harness: DeepSeek Harness: The Agent Framework That Makes Itself Keep the Receipts
dsh turns documentation, translation, and architectural decisions into enforced runtime discipline, then wraps it in a plugin system built for modular agent infrastructure.

Hi I'm one of the authors of DeepSeek Harness. It's just an early developer preview version we're presenting in MIT license currently. Expect lots of rough edges and compatibility-breaking changes.
- DeepSeek Harness is unusual because it makes the agent explain and version its own decisions, turning documentation into a governance layer.
- Its Cordis-based plugin model treats the runtime as a shared context, so major pieces can be swapped without collapsing the whole system.
- The project is less a polished chatbot than an execution environment with Web UI, TUI, session replay, and ACP plumbing.
- Its bilingual note system and archived manifests make traceability part of the repository’s design, not a separate process.
Most agent frameworks optimize for the output. DeepSeek Harness optimizes for control. Its strangest move is the Agent Note system, where the agent is expected to write down, translate, version, and move its own decisions through a strict lifecycle.
That changes the job of the repository. It is not just a runtime for tools and prompts. It is a system for keeping the agent accountable to its own history.
The strangest thing about dsh: the agent has to document itself
The note system is structured like a decision ledger. Notes move through proposed, implemented, rejected, and archived, and the archived side keeps a frozen manifest so the rationale does not drift after the fact.
That is not normal repo hygiene. It is process enforced at the file level. The project treats explanation as something the agent must produce before it gets to change the system.
.agents/notes/
proposed/
implemented/
rejected/
archived/
.agents/notes/README.md
.agents/notes/README.i18n.yaml
.agents/notes/archived/manifest.json
The sidecar translation files matter for the same reason. English and Chinese are tied together with hash checks, so the repo does not just ask for bilingual documentation. It verifies that the two versions stay in sync.
Everything is a plugin is not a slogan here
The core architecture runs on Cordis, a service-based context model. Plugins provide and consume capabilities through shared context instead of hard-coding dependencies into one central loop.
That makes the system feel less like an app and more like a chassis. Sessions, extensions, command adapters, and interfaces all sit on the same conceptual footing because the context orchestrates the relationships.
The harness has a runtime, not just a prompt loop
The repo is not only about prompts and tool calls. It includes a Web UI, a terminal UI, ACP support, session replay, and trajectory-style workflows for inspecting how an agent arrived at an outcome.
That matters because it changes what the product is. A chat wrapper shows text. A runtime shows state, history, and control surfaces.
Web UI 3080
-> session log
-> trajectory view
-> tool execution
-> replay / fork
TUI
-> interactive session control
-> reloadable commands
-> terminal-first agent work
Why the comparison is the point
The easiest way to misread dsh is to compare it to polished agent products and expect the same goal. It is not trying to be the friendliest surface. It is trying to be the most composable and auditable runtime.
| Project | Primary goal | What is swappable | Model dependence | Runtime style | Governance / traceability | Best fit |
|---|---|---|---|---|---|---|
| DeepSeek Harness | Composable, auditable agent runtime | Models, tools, sessions, UI, adapters, storage | Low to medium | Context-bound plugin mesh | Strong | Teams that want control and provenance |
| Claude Code | Refined coding assistant | Limited | High | Productized CLI | Moderate | Users who want an opinionated tool |
| OpenClaw | Broad open-source agent surface | Varies by integration | Medium | Multi-channel application | Mixed | General-purpose deployment |
| Hermes Agent | Memory-centric agent design | Memory and workflow components | Medium | Pipeline-oriented | Moderate | Experiments centered on long-term memory |
| OpenCode / Codex-style harnesses | Developer coding workflow | Some integrations | Medium to high | CLI-first harness | Basic to moderate | Straightforward coding assistance |
The contrast is philosophical as much as technical. Claude Code is a finished product. dsh looks more like infrastructure you can rewire.
Bilingual governance as a scaling strategy
The translation system is not cosmetic. The sidecar .i18n.yaml files and hash checks turn bilingual documentation into a verifiable part of the repo, while the archived manifest freezes decisions after they have been settled.
That is the broader pattern in the project. DeepSeek is encoding consistency into the repository itself, so the harness can grow without losing its own rules.
Everything is a plugin, literally everything. The core idea driving the whole project: every component of the agent — model adapter, tool registry, session log, agent loop, storage, scheduling, sandboxing, even the UI itself — is a swappable plugin conforming to the same interface.
What dsh is really optimized for
The project is optimized for sovereignty, traceability, and change over time. It assumes the agent is not a black box to be trusted blindly, but a participant whose decisions must be explainable and replayable.
That is why the note system matters as much as the runtime. dsh does not just orchestrate agents. It orchestrates the rules by which agents explain themselves.