dsh-anchored-standard: How One Tiny First Turn Steers DeepSeek Into a Better Mode
A DeepSeek Harness preset that starts with restraint, locks in a cleaner reasoning trajectory, then unlocks the full toolset only after the model is already on track.
- `dsh-anchored-standard` treats the first turn as a behavioral steering problem, not a tooling problem.
- The repo separates trajectory selection from later capability, so the model stabilizes before the full tool surface appears.
- Its context gate and tool bootstrap logic make the phase change durable across compaction and resume.
- The project is also a warning: sometimes fewer visible tools produce better agent behavior than a richer preset.
The surprising part of `dsh-anchored-standard` is not that it trims tools. It is that it uses a tiny first turn to shape how DeepSeek reasons for the rest of the session. The preset treats the model less like a static API consumer and more like a system whose behavior can be nudged by initial conditions.
That is a different kind of control. Standard agent design tries to expose more context, more tools, and more affordances. This repo argues that the first step is often the opposite: hide power until the model has stopped drifting.
The first turn decides more than you think
The project starts from a simple observation in the README: DeepSeek V4 Pro conditions strongly on the API-visible tool catalog. In practice, the model can lean toward a more efficient, engineering-minded trajectory when it sees a minimal environment first, then it can be promoted into the full resident toolset later.
DeepSeek V4 Pro conditions strongly on the API-visible tool catalog. In the Project2 evaluation, Standard and PTC produced scores of 91 and 92, while the official Minimal preset produced 99 and 96. Permanently staying on Minimal, however, gives up the Standard preset's broader tool set. Anchored Standard separates initial trajectory selection from later tool use.
The repo is built around a two-phase bootstrap. First, the session is anchored in a constrained environment. Then, once a promotion signal appears, the full standard catalog comes online and the session continues as a resident agent instead of a bootstrap experiment.
| Mode | Visible tools at start | Context injection | Behavioral goal |
|---|---|---|---|
| Minimal | Very few | Suppressed or reduced | Anchor the model in a concise trajectory |
| Standard | Full catalog | Normal | Maximize capability immediately |
| Anchored Standard | Few at first, then full | Blocked until promotion | Separate trajectory selection from later tool use |
The control plane is where the trick lives
Two files do most of the work: `context-gate.mjs` and `tool-bootstrap.mjs`. The first suppresses auto-injected context when the session is still unpromoted, so the model does not get flooded with the usual digests and catalogs too early. The second limits the visible tool surface to a tiny set during bootstrap, then expands it after the session crosses the promotion line.
// Conceptual flow, simplified
if (!session.promoted) {
hideAutoInjectedContext();
exposeTools(['bash', 'str_replace_editor']);
} else {
restoreStandardContext();
exposeResidentToolCatalog();
}
// Promotion is durable, not temporary
if (eventLog.includes('tool/call') || eventLog.includes('assistant/message')) {
session.promoted = true;
}
This is why the preset feels sturdier than a one-off prompt trick. It is not just asking the model to behave differently. It changes what the model can see, when it can see it, and how that state survives session lifecycle events like compaction.
Why the anchor turn matters
The hidden anchor turn is the sharpest move in the repo. Before the user’s real request arrives, the system feeds the model a controlled first exchange. That burns the first turn on purpose, so the model settles into the bootstrap trajectory before it has a chance to wander.
DSH Minimal is likely not a "stripped-down" version of Standard at all. Its real purpose is to recreate—as closely as possible—the exact environment the model experienced during V4 Pro Agent RL training. Looking back at the so-called "God-tier V4 Pro" with this context, everything starts to click.
That claim matters because it reframes Minimal mode. It is not merely a stripped-down preset. It is a behavioral reconstruction tool, designed to recreate the conditions under which the model seems to reason best.
How promotion survives real sessions
The repo does not stop at the first promotion. It tracks durable session events so the system can tell whether it is still in bootstrap mode or already resident. That matters after compaction, because the agent should not lose its phase state just because the context was summarized.
In other words, the preset is stateful. It remembers whether a session has crossed the line, and it can rebuild that judgment from the event log. That is a more serious design than a prompt wrapper, because it treats phase as a recoverable property of the session rather than a transient variable.
What this says about agent design
The comparison with standard DeepSeek Harness presets is straightforward. Standard gives you the full tool surface immediately. Minimal gives you a smaller environment that may produce better early reasoning but leaves capability on the table. `dsh-anchored-standard` tries to take the best parts of both by changing them in sequence.
| Approach | Strength | Weakness | Best use case |
|---|---|---|---|
| Minimal preset | Strong early trajectory | Low immediate utility | Anchoring the model before real work |
| Standard preset | Full capability from turn one | More context pollution | Tasks that need broad tool access immediately |
| dsh-anchored-standard | Anchors first, then expands | More moving parts | Sessions where early reasoning quality matters more than raw tool count |
That is the larger lesson. More tools can make a model worse at the start of a task. Sometimes the right intervention is not to expose more power, but to delay it until the model has stabilized.
The cost of staying clever
The maintenance-only status gives the article its ending. This is a smart system, but it exists inside a moving economic constraint. API pricing, evaluation costs, and model churn all shape whether this kind of research can keep going. The code may be elegant, but the operating environment still decides how long the experiment survives.