OpenSpace: The Skill Layer That Lets AI Agents Inherit Their Own Work
A deep dive into a self-evolving middleware for Claude Code, Cursor, and other MCP agents, where skills are repaired, reused, and ranked like living code.
- OpenSpace treats agent competence as infrastructure, not as disposable prompting, so successful work can be stored, ranked, and reused.
- Its three-mode loop, FIX, DERIVED, and CAPTURED, turns runs into lineage instead of one-off executions.
- The project’s real claim is economic, not just technical: warm skills should cut tokens and increase output across repeated tasks.
- MCP is what keeps the system from becoming a walled garden, because the skill layer can serve many hosts rather than replacing them.
OpenSpace is easiest to understand if you stop thinking about prompts as the unit of progress. The repo is trying to make skills the reusable thing, which means an agent can inherit a workflow instead of rebuilding it every time. That is a bigger shift than a better prompt library.
The project’s own framing is blunt. As the README puts it, “Today's AI agents ... are powerful, but they have a critical weakness: they never Learn, Adapt, and Evolve from real-world experience — let alone Share with each other.” HKUDS/OpenSpace README
One agent learns, all agents benefit — collective intelligence at scale.
Why skills, not prompts
A prompt is cheap to write and cheap to forget. A skill is supposed to persist after the moment that discovered it. OpenSpace leans into that distinction by treating good behavior as an asset that can be selected again, instead of as a text blob buried in a repo.
That matters because agent work repeats. The same file format, the same API shape, the same recovery steps, the same boring bug class. If an agent has to reason from scratch each time, you pay for the same thinking over and over.
| Static prompt library | OpenSpace skill layer |
|---|---|
| Stores instructions | Stores reusable behavior |
| Manual reuse | Automatic selection |
| No lineage | Versioned skill graph |
| No scoring | Ranked by usage and outcome |
| Single host mindset | Works across MCP hosts |
The self-evolution loop
This is where OpenSpace stops looking like a skill catalog and starts looking like a learning system. The loop has three named outcomes: FIX, DERIVED, and CAPTURED. That vocabulary matters because it describes what happened to the skill after the run, not just what the model output looked like.
FIX repairs a broken skill. DERIVED forks a specialized version when the original is too broad. CAPTURED extracts a new pattern from a successful run and turns it into a reusable asset. Together, they make skill evolution explicit instead of accidental.
task -> skill lookup -> execution -> outcome
├─ failure -> FIX
├─ success with novelty -> CAPTURED
└─ specialization needed -> DERIVED
updated skill lineage -> future selection
How the skill graph actually works
The frontend hook tells the story clearly enough: skills are not just listed, they are turned into a graph with generation, score, and usage ratio. That makes evolution visible. You can see which skills keep getting selected, which ones branch, and which ones are fading out.
That is why the metaphor of Darwinian selection keeps showing up around the project. The system is not pretending every skill is equally useful. It is trying to measure which ones survive repeated contact with real work.
| Graph signal | What it means |
|---|---|
| Generation | How far a skill has branched from the original |
| Score | A quality signal from use or evaluation |
| Usage ratio | How often the skill is selected relative to peers |
| Lineage | Which prior skill it descended from |
| Selection | Which version gets reused next |
The architecture is straightforward, which is part of the appeal. A Python backend manages skill state, SQLite stores the graph, and a React frontend visualizes lineage. The interesting part is not the stack. It is the data model that lets the stack behave like a memory system.
What GDPVal is really measuring
The benchmark is more than a scoreboard. It is a before-and-after test for accumulated skill memory. Cold runs force the system to discover or repair the workflow. Warm runs reuse what was already learned. That is the economic claim: less repeated reasoning, fewer tokens, more output per unit of work.
On 50 professional tasks (📈 GDPVal Economic Benchmark) across 6 industries, OpenSpace agents earn 4.2× more money than baseline (ClawWork) agents using the same backbone LLM (Qwen 3.5-Plus). While cutting 46% of costly tokens through skill evolution.
| Cold run | Warm run |
|---|---|
| Finds or creates skills | Reuses evolved skills |
| Pays full reasoning cost | Pays only selection and execution cost |
| Produces a first version of the graph | Compounds the graph |
| Measures discovery | Measures reuse |
The benchmark framing also raises the right question. If a skill layer gets better every time it is used, then the value is not just in the model. It is in the accumulated memory around the model.
Why MCP matters here
OpenSpace is not trying to become the only agent you use. It is trying to become the layer that other agents can plug into. MCP is the transport that makes that possible, because it lets the skill system sit alongside Claude Code, Cursor, OpenClaw, and others instead of replacing them.
That is a subtle but important product decision. The project is not selling a new walled garden. It is selling a shared substrate for tools that already exist.
What the product shape suggests
The repo is not only a backend. It also ships a dashboard, a community cloud, and newer communication adapters like WhatsApp and Feishu. That combination hints at a larger ambition: skills as a persistent object that can move through different surfaces, not just through one CLI.
That is where OpenSpace starts to feel less like a framework and more like an operating surface. The agent can learn a skill, publish it, inherit it elsewhere, and eventually be reached through a messaging gateway.
If that sounds broad, it is. But the core idea stays narrow: make the useful part of agent work durable.
| Adjacent category | How it differs from OpenSpace |
|---|---|
| Agent framework | Runs the agent, not the skill memory |
| Prompt library | Stores instructions, not lineage |
| Workflow runner | Executes steps, not evolution |
| Tool registry | Lists tools, but does not score inherited behavior |
There is a real open question under all this. If skills evolve automatically, who audits them? A system that rewards reuse can also entrench the wrong habit if the benchmark is too narrow. The most selected skill is not always the most trustworthy one.
That tension is why OpenSpace is interesting. It is not just a promise of cheaper tokens. It is an attempt to turn agent work into a shared memory layer, with all the power and mess that implies.