OpenSpace: The Skill Layer That Lets AI Agents Inherit Their Own Work

A deep dive into a self-evolving middleware for Claude Code, Cursor, and other MCP agents, where skills are repaired, reused, and ranked like living code.

8 min read • View on GitHub • More from HKUDS

A wide black-ink illustration of a skill tree rising out of a terminal window, with branches labeled as evolving paths and several stems reinforced while others are pruned. It explains how OpenSpace treats agent skills as persistent infrastructure that can branch, survive, and spread across hosts.
OpenSpace frames competence as a living system: skills branch, survive, and get reused instead of being rediscovered from scratch.
Key Takeaways

OpenSpace is easiest to understand if you stop thinking about prompts as the unit of progress. The repo is trying to make skills the reusable thing, which means an agent can inherit a workflow instead of rebuilding it every time. That is a bigger shift than a better prompt library.

The project’s own framing is blunt. As the README puts it, “Today's AI agents ... are powerful, but they have a critical weakness: they never Learn, Adapt, and Evolve from real-world experience — let alone Share with each other.” HKUDS/OpenSpace README

One agent learns, all agents benefit — collective intelligence at scale.

xlrrrr, Maintainer (Top Contributor) · HKUDS/OpenSpace README

Why skills, not prompts

A prompt is cheap to write and cheap to forget. A skill is supposed to persist after the moment that discovered it. OpenSpace leans into that distinction by treating good behavior as an asset that can be selected again, instead of as a text blob buried in a repo.

That matters because agent work repeats. The same file format, the same API shape, the same recovery steps, the same boring bug class. If an agent has to reason from scratch each time, you pay for the same thinking over and over.

Static prompt libraryOpenSpace skill layer
Stores instructionsStores reusable behavior
Manual reuseAutomatic selection
No lineageVersioned skill graph
No scoringRanked by usage and outcome
Single host mindsetWorks across MCP hosts
A close-up black-ink scene of two hands at a workbench. One hand rewrites the same procedure on a fresh sheet, while the other pulls a card from a labeled skill archive drawer. It shows the practical difference between recomputing a workflow and reusing a battle-tested one.
The core bet is that a good workflow should be retrieved, not reinvented.

The self-evolution loop

This is where OpenSpace stops looking like a skill catalog and starts looking like a learning system. The loop has three named outcomes: FIX, DERIVED, and CAPTURED. That vocabulary matters because it describes what happened to the skill after the run, not just what the model output looked like.

The interesting part is not that a task runs. It is that the run feeds the next version of the skill graph.

FIX repairs a broken skill. DERIVED forks a specialized version when the original is too broad. CAPTURED extracts a new pattern from a successful run and turns it into a reusable asset. Together, they make skill evolution explicit instead of accidental.

task -> skill lookup -> execution -> outcome
                     ├─ failure -> FIX
                     ├─ success with novelty -> CAPTURED
                     └─ specialization needed -> DERIVED

updated skill lineage -> future selection

How the skill graph actually works

The frontend hook tells the story clearly enough: skills are not just listed, they are turned into a graph with generation, score, and usage ratio. That makes evolution visible. You can see which skills keep getting selected, which ones branch, and which ones are fading out.

That is why the metaphor of Darwinian selection keeps showing up around the project. The system is not pretending every skill is equally useful. It is trying to measure which ones survive repeated contact with real work.

Graph signalWhat it means
GenerationHow far a skill has branched from the original
ScoreA quality signal from use or evaluation
Usage ratioHow often the skill is selected relative to peers
LineageWhich prior skill it descended from
SelectionWhich version gets reused next

The architecture is straightforward, which is part of the appeal. A Python backend manages skill state, SQLite stores the graph, and a React frontend visualizes lineage. The interesting part is not the stack. It is the data model that lets the stack behave like a memory system.

What GDPVal is really measuring

The benchmark is more than a scoreboard. It is a before-and-after test for accumulated skill memory. Cold runs force the system to discover or repair the workflow. Warm runs reuse what was already learned. That is the economic claim: less repeated reasoning, fewer tokens, more output per unit of work.

On 50 professional tasks (📈 GDPVal Economic Benchmark) across 6 industries, OpenSpace agents earn 4.2× more money than baseline (ClawWork) agents using the same backbone LLM (Qwen 3.5-Plus). While cutting 46% of costly tokens through skill evolution.

xlrrrr, Maintainer (Top Contributor) · HKUDS/OpenSpace README
Cold runWarm run
Finds or creates skillsReuses evolved skills
Pays full reasoning costPays only selection and execution cost
Produces a first version of the graphCompounds the graph
Measures discoveryMeasures reuse

The benchmark framing also raises the right question. If a skill layer gets better every time it is used, then the value is not just in the model. It is in the accumulated memory around the model.

Why MCP matters here

OpenSpace is not trying to become the only agent you use. It is trying to become the layer that other agents can plug into. MCP is the transport that makes that possible, because it lets the skill system sit alongside Claude Code, Cursor, OpenClaw, and others instead of replacing them.

That is a subtle but important product decision. The project is not selling a new walled garden. It is selling a shared substrate for tools that already exist.

What the product shape suggests

The repo is not only a backend. It also ships a dashboard, a community cloud, and newer communication adapters like WhatsApp and Feishu. That combination hints at a larger ambition: skills as a persistent object that can move through different surfaces, not just through one CLI.

That is where OpenSpace starts to feel less like a framework and more like an operating surface. The agent can learn a skill, publish it, inherit it elsewhere, and eventually be reached through a messaging gateway.

If that sounds broad, it is. But the core idea stays narrow: make the useful part of agent work durable.

Adjacent categoryHow it differs from OpenSpace
Agent frameworkRuns the agent, not the skill memory
Prompt libraryStores instructions, not lineage
Workflow runnerExecutes steps, not evolution
Tool registryLists tools, but does not score inherited behavior

There is a real open question under all this. If skills evolve automatically, who audits them? A system that rewards reuse can also entrench the wrong habit if the benchmark is too narrow. The most selected skill is not always the most trustworthy one.

That tension is why OpenSpace is interesting. It is not just a promise of cheaper tokens. It is an attempt to turn agent work into a shared memory layer, with all the power and mess that implies.