`os-taxonomy`: The Open Curriculum Graph That Treats Learning Like Infrastructure
A machine-checkable map of primary education, with prerequisites, standards mapping, semantic topic types, and legal provenance baked into the data model.
- `os-taxonomy` treats learning as a validated graph, not a static syllabus, so downstream software can reason about sequencing instead of just displaying topics.
- The repo’s real novelty is not scale alone, but the combination of prerequisite logic, standards mapping, semantic topic types, and machine checks that keep the graph trustworthy.
- Its codes-only provenance pattern turns licensing constraints into part of the architecture, which makes the data safer to reuse in open and commercial products.
- For builders, the dataset reads like a substrate for tutoring, assessment, and personalization systems, not a finished curriculum product.
There are plenty of curriculum datasets. This is not one of them. os-taxonomy is built like a substrate: typed nodes, explicit prerequisite edges, standards mappings, and a validation script that behaves less like housekeeping and more like a compiler.
That is why the repo feels unusual. It does not just describe what children should learn. It encodes what can depend on what, what can be said about external standards, and what must fail validation before the data ships.
Why this feels like a kernel, not a syllabus
The best way to understand Marble’s taxonomy is to stop thinking about content and start thinking about infrastructure. A syllabus is a document. A kernel is a coordinating layer. This repository is trying to be the latter for elementary education.
이 저장소는 초등 학습 쪽에서 거의 `AI 치트키`처럼 보일 수도 있음 👀 `os-taxonomy`는 8개 주요 과목을 1590개 핵심 개념과 3221개 논리 연결로 쪼개서, `무엇을 먼저 배우고, 무엇을 나중에 배워야 하는지`를 계산 가능한 지식 그래프로 만들어줌. 중요한 건 이게 그냥 커리큘럼
That reaction gets the instinct right, even if it stops short of the architecture. The point is not that the graph is large. The point is that the graph is operational. It is meant to support sequencing, mastery checks, curriculum alignment, and any product that needs to know what comes next.
What lives in the graph
The repository is organized like a data product, not a content dump. The core lives in JSON files, with schemas and scripts enforcing the rules around them. That structure matters because every layer has a job: topics describe nodes, dependencies describe edges, standards map the graph to outside frameworks, and validation checks that the whole thing still makes sense.
{
"topic": {
"id": "math.fractions.compare",
"type": "CONCEPTUAL",
"ageRange": "7-8",
"evidence": ["Compares unit fractions using visual models"],
"assessmentPrompt": "Can the learner explain which fraction is larger and why?"
},
"dependency": {
"prerequisiteId": "math.number.sense.part-whole",
"kind": "hard",
"reason": "Must understand part-whole relationships before comparing fractions"
}
}
That shape tells you a lot. A topic is not just a label. It carries pedagogical evidence, a prompt, a type, and a place in the graph. A dependency is not just an arrow. It has a kind and a reason, which makes the learning path explainable instead of opaque.
The strange power of codes-only
This is the most subtle part of the repo. Educational standards often live in a legal gray zone, where identifiers are reusable but verbatim text can be restricted. Marble handles that by separating factual codes from text-bearing records, then enforcing the distinction in validation.
我的天神爷,有人居然把小学课程开源知识图谱化了,后续想象空间太大了😮🧠 想学习和开发AI教育工具、个性化学习App,却苦于没有结构化、靠谱的小学课程数据?手动整理太累,买商业数据又贵?😭 今天挖到这个 https://t.co/Y1za6fEpKk https://t.co/QGP7twgpW6
That matters for two reasons. First, it keeps the repository open without pretending every source is equally open. Second, it makes reuse realistic. A downstream builder can rely on the identifiers and graph structure even when the source text must stay out of the bundle.
How the prerequisite machine works
The graph is a directed acyclic graph, which is exactly the right shape for learning. You want a path forward. You do not want cycles that let a learner depend on knowledge that only exists after the thing it is supposed to unlock.
The hard and soft prerequisite split is especially useful. Hard edges say a topic is genuinely blocked without the earlier concept. Soft edges are recommendations, context, or likely support skills. That difference gives the graph a sense of priority without flattening pedagogy into a single binary.
The reason field is the quiet breakthrough. It makes the dependency legible to humans. A machine can sort the graph. A teacher, product manager, or curriculum designer can read why the edge exists.
| Layer | Flat curriculum document | `os-taxonomy` |
|---|---|---|
| Structure | Static, linear, hard to query | Graph of typed nodes and edges |
| Prerequisites | Implicit or buried in prose | Explicit hard and soft dependencies |
| Standards mapping | Usually embedded in tables or appendices | Separate mapping layer tied to nodes |
| Legal provenance | Rarely machine-enforced | Codes-only rules checked in validation |
| Reuse by developers | Manual extraction | Direct ingestion into software |
| Explainability | Low | High, because every edge can carry a reason |
Why the schema matters more than the data dump
A large JSON file is not infrastructure by itself. The schema is what gives the data a contract. Validation is what keeps that contract honest. In this repo, those two things do the work that an API server would normally do for a live product.
import crypto from 'node:crypto';
import fs from 'node:fs';
const data = JSON.parse(fs.readFileSync('data/topics.json', 'utf8'));
const hash = crypto.createHash('sha256').update(JSON.stringify(data)).digest('hex');
if (hash !== manifest.topicsSha256) {
throw new Error('topics.json checksum mismatch');
}
for (const edge of dependencies) {
if (!topicsById.has(edge.prerequisiteId)) {
throw new Error(`Missing prerequisite: ${edge.prerequisiteId}`);
}
}
That is the compiler mindset. Validate references. Check integrity. Fail fast. The effect is bigger than correctness. It tells downstream systems they can trust the graph enough to automate on top of it.
What makes this useful for builders
If you are building tutoring, assessment, sequencing, or analytics software, this repo saves you from a brutal first step: inventing your own curriculum backbone. It gives you a graph that is already structured for dependency resolution and standards alignment.
| Builder use case | What the graph enables |
|---|---|
| Adaptive tutoring | Choose the next topic only when prerequisites are satisfied |
| Assessment generation | Map prompts to specific skill nodes |
| Curriculum alignment | Trace a local lesson plan back to external standards |
| Learner analytics | Spot bottlenecks and central topics |
| Personalization | Branch along hard and soft paths without losing structure |
That is why the taxonomy feels more like substrate than product. It does not answer every learning question. It gives other systems a reliable map to answer them with.
The bigger bet
The strongest version of the Marble thesis is simple. Learning content can be versioned, validated, and reasoned about like software. Once that is true, the curriculum stops being a PDF you consult and becomes an operating layer you can build on.
That is a serious bet. It suggests that education infrastructure should be open, typed, provenance-aware, and machine-checkable from the start. If Marble is right, the next generation of EdTech will not begin with lesson plans. It will begin with graphs.





