`os-taxonomy`: The Open Curriculum Graph That Treats Learning Like Infrastructure

A machine-checkable map of primary education, with prerequisites, standards mapping, semantic topic types, and legal provenance baked into the data model.

8 min read • View on GitHub • More from withmarbleapp

A wide editorial scene showing a monumental control panel made of topic nodes, arrows, stamps, and validation rails. It explains that the repository treats education like a machine-readable infrastructure layer rather than a flat curriculum document.
Marble turns curriculum into a graph that can be validated, queried, and reused by software.
Key Takeaways

There are plenty of curriculum datasets. This is not one of them. os-taxonomy is built like a substrate: typed nodes, explicit prerequisite edges, standards mappings, and a validation script that behaves less like housekeeping and more like a compiler.

That is why the repo feels unusual. It does not just describe what children should learn. It encodes what can depend on what, what can be said about external standards, and what must fail validation before the data ships.

Why this feels like a kernel, not a syllabus

The best way to understand Marble’s taxonomy is to stop thinking about content and start thinking about infrastructure. A syllabus is a document. A kernel is a coordinating layer. This repository is trying to be the latter for elementary education.

이 저장소는 초등 학습 쪽에서 거의 `AI 치트키`처럼 보일 수도 있음 👀 `os-taxonomy`는 8개 주요 과목을 1590개 핵심 개념과 3221개 논리 연결로 쪼개서, `무엇을 먼저 배우고, 무엇을 나중에 배워야 하는지`를 계산 가능한 지식 그래프로 만들어줌. 중요한 건 이게 그냥 커리큘럼

AI 카페인 ☕️, AI_Caffeine · @AI_Caffeine on X

That reaction gets the instinct right, even if it stops short of the architecture. The point is not that the graph is large. The point is that the graph is operational. It is meant to support sequencing, mastery checks, curriculum alignment, and any product that needs to know what comes next.

What lives in the graph

A topic becomes computable only after it is connected to prerequisites, standards, and validation rules.

The repository is organized like a data product, not a content dump. The core lives in JSON files, with schemas and scripts enforcing the rules around them. That structure matters because every layer has a job: topics describe nodes, dependencies describe edges, standards map the graph to outside frameworks, and validation checks that the whole thing still makes sense.

{
  "topic": {
    "id": "math.fractions.compare",
    "type": "CONCEPTUAL",
    "ageRange": "7-8",
    "evidence": ["Compares unit fractions using visual models"],
    "assessmentPrompt": "Can the learner explain which fraction is larger and why?"
  },
  "dependency": {
    "prerequisiteId": "math.number.sense.part-whole",
    "kind": "hard",
    "reason": "Must understand part-whole relationships before comparing fractions"
  }
}

That shape tells you a lot. A topic is not just a label. It carries pedagogical evidence, a prompt, a type, and a place in the graph. A dependency is not just an arrow. It has a kind and a reason, which makes the learning path explainable instead of opaque.

The strange power of codes-only

A close-up black-ink scene showing a standards document being reduced to a code token, passing through a validation gate, and emerging as structured metadata attached to a topic node. It explains how legal constraints become part of the data pipeline.
Marble turns licensing constraints into a design feature by separating codes from verbatim standards text when needed.

This is the most subtle part of the repo. Educational standards often live in a legal gray zone, where identifiers are reusable but verbatim text can be restricted. Marble handles that by separating factual codes from text-bearing records, then enforcing the distinction in validation.

我的天神爷,有人居然把小学课程开源知识图谱化了,后续想象空间太大了😮🧠 想学习和开发AI教育工具、个性化学习App,却苦于没有结构化、靠谱的小学课程数据?手动整理太累,买商业数据又贵?😭 今天挖到这个 https://t.co/Y1za6fEpKk https://t.co/QGP7twgpW6

Denzii 🕊️, denziideng · @denziideng on X

That matters for two reasons. First, it keeps the repository open without pretending every source is equally open. Second, it makes reuse realistic. A downstream builder can rely on the identifiers and graph structure even when the source text must stay out of the bundle.

How the prerequisite machine works

The graph is a directed acyclic graph, which is exactly the right shape for learning. You want a path forward. You do not want cycles that let a learner depend on knowledge that only exists after the thing it is supposed to unlock.

The hard and soft prerequisite split is especially useful. Hard edges say a topic is genuinely blocked without the earlier concept. Soft edges are recommendations, context, or likely support skills. That difference gives the graph a sense of priority without flattening pedagogy into a single binary.

The reason field is the quiet breakthrough. It makes the dependency legible to humans. A machine can sort the graph. A teacher, product manager, or curriculum designer can read why the edge exists.

LayerFlat curriculum document`os-taxonomy`
StructureStatic, linear, hard to queryGraph of typed nodes and edges
PrerequisitesImplicit or buried in proseExplicit hard and soft dependencies
Standards mappingUsually embedded in tables or appendicesSeparate mapping layer tied to nodes
Legal provenanceRarely machine-enforcedCodes-only rules checked in validation
Reuse by developersManual extractionDirect ingestion into software
ExplainabilityLowHigh, because every edge can carry a reason

Why the schema matters more than the data dump

A large JSON file is not infrastructure by itself. The schema is what gives the data a contract. Validation is what keeps that contract honest. In this repo, those two things do the work that an API server would normally do for a live product.

import crypto from 'node:crypto';
import fs from 'node:fs';

const data = JSON.parse(fs.readFileSync('data/topics.json', 'utf8'));
const hash = crypto.createHash('sha256').update(JSON.stringify(data)).digest('hex');

if (hash !== manifest.topicsSha256) {
  throw new Error('topics.json checksum mismatch');
}

for (const edge of dependencies) {
  if (!topicsById.has(edge.prerequisiteId)) {
    throw new Error(`Missing prerequisite: ${edge.prerequisiteId}`);
  }
}

That is the compiler mindset. Validate references. Check integrity. Fail fast. The effect is bigger than correctness. It tells downstream systems they can trust the graph enough to automate on top of it.

What makes this useful for builders

If you are building tutoring, assessment, sequencing, or analytics software, this repo saves you from a brutal first step: inventing your own curriculum backbone. It gives you a graph that is already structured for dependency resolution and standards alignment.

Builder use caseWhat the graph enables
Adaptive tutoringChoose the next topic only when prerequisites are satisfied
Assessment generationMap prompts to specific skill nodes
Curriculum alignmentTrace a local lesson plan back to external standards
Learner analyticsSpot bottlenecks and central topics
PersonalizationBranch along hard and soft paths without losing structure

That is why the taxonomy feels more like substrate than product. It does not answer every learning question. It gives other systems a reliable map to answer them with.

The bigger bet

The strongest version of the Marble thesis is simple. Learning content can be versioned, validated, and reasoned about like software. Once that is true, the curriculum stops being a PDF you consult and becomes an operating layer you can build on.

That is a serious bet. It suggests that education infrastructure should be open, typed, provenance-aware, and machine-checkable from the start. If Marble is right, the next generation of EdTech will not begin with lesson plans. It will begin with graphs.