IASC: The Compiler for Constructed Languages
Sakana AI's open-source pipeline turns phonology, grammar, lexicon, and orthography into a reproducible language build, not a one-shot prompt.
- IASC turns conlanging into a staged build process that compiles phonology, grammar, lexicon, orthography, and documentation into one coherent artifact.
- The repo's real trick is not generation alone, but the way classical constraints and shell orchestration keep the model from drifting.
- Each stage emits an artifact that the next stage must respect, which makes the final language more reproducible than prompt-only workflows.
- The project matters most where consistency beats improvisation, especially in worldbuilding, games, and research tooling.
A language that gets compiled
IASC is easiest to understand if you stop thinking about it as a chatbot and start thinking about it as a compiler. It takes language creation apart into stages, then forces each stage to commit before the next one begins.
That matters because languages are systems, not vibes. If phonology, morphosyntax, lexicon, and orthography do not agree, the result feels fake even when the individual pieces sound clever. IASC's wager is that consistency is a pipeline problem.
The repository is rooted in Sakana AI's codebase, with visible contributions from rwsproat and ctaguchi. That is a useful clue: this is not productized AI garnish, it is computational linguistics with orchestration layered on top.
Why a one-shot prompt is not enough
A single prompt can produce something language-shaped, but it is bad at staying faithful to its own rules. The model may invent a sound inventory, drift on agreement, forget a negation rule, or produce an orthography that no longer matches the phonology.
IASC addresses that by splitting the work into smaller decisions. Each decision narrows the space for the next one, which is exactly what a compiler does when it turns a loose source into a runnable artifact.
Inside the pipeline
IASC/
agentic_phonology/data/ipas/
Abkhaz.csv
Ainu.txt
Abkhaz.cnts
Abkhaz.mod
modular_experiments/
controlled_morphosyntax/
claude/french.sh
modular_experiment_outputs_controlled/
.../metascripts/morphosyntax_0_0_0.sh
eval_morphosyntax.py
That tree captures the repo's logic. The phonology layer starts from IPA data and frequency artifacts, the experiment scripts assemble controlled morphosyntax, the generated metascripts lock in feature choices, and eval_morphosyntax.py gives the team a way to check whether the resulting grammar behaves like a system instead of a collage.
The hidden trick: constraints before creativity
This is where the hybrid design matters. The repository leans on classical NLP artifacts such as IPA data and finite-state style files, then lets an LLM work inside that scaffold. The result is less expressive freedom in the abstract, but more real control in the output.
That trade-off is the point. A language generator that can do anything is usually a language generator that cannot stay coherent for long. IASC gives the model a smaller room to work in, and the room has walls.
How IASC compares to other ways of making languages
| Workflow | Manual conlanging | Prompt-only LLM generation | IASC |
|---|---|---|---|
| Speed | Slow, often months or years | Fast at first, then unstable | Fast enough to iterate, structured enough to finish |
| Reproducibility | High only if the creator is disciplined | Low, because prompts do not guarantee repeatability | Higher, because each stage emits a constrained artifact |
| Internal consistency | Usually strong, but labor intensive | Often drifts across grammar and writing system | Built in by chaining dependencies stage by stage |
| Linguistic rigor | Depends on the creator's expertise | Usually shallow unless heavily guided | Anchored in IPA data, feature modules, and evaluation scripts |
| Creator control | Maximum control, maximum effort | Loose control, high surprise | Fine-grained control with less manual repetition |
| Best fit | Deep craft projects | Quick experiments and demos | Research workflows and serious worldbuilding pipelines |
| Main failure mode | Time sink | Drift and incoherence | Setup overhead and a steeper learning curve |
The comparison is not about declaring a winner. Manual conlanging still wins on artistry, prompt-only generation still wins on speed, and IASC sits in the middle as a systems answer to a systems problem. If you need consistency at scale, the middle ground is the interesting one.
What this unlocks
The obvious use case is worldbuilding. Games, fiction franchises, and speculative projects can get a language that sounds different, behaves differently, and still obeys itself from the first page of the handbook to the last.
The deeper value is methodological. IASC shows a pattern that can travel beyond conlangs: when generation needs rigor, do not ask the model for the whole answer at once. Break the problem into constrained artifacts, let each artifact narrow the next choice, and evaluate the result before you move on.
That is why this repo feels bigger than its niche. It is a compact example of how LLMs become more useful when they stop pretending to be monoliths and start acting like components in a controlled build system.