IASC: The Compiler for Constructed Languages

Sakana AI's open-source pipeline turns phonology, grammar, lexicon, and orthography into a reproducible language build, not a one-shot prompt.

8 min read • View on GitHub • More from SakanaAI

A wide editorial scene shows a mechanical language press turning raw phonetic symbols into a finished handbook, with sheets, gears, and symbol plates moving through a controlled machine. It explains the article's main idea: IASC treats language creation like a compiled artifact, where each stage produces an output the next stage must obey.
IASC does not ask an LLM to improvise a language in one pass. It stages the work so later decisions inherit earlier constraints.
Key Takeaways

A language that gets compiled

IASC is easiest to understand if you stop thinking about it as a chatbot and start thinking about it as a compiler. It takes language creation apart into stages, then forces each stage to commit before the next one begins.

That matters because languages are systems, not vibes. If phonology, morphosyntax, lexicon, and orthography do not agree, the result feels fake even when the individual pieces sound clever. IASC's wager is that consistency is a pipeline problem.

The repository is rooted in Sakana AI's codebase, with visible contributions from rwsproat and ctaguchi. That is a useful clue: this is not productized AI garnish, it is computational linguistics with orchestration layered on top.

Why a one-shot prompt is not enough

A single prompt can produce something language-shaped, but it is bad at staying faithful to its own rules. The model may invent a sound inventory, drift on agreement, forget a negation rule, or produce an orthography that no longer matches the phonology.

IASC addresses that by splitting the work into smaller decisions. Each decision narrows the space for the next one, which is exactly what a compiler does when it turns a loose source into a runnable artifact.

Inside the pipeline

The diagram makes the project's central idea legible at a glance. IASC is a staged build, where every output becomes the input to the next constraint.

IASC/
  agentic_phonology/data/ipas/
    Abkhaz.csv
    Ainu.txt
    Abkhaz.cnts
    Abkhaz.mod
  modular_experiments/
    controlled_morphosyntax/
      claude/french.sh
  modular_experiment_outputs_controlled/
    .../metascripts/morphosyntax_0_0_0.sh
  eval_morphosyntax.py

That tree captures the repo's logic. The phonology layer starts from IPA data and frequency artifacts, the experiment scripts assemble controlled morphosyntax, the generated metascripts lock in feature choices, and eval_morphosyntax.py gives the team a way to check whether the resulting grammar behaves like a system instead of a collage.

The hidden trick: constraints before creativity

A close-up workbench shows several locked drawers, each holding different linguistic tools and symbolic pieces, while a mechanical hand drops one rule card into a slot that seals the next drawer open. It explains how IASC uses constraints to shape creativity instead of replacing it.
The point is not to eliminate the model. The point is to fence it in so its creativity stays internally consistent.

This is where the hybrid design matters. The repository leans on classical NLP artifacts such as IPA data and finite-state style files, then lets an LLM work inside that scaffold. The result is less expressive freedom in the abstract, but more real control in the output.

That trade-off is the point. A language generator that can do anything is usually a language generator that cannot stay coherent for long. IASC gives the model a smaller room to work in, and the room has walls.

How IASC compares to other ways of making languages

WorkflowManual conlangingPrompt-only LLM generationIASC
SpeedSlow, often months or yearsFast at first, then unstableFast enough to iterate, structured enough to finish
ReproducibilityHigh only if the creator is disciplinedLow, because prompts do not guarantee repeatabilityHigher, because each stage emits a constrained artifact
Internal consistencyUsually strong, but labor intensiveOften drifts across grammar and writing systemBuilt in by chaining dependencies stage by stage
Linguistic rigorDepends on the creator's expertiseUsually shallow unless heavily guidedAnchored in IPA data, feature modules, and evaluation scripts
Creator controlMaximum control, maximum effortLoose control, high surpriseFine-grained control with less manual repetition
Best fitDeep craft projectsQuick experiments and demosResearch workflows and serious worldbuilding pipelines
Main failure modeTime sinkDrift and incoherenceSetup overhead and a steeper learning curve

The comparison is not about declaring a winner. Manual conlanging still wins on artistry, prompt-only generation still wins on speed, and IASC sits in the middle as a systems answer to a systems problem. If you need consistency at scale, the middle ground is the interesting one.

What this unlocks

The obvious use case is worldbuilding. Games, fiction franchises, and speculative projects can get a language that sounds different, behaves differently, and still obeys itself from the first page of the handbook to the last.

The deeper value is methodological. IASC shows a pattern that can travel beyond conlangs: when generation needs rigor, do not ask the model for the whole answer at once. Break the problem into constrained artifacts, let each artifact narrow the next choice, and evaluate the result before you move on.

That is why this repo feels bigger than its niche. It is a compact example of how LLMs become more useful when they stop pretending to be monoliths and start acting like components in a controlled build system.