Shachi: A Laboratory for LLM Societies

Sakana AI’s framework turns brittle agent-based modeling into repeatable experiments, with tools, memory, and a second parsing pass when models return messy output.

11 min read • View on GitHub • More from SakanaAI

A white laboratory bench holds several transparent simulation chambers, each containing a tiny cluster of abstract figures and tokens. A small memory capsule moves from one chamber to the next, showing that something from one run can survive into the next experiment. The image explains that Shachi is designed for repeatable social simulations, not one-off chat demos.
Shachi treats memory as a transferable experimental variable. That makes each run less like a disposable prompt and more like a controlled society you can revisit.
Key Takeaways

Most agent frameworks optimize for breadth. Shachi optimizes for control. That sounds modest until you notice what it unlocks: a simulation that can be paused, resumed, compared, and even carried into a new run without losing its earlier state.

Shachi is a modular framework designed to simplify building **LLM-based agents** for **Agent-Based Modeling (ABM)**. By separating each agent into four core components—**LLM**, **Tools**, **Memory**, and **Configuration**—Shachi enables reproducible experiments across a variety of social, economic, and cognitive simulation tasks.

SakanaAI/shachi Repository, Project Documentation · SakanaAI/shachi README

A society that remembers

The most interesting thing in Shachi is not that agents talk. It is that they can remember across lives. In scripts/main_carrying_memory.py, memory can persist from one simulation to the next, which turns a run into a longitudinal experiment instead of a disposable demo.

A hedcut-style portrait of Kuroki1931 rendered in black ink on white. It grounds the origin story in the main contributor visible in the repository history, rather than in an abstract organization name.

That matters because Shachi comes from Sakana AI, but the repo itself is clearly shaped like a research instrument, not a broad community platform. The codebase is modular, typed, and opinionated, with Hydra for configuration, pydantic for structure, litellm for model routing, and uv for modern dependency management. The footprint is small enough to feel focused, which is exactly why the framework reads like an experiment designed to survive scrutiny.

Why LLM ABM gets brittle fast

Agent-based modeling gets fragile the moment the model stops behaving like a tidy API. A response drifts out of schema, a tool call fails, or one benchmark needs a different prompt shape, and suddenly the whole simulation becomes hard to compare with anything else. Shachi is built around that failure mode, which is why its value is less about flash and more about disciplined repetition.

DimensionGeneral-purpose agent frameworksOne-off ABM reposShachi
Primary goalBuild autonomous assistants and appsReproduce a single paper or benchmarkRun controlled LLM societies as repeatable experiments
InterfaceBroad, flexible abstractionsCustom code around one taskGym-style agent, environment, tool, memory contract
ReproducibilityDepends on the builderOften fragile outside the original notebookStandardized configs and structured outputs
MemoryOften app-centric or externalUsually ad hocFirst-class and persistable across runs
Output handlingVaries by frameworkUsually assumes the model behavesCan fall back to a second parsing pass
Best fitProduct prototypesOne-off automationsResearch on emergent collective behavior

The four-core contract

Shachi’s architecture is simple on purpose. Each agent is defined by four parts, LLM, Tools, Memory, and Configuration, and each environment speaks a common contract back to the agent. That is the trick, because once the interfaces are stable, the framework can swap in different benchmark worlds without rewriting the entire simulation loop.

Shachi standardizes the agent-environment loop so the same core logic can run across multiple benchmark worlds, while a parser fallback catches messy model output before it breaks the experiment.

That design shows up in the execution loop too. The main runner uses asynchronous steps and batching, which means the framework can scale to many agents without turning the code into a tangle of special cases. For a simulation tool, that is the right kind of boring.

One agent, many lives

The carrying-memory flow is Shachi’s sharpest idea. Instead of treating each run as an isolated sandbox, it lets an agent accumulate state and return with something intact, which opens the door to studying adaptation, transfer, and culture-like effects across disconnected simulations.

That benchmark spread matters because it is not one world with one metric. It is a shared contract applied to different social and cognitive settings, from market behavior to online forum dynamics to psychological tasks. Once the same agent logic can move between those worlds, the question stops being whether the model can chat and becomes what kind of society it produces under pressure.

When the model hands you broken JSON

This is where Shachi feels production-minded. In src/shachi/agent/sotopia.py, the framework supports several parsing modes, including a two_steps_structured_output path that sends bad output through a second cleaning pass. If the model returns markdown, stray prose, or malformed structure, the simulation does not have to die with it.

A close-up shows a tangled mass of malformed output, torn paper edges, stray brackets, and scattered markdown fragments being pressed through a narrow mechanical sorter. On the other side, the same material emerges as a clean structured record. The image explains Shachi’s two-step parsing fallback, which turns unreliable model output into schema-compliant data.
Shachi does not assume the model will behave. It builds a recovery path so the run can keep going when the first output is messy.

That fallback is the difference between a demo and an instrument. It says the framework expects stochastic behavior, then contains it just enough to keep the experiment coherent.

How Shachi compares

Shachi does not try to beat LangChain or AutoGPT at being broad. It narrows the problem until the hard parts become visible, then puts guardrails around them. That makes it less useful for generic agent apps, and more useful for the kind of research where you want to know whether behavior changed because the world changed, not because the plumbing broke.

Compared with one-off academic repos, Shachi is also more reusable. The benchmark worlds differ, but the contract stays put, which is the point of the whole project. It turns a pile of scenario code into a comparative lab for LLM societies.