Shachi: A Laboratory for LLM Societies
Sakana AI’s framework turns brittle agent-based modeling into repeatable experiments, with tools, memory, and a second parsing pass when models return messy output.
- Shachi matters because it treats memory as an experimental variable, not a convenience feature.
- Its four-part agent contract makes LLM societies easier to swap, compare, and reproduce across benchmark worlds.
- The two-step parsing fallback is a quiet admission that models fail, and the framework is built to recover without breaking the run.
- Shachi is narrower than general agent frameworks, which is exactly why it is more useful for controlled simulation.
Most agent frameworks optimize for breadth. Shachi optimizes for control. That sounds modest until you notice what it unlocks: a simulation that can be paused, resumed, compared, and even carried into a new run without losing its earlier state.
Shachi is a modular framework designed to simplify building **LLM-based agents** for **Agent-Based Modeling (ABM)**. By separating each agent into four core components—**LLM**, **Tools**, **Memory**, and **Configuration**—Shachi enables reproducible experiments across a variety of social, economic, and cognitive simulation tasks.
A society that remembers
The most interesting thing in Shachi is not that agents talk. It is that they can remember across lives. In scripts/main_carrying_memory.py, memory can persist from one simulation to the next, which turns a run into a longitudinal experiment instead of a disposable demo.
That matters because Shachi comes from Sakana AI, but the repo itself is clearly shaped like a research instrument, not a broad community platform. The codebase is modular, typed, and opinionated, with Hydra for configuration, pydantic for structure, litellm for model routing, and uv for modern dependency management. The footprint is small enough to feel focused, which is exactly why the framework reads like an experiment designed to survive scrutiny.
Why LLM ABM gets brittle fast
Agent-based modeling gets fragile the moment the model stops behaving like a tidy API. A response drifts out of schema, a tool call fails, or one benchmark needs a different prompt shape, and suddenly the whole simulation becomes hard to compare with anything else. Shachi is built around that failure mode, which is why its value is less about flash and more about disciplined repetition.
| Dimension | General-purpose agent frameworks | One-off ABM repos | Shachi |
|---|---|---|---|
| Primary goal | Build autonomous assistants and apps | Reproduce a single paper or benchmark | Run controlled LLM societies as repeatable experiments |
| Interface | Broad, flexible abstractions | Custom code around one task | Gym-style agent, environment, tool, memory contract |
| Reproducibility | Depends on the builder | Often fragile outside the original notebook | Standardized configs and structured outputs |
| Memory | Often app-centric or external | Usually ad hoc | First-class and persistable across runs |
| Output handling | Varies by framework | Usually assumes the model behaves | Can fall back to a second parsing pass |
| Best fit | Product prototypes | One-off automations | Research on emergent collective behavior |
The four-core contract
Shachi’s architecture is simple on purpose. Each agent is defined by four parts, LLM, Tools, Memory, and Configuration, and each environment speaks a common contract back to the agent. That is the trick, because once the interfaces are stable, the framework can swap in different benchmark worlds without rewriting the entire simulation loop.
That design shows up in the execution loop too. The main runner uses asynchronous steps and batching, which means the framework can scale to many agents without turning the code into a tangle of special cases. For a simulation tool, that is the right kind of boring.
One agent, many lives
The carrying-memory flow is Shachi’s sharpest idea. Instead of treating each run as an isolated sandbox, it lets an agent accumulate state and return with something intact, which opens the door to studying adaptation, transfer, and culture-like effects across disconnected simulations.
- StockAgent
- PsychoBench
- LM_Caricature
- Cognitive Biases
- EmotionBench
That benchmark spread matters because it is not one world with one metric. It is a shared contract applied to different social and cognitive settings, from market behavior to online forum dynamics to psychological tasks. Once the same agent logic can move between those worlds, the question stops being whether the model can chat and becomes what kind of society it produces under pressure.
When the model hands you broken JSON
This is where Shachi feels production-minded. In src/shachi/agent/sotopia.py, the framework supports several parsing modes, including a two_steps_structured_output path that sends bad output through a second cleaning pass. If the model returns markdown, stray prose, or malformed structure, the simulation does not have to die with it.
That fallback is the difference between a demo and an instrument. It says the framework expects stochastic behavior, then contains it just enough to keep the experiment coherent.
How Shachi compares
Shachi does not try to beat LangChain or AutoGPT at being broad. It narrows the problem until the hard parts become visible, then puts guardrails around them. That makes it less useful for generic agent apps, and more useful for the kind of research where you want to know whether behavior changed because the world changed, not because the plumbing broke.
Compared with one-off academic repos, Shachi is also more reusable. The benchmark worlds differ, but the contract stays put, which is the point of the whole project. It turns a pile of scenario code into a comparative lab for LLM societies.