Inside MiroFish: The Open-Source Engine for Predictive Social Simulation

While most AI agents are built to execute tasks, MiroFish uses GraphRAG and multi-agent swarms to simulate how thousands of humans will react to new information.

9 min read • 666ghj/MiroFish

A large antique glass terrarium resting on a wooden desk. Inside the terrarium is a miniature, bustling city square with tiny figures interacting. A large magnifying glass hovers outside the glass, focusing light onto one specific interaction representing predictive emergence.
MiroFish flips the multi-agent paradigm from executing discrete tasks to observing emergent social behavior in a contained environment.
Editorial portrait of Guo Hangjiang

I was really excited at first. After reaching 10k, I kind of lost the feeling.

— Guo Hangjiang, Creator. blocmates
Key Takeaways

The Shift to Simulative AI

Most multi-agent frameworks are designed for execution. You give them a goal, and they write code, browse the web, or book a flight. MiroFish flips this paradigm entirely. It is designed for prediction. By combining GraphRAG with thousands of autonomous agents, it creates a digital twin of a social network to forecast how public opinion, market trends, or rumors will spread.

Asking a single large language model what will happen in a complex social scenario usually fails. The model averages out human behavior into a bland consensus. MiroFish takes a different approach. It builds 1,000 distinct agents, gives them different personalities and biases, and watches them interact. The story here is the shift from Generative AI to Simulative AI. It operates as a pre-rehearsal laboratory for real-world social dynamics.

Grounding the Swarm with GraphRAG

The biggest risk in running a massive simulation is hallucination. If agents invent their own facts, the simulation derails. MiroFish prevents this by grounding its swarm using Zep Cloud and a GraphRAG architecture. The `graph_builder.py` service handles the ingestion of raw text (seed documents like news reports or literature) and transforms it into a structured knowledge graph.

This graph acts as the physical laws of the simulation. Without this Extract, Transform, and Load (ETL) layer, the agents would have no world knowledge or relational context to act upon. The system uses an `InsightForge` pattern to perform multi-dimensional retrieval. It simulates requirements and generates relationship chains so the agents understand exactly how entities are connected over time.

A five-step circular lifecycle of a MiroFish simulation. The nodes are 1. Ingestion (Seed Data)

Hallucinating Reality

Converting a static graph node into a living agent requires a translation layer. The `oasis_profile_generator.py` script uses a language model to hallucinate rich personas based on minimal graph data. It assigns each agent a profession, a Myers-Briggs personality type, and a social standing.

The configuration logic is deeply culturally grounded. For example, the engine includes a hardcoded Chinese timezone configuration that defines dead hours (midnight to 5 AM) and peak hours (7 PM to 10 PM). This ensures the simulated agents sleep, post, and react on schedules that mimic real human populations, rather than acting as tireless bots.

A close-up of a drafting table where a mechanical hand uses a compass and fountain pen to draw detailed human faces onto blank wooden game pawns.
The profile generator translates static database nodes into rich, opinionated personas with distinct daily routines and biases.

The File-System Lifeline

Social simulations are computationally expensive and can run for hours. If the web server managing the UI crashes, the simulation should not die with it. To solve this, MiroFish uses a file-based Inter-Process Communication (IPC) architecture defined in `simulation_ipc.py`.

This mechanism decouples the Flask backend from the heavy simulation subprocess. The system uses an asynchronous command and response pattern via the file system. The Flask app drops JSON commands into a directory, and the simulation engine polls that directory, executes the command, and writes a response file back. It is a pragmatic, battle-tested pattern that keeps massive simulations stable.

Synthesizing the Chaos

Watching thousands of agents argue on a simulated timeline is overwhelming. To make the data useful, MiroFish employs a specialized `ReportAgent`. This agent uses the ReACT (Reasoning and Acting) pattern to observe the simulation, interview key node clusters, and write a final analytical brief.

Framework Core Focus Architecture Model Primary Output
MiroFish Predictive emergence GraphRAG + OASIS swarm Analytical forecasts and social simulation reports
Hive Self-healing execution Queen and Worker hierarchy Completed digital tasks and refined agent logic
ChatDev Software factory Sequential waterfall simulation Compiled code repositories and software documentation
Stanford Smallville Academic observation Memory stream and reflection Qualitative behavioral research and 2D sandbox logs

From Stanford to Beijing

MiroFish is a direct spiritual successor to the famous Stanford Smallville experiment, which proved that agents could maintain memory streams and live simulated lives. Created by Guo Hangjiang, an undergraduate student at Beijing University of Posts and Telecommunications, the project launched in March 2026 and rapidly gained traction in the open-source community.

Backed by institutional investment, the engine is now being adapted for everything from literary analysis to financial market forecasting. By focusing on consequences rather than task completion, MiroFish offers developers a zero-risk environment to test how reality might unfold.


Sources: Codebase analysis of 666ghj/MiroFish; architectural documentation; and interviews with the creator via blocmates.