The Digital Courtroom: How BettaFish Uses Agentic Debate to Solve the Hallucination Problem
Moving beyond linear pipelines, this framework-less system reconstructs public opinion by forcing specialized AI agents to challenge each other's biases.

始于舆情,而不止于舆情。
- BettaFish uses a simulated forum of specialized agents to eliminate the artificial consensus and hallucinations found in linear AI pipelines.
- The system avoids heavy agentic frameworks in favor of a raw Python state machine to ensure total transparency in the reasoning loop.
- K-Means clustering and semantic sampling allow the engine to feed a diverse cross-section of data into the context window without missing minority opinions.
- Five distinct modules coordinate to transform raw multi-modal data from social platforms into professional research reports.
The Hallucination of Consensus
The current era of artificial intelligence is obsessed with linear pipelines. Data goes in, an autonomous agent processes it, and a summary comes out. This works for simple tasks. It fails catastrophically for public opinion.
Social media is inherently noisy and contradictory. When a single large language model attempts to summarize a trending event, it acts as a yes-man. It averages out the noise, creating an artificial consensus that strips away the actual tension of the discourse. BettaFish, an open-source public opinion analysis system, rejects this linear approach entirely.
Instead of a single smart prompt, BettaFish implements a Simulated Forum. It forces specialized agents to debate conflicting social media signals before reaching a conclusion. By treating intelligence as a Society of Mind rather than a solitary oracle, it actively uses conflict to mitigate hallucination.
Five Engines, Zero Frameworks
Building a multi-agent system usually involves reaching for heavy frameworks like LangChain or CrewAI. BettaFish deliberately avoids them. The entire architecture is a raw Python implementation of state machines.
The system is divided into five highly specialized modules. The MindSpider handles raw data acquisition across dozens of platforms, from Weibo to TikTok. The Query Engine executes external web searches, while the Media Engine analyzes multi-modal content like short videos. The Insight Engine mines internal databases. Finally, the Report Engine formats the debated consensus into professional Markdown or PDF documents.
This framework-less approach solves a major pain point in agentic development: debugging. By avoiding deep library abstractions, developers can see exactly how prompts are constructed and how state dictionaries are passed between nodes. It provides total transparency into the reasoning loop.
Solving the Context Window with K-Means
Gathering data is easy. Feeding it to an LLM is hard. A typical trending event might generate thousands of posts across different platforms. Pushing all of that into a context window is both expensive and ineffective due to the needle-in-a-haystack problem.
BettaFish solves this in the Insight Engine using Semantic Representative Sampling. Rather than simply truncating the data chronologically, it uses SentenceTransformers to embed the search results into a vector space. It then applies K-Means clustering to group similar opinions together.
The agent then samples the hottest content from each cluster. This ensures the LLM receives a perfectly balanced, diverse cross-section of public opinion, acting as a brilliant bridge between traditional machine learning and modern generative AI.
| Approach | Data Selection | Risk Profile | Best For |
|---|---|---|---|
| Linear Truncation | Top 50 chronological results. | High risk of missing minority opinions. | Simple keyword alerts. |
| Semantic Sampling (BettaFish) | One representative from each K-Means cluster. | Requires local embedding compute. | Complex crisis management. |
From Classroom to Scale
The sophistication of the architecture belies its origin. The project was initially created by BaiFu, a university student, as a course assignment for public opinion monitoring. It quickly evolved beyond academia, gaining massive community traction and investment.
The evolution of the project points toward a broader ambition. The goal is no longer just to monitor what is happening, but to use collective agentic intelligence to predict what will happen next. By proving that agents must argue to be accurate, BettaFish sets a new standard for automated research.