minimax_search: MiniMax-AI: The Three-Stage Engine for Agentic Web Intelligence

Moving beyond raw HTML with a Model Context Protocol server that parallelizes discovery, sanitizes content, and reasons over the results.

MiniMax-AI/minimax_search

A massive, chaotic waterfall of binary code and messy UI elements pouring into a refined clockwork funnel, emerging as a single perfectly inscribed scroll of parchment.
MiniMax Search transforms noisy web data into highly structured context.

Key Takeaways

The standard approach to agentic web search is fundamentally flawed. When a Large Language Model needs information, it typically triggers a single search query, scrapes the first few URLs, and ingests a massive wall of raw HTML. This linear process wastes tokens, dilutes the context window with navigation menus and ads, and often derails the model's reasoning process entirely.

The MiniMax Search repository solves this bottleneck by treating web research as a high-performance, multi-stage pipeline. Built as a Model Context Protocol (MCP) server, it replaces the naive scrape-and-pray method with a structured orchestration layer. It delegates discovery to Serper, delegates cleaning to Jina Reader, and uses the MiniMax-M2 model to synthesize the final answer.

The High-Fidelity Web Pipeline

Most search tools act as simple wrappers around an API. MiniMax Search acts as a conductor. The architecture is a discrete Search-Read-Understand loop.

First, the agent uses the search tool to find relevant URLs. Next, it uses the browse tool. Instead of fetching raw HTML, the server routes the target URLs through Jina Reader. This step converts chaotic web pages into clean, structured Markdown. Finally, the Markdown is passed directly into a MiniMax LLM, which is instructed to extract only the information relevant to the user's specific question.

The three-stage pipeline orchestrates search, scraping, and reasoning into a single agentic tool call.

Thinking in Parallel

The most powerful feature of the MiniMax Search server is its native multi-querying capability. Traditional agents suffer from linear thinking. They ask a question, wait for the result, and then ask a follow-up question. This sequential ping-pong introduces severe latency.

MiniMax Search accepts an array of queries simultaneously. The underlying Python logic utilizes thread pools to fan out these requests. An agent can explore the pros of a framework, the cons of a framework, and three alternative solutions in a single network round-trip. This parallel execution fundamentally changes how an agent gathers context.

A mechanical hand holding five different magnifying glasses, each focused on a different unique stamp on a single large map.
Parallel multi-querying allows agents to explore multiple hypotheses simultaneously.

Sanitizing the Stream

Feeding raw HTML to an LLM is like asking a human to read a book by looking at its printing press plates. The integration with Jina Reader is a critical architectural decision. By standardizing all web content into Markdown, the server ensures the context window is filled exclusively with signal.

FeatureTraditional Search ToolMiniMax Search MCP
Query ExecutionSequential (High Latency)Parallel Array (Low Latency)
Context FormatRaw HTML or DOM TextStructured Markdown via Jina
Reasoning StepNone (Client side)Integrated MiniMax-M2 Synthesis
IntegrationCustom API WrappersStandardized Model Context Protocol

Reasoning and the Protocol Bridge

The final stage of the pipeline relies heavily on the reasoning capabilities of the MiniMax models. The internal logic actively strips out the chain-of-thought blocks before returning the final synthesized answer to the client. This ensures the calling agent only receives the finalized, verified data.

Currently, turning off thinking is not supported. As we pay more attention to the final results, we do not focus much on non-thinking modes at this stage.

By wrapping this entire three-stage pipeline in the Model Context Protocol, MiniMax has created a universal, plug-and-play skill. Any MCP-compliant environment (from Claude Desktop to Cursor) can immediately leverage this advanced research engine over standard stdio communication. It shifts the burden of web navigation away from the primary agent and onto a specialized, high-performance microservice.