Auto-claude-code-research-in-sleep: The End of the Agent Framework: Unpacking ARIS
How a folder of plain Markdown files and adversarial LLM psychology turned Claude Code into an autonomous ML researcher.

Radically lightweight — zero dependencies, zero lock-in. The entire system is plain Markdown files.
- ARIS orchestrates complex, multi-agent scientific research using plain Markdown files instead of heavy Python dependencies.
- The system prevents hallucination by enforcing adversarial cross-model reviews, forcing Claude and GPT to critique each other.
- Python is relegated to the edges, functioning only as Model Context Protocol (MCP) servers bridging the Markdown logic to external APIs.
The Markdown Microkernel
The era of bloated, dependency-heavy Python AI agent frameworks is facing a pragmatic rebellion. ARIS (Auto-Research-In-Sleep) proves that you can orchestrate complex workflows without a traditional framework. The secret weapon is not a massive new library, but treating Markdown as a Domain Specific Language (DSL).
Most agent frameworks require heavy Python dependencies to manage state, memory, and tool execution. ARIS discards this orchestration layer entirely. It treats SKILL.md files as executable code. This makes the system entirely agent-agnostic and readable by both humans and LLMs.
---
allowed-tools: [read_file, write_file, execute_command]
argument-hint: "Provide the path to the research hypothesis."
---
# Auto Review Loop
## State Logic
- MAX_ROUNDS = 4
- TARGET_SCORE = 7.0
## Execution Steps
1. Read the provided hypothesis document.
2. Call the `claude-review` MCP tool to get a critique.
3. If `score < TARGET_SCORE` and `current_round < MAX_ROUNDS`, revise the document and repeat.
Breaking the Mirror of Self-Play
The core theoretical innovation of ARIS is its approach to evaluation. A single model reviewing its own work inevitably falls into a "local minimum" of compounding errors. The model predicts what it expects to see, effectively grading its own homework with a built-in bias.
ARIS implements an adversarial review system. Claude executes the research, but an external model (like GPT-5.4 via MCP) acts as the ruthless reviewer. This simulates a Nash equilibrium for higher-quality scientific output. The friction between the two models prevents hallucination and forces rigorous refinement.
a single model self-reviewing is the stochastic case (predictable reward noise), while cross-model review is adversarial (the reviewer actively probes weaknesses the executor didn't anticipate).
Bridging the Protocols
While the core logic lives in Markdown, Python is still required to interact with the outside world. The /mcp-servers directory acts as the translation layer between the LLM's text output and external APIs.
The claude-review server wraps CLI commands into MCP calls, enabling the cross-model adversarial loop. The Feishu (Lark) bridge enables asynchronous human-in-the-loop approvals via long-polling. This allows the agent to pause execution, send an interactive card to a researcher's phone, and wait for a decision while the researcher sleeps.
Let Claude Code do research while you sleep. Wake up to find your paper scored, weaknesses identified, experiments run, and narrative rewritten — autonomously.
The Heavyweight Divide
ARIS represents a philosophical divide in agent design. Traditional approaches like AutoResearchClaw build complex, monolithic Python pipelines with 23-stage sandboxes and self-healing mechanisms. ARIS relies on the LLM's native ability to follow structured text.
| Feature | ARIS | AutoResearchClaw |
|---|---|---|
| Orchestration Engine | Plaintext Markdown DSL | Monolithic Python Pipeline |
| Review Mechanism | Cross-Model Adversarial via MCP | Intra-framework Multi-Agent Debate |
| Extensibility | IDE-agnostic (Cursor, Trae, CLI) | Requires full Python environment |
| Target User | Prompt hackers and researchers | Python engineers |
This is a debate between rigid engineering and organic, prompt-driven state machines. By pushing the complexity into the prompt and using MCP solely as a dumb pipe, ARIS proves that sometimes the best framework is no framework at all.