Auto-claude-code-research-in-sleep: The End of the Agent Framework: Unpacking ARIS

How a folder of plain Markdown files and adversarial LLM psychology turned Claude Code into an autonomous ML researcher.

8 min read • View on GitHub • More from wanshuiyin

A massive, overly complex Rube Goldberg machine covered in tangled wires and heavy gears sits idle on the left. On the right, a simple, elegant Newton's cradle pendulum swings flawlessly in empty space.
ARIS discards complex Python state machines for a radically lightweight approach: treating plain Markdown files as executable orchestration logic.

Radically lightweight — zero dependencies, zero lock-in. The entire system is plain Markdown files.

wanshuiyin, Creator/Maintainer · ARIS GitHub Repository
Key Takeaways

The Markdown Microkernel

The era of bloated, dependency-heavy Python AI agent frameworks is facing a pragmatic rebellion. ARIS (Auto-Research-In-Sleep) proves that you can orchestrate complex workflows without a traditional framework. The secret weapon is not a massive new library, but treating Markdown as a Domain Specific Language (DSL).

Most agent frameworks require heavy Python dependencies to manage state, memory, and tool execution. ARIS discards this orchestration layer entirely. It treats SKILL.md files as executable code. This makes the system entirely agent-agnostic and readable by both humans and LLMs.

---
allowed-tools: [read_file, write_file, execute_command]
argument-hint: "Provide the path to the research hypothesis."
---

# Auto Review Loop

## State Logic
- MAX_ROUNDS = 4
- TARGET_SCORE = 7.0

## Execution Steps
1. Read the provided hypothesis document.
2. Call the `claude-review` MCP tool to get a critique.
3. If `score < TARGET_SCORE` and `current_round < MAX_ROUNDS`, revise the document and repeat.
Portrait of wanshuiyin, creator of ARIS.

Breaking the Mirror of Self-Play

The core theoretical innovation of ARIS is its approach to evaluation. A single model reviewing its own work inevitably falls into a "local minimum" of compounding errors. The model predicts what it expects to see, effectively grading its own homework with a built-in bias.

ARIS implements an adversarial review system. Claude executes the research, but an external model (like GPT-5.4 via MCP) acts as the ruthless reviewer. This simulates a Nash equilibrium for higher-quality scientific output. The friction between the two models prevents hallucination and forces rigorous refinement.

a single model self-reviewing is the stochastic case (predictable reward noise), while cross-model review is adversarial (the reviewer actively probes weaknesses the executor didn't anticipate).

A close-up of a chessboard where two mechanical chess pieces, one sleek and one brutalist, are locked in a tense standoff, cracking the board beneath them.
The cross-model adversarial review creates necessary friction, preventing the 'local minima' hallucinations inherent in self-grading.

The Adversarial Review Loop: Forcing cross-model feedback to escape the self-play trap.

Bridging the Protocols

While the core logic lives in Markdown, Python is still required to interact with the outside world. The /mcp-servers directory acts as the translation layer between the LLM's text output and external APIs.

The claude-review server wraps CLI commands into MCP calls, enabling the cross-model adversarial loop. The Feishu (Lark) bridge enables asynchronous human-in-the-loop approvals via long-polling. This allows the agent to pause execution, send an interactive card to a researcher's phone, and wait for a decision while the researcher sleeps.

Let Claude Code do research while you sleep. Wake up to find your paper scored, weaknesses identified, experiments run, and narrative rewritten — autonomously.

wanshuiyin, Creator/Maintainer · ARIS GitHub Repository

The Heavyweight Divide

ARIS represents a philosophical divide in agent design. Traditional approaches like AutoResearchClaw build complex, monolithic Python pipelines with 23-stage sandboxes and self-healing mechanisms. ARIS relies on the LLM's native ability to follow structured text.

FeatureARISAutoResearchClaw
Orchestration EnginePlaintext Markdown DSLMonolithic Python Pipeline
Review MechanismCross-Model Adversarial via MCPIntra-framework Multi-Agent Debate
ExtensibilityIDE-agnostic (Cursor, Trae, CLI)Requires full Python environment
Target UserPrompt hackers and researchersPython engineers

This is a debate between rigid engineering and organic, prompt-driven state machines. By pushing the complexity into the prompt and using MCP solely as a dumb pipe, ARIS proves that sometimes the best framework is no framework at all.