predict-rlm: The Agent Framework That Lets the Model Write Its Own Runtime
Inside Trampoline AI’s recursive language model runtime, the LLM stops being a tool caller and starts acting like a program that can recurse, parallelize, and keep state without drowning in context.
- predict-rlm treats the model as a program author inside a sandboxed Python runtime, not as a chat agent glued to a human-written loop.
- Its main trick is to keep task state in Python variables, which makes long-running work less fragile than depending on token history alone.
- JSPI and Pyodide turn the sandbox into a place where async tool calls can fan out in parallel without breaking the control flow.
- The repo is best understood as a smaller harness with a smarter model, not as yet another layer of agent orchestration.
The model is not just answering. It is programming
predict-rlm flips the usual agent pattern. The model gets a sandboxed Python REPL, then writes the loops, branches, and sub-calls it needs to finish the job. That means the interesting unit is no longer a single completion. It is a little program that can inspect state, recurse, and call other models from inside its own runtime.
Production focused port of RLMs that allows the LM to call its sub-lm with DSPy signatures. Define your inputs, outputs, and tools — the model handles its own control flow. Get fully interpretable trajectories and performance that scales directly with model improvements. Without context rot.
That inversion matters because it changes where reasoning lives. Traditional agents orchestrate from the outside. Here, the orchestration is authored from the inside, in Python variables that can persist across steps and be inspected directly.
Why recursive reasoning beats a long prompt
The repo is trying to solve a simple failure mode with a blunt name: context rot. As the chat grows, the model gets more text but less reliable access to the parts that matter. `predict-rlm` sidesteps that by using the REPL as working memory. State lives in variables, not in an ever-expanding prompt.
That makes the system more like a scratchpad with execution privileges than a conversation. The model can chunk a problem, store intermediate results, and come back to them later without paying the tax of rewriting the whole situation back into tokens.
The runtime trick: Pyodide, Deno, and JSPI
This is the technical hinge. The sandbox is built on Deno and Pyodide, which means the model’s Python runs in a WASM environment instead of on the host. That gives isolation, but isolation alone would still make async tool use awkward. JSPI is the bridge that lets Python code look synchronous while actually waiting on JavaScript promises underneath.
async def solve(tasks):
results = await asyncio.gather(*[call_sub_lm(task) for task in tasks])
shared_state["results"] = results
return summarize(shared_state)
The point is not just safety. It is that the runtime can support fan-out. Once the sandbox can await JavaScript promises, recursive Python code can launch multiple sub-calls, collect them, and continue with the same state intact.
How PredictRLM turns schemas into behavior
`PredictRLM` is the orchestration layer that turns user intent into typed behavior. It dynamically builds DSPy signatures for action and extraction, then reconstructs nested Pydantic models from schema. In plain English, it gives the model a typed shape for what to do and what to return.
That matters because freeform agent output is expensive to clean up. Here, the runtime already knows whether it is dealing with a scalar, a list of files, or a nested object. The model can still explore, but the edges are structured.
| Traditional agent | predict-rlm | Why it matters |
|---|---|---|
| Human-coded orchestration loop | Model-written recursion in a REPL | The model can adapt its own control flow to the task. |
| State lives in chat history | State lives in Python variables | Working memory becomes persistent and inspectable. |
| Mostly serial tool calls | Async sub-calls can fan out in parallel | The runtime can scale exploration without linear slowdown. |
| Freeform outputs need cleanup | DSPy signatures and Pydantic schemas | Typed results are easier to trust and integrate. |
| Failure often means lost context | Failure is contained in a sandbox | The host stays isolated from messy execution. |
| Best for simple step-by-step tasks | Best for long, structured, recursive tasks | The control model matches the problem shape. |
Skills, files, and the sandbox as a working environment
The repo gets practical fast in `rlm_skills.py` and `files.py`. Skills bundle instructions, dependencies, and tools into reusable units. Files are treated as first-class inputs and outputs, so the sandbox can mount input artifacts and sync results back out when the job is done.
That makes the runtime useful for document-heavy workflows. The model does not have to pretend a PDF is just more text. It can inspect, chunk, transform, and emit files as part of the same loop.
skills = merge_skills(base_skill, pdf_skill, pii_skill)
inputs = mount_files([invoice_pdf, supporting_docs])
result = await run_rlm(goal, skills=skills, files=inputs)
The design is small but consequential. Once files, dependencies, and tools are packaged cleanly, the sandbox stops being a demo and starts behaving like a workbench.
What it replaces, and what it does not
`predict-rlm` is not a claim that every agent framework is obsolete. It is a claim that some tasks deserve a different control model. ReAct loops, Claude Code-style orchestration, and broader frameworks like LangChain or CrewAI all optimize for a human-authored flow. This repo optimizes for a model-authored one.
| Pattern | Strength | Limitation |
|---|---|---|
| ReAct-style loop | Simple and familiar | The loop is still fixed by the human designer. |
| Claude Code-style orchestration | Useful for code-centric tasks | Results often come back as text blobs that re-enter context. |
| LangChain or CrewAI | Broad integration surface | Control flow usually remains externally managed. |
| predict-rlm | Recursive, typed, sandboxed execution | Best when the problem is long, structured, and stateful. |
The real comparison is not about features. It is about where intelligence is allowed to live. The old model says humans author the orchestration and the model fills in the steps. `predict-rlm` says the model should author more of the steps itself, inside a safer runtime.
The bigger bet: a smaller harness, a smarter model
This repo is a bet on the Bitter Lesson. Instead of building thicker orchestration, it gives the model a better sandbox and better primitives. If the base model gets better at code, the runtime gets better too.
That is why `predict-rlm` feels less like a wrapper and more like a compact distributed system for language models. It is not trying to outsmart the model. It is trying to remove the parts of the agent stack that get in the model’s way.