predict-rlm: The Agent Framework That Lets the Model Write Its Own Runtime

Inside Trampoline AI’s recursive language model runtime, the LLM stops being a tool caller and starts acting like a program that can recurse, parallelize, and keep state without drowning in context.

9 min read • View on GitHub • More from Trampoline-AI

A model sits inside a sealed glass sandbox at a drafting table, writing Python that branches into a widening execution graph. On one side of the glass is a cluttered pile of chat history, and on the other is a clean workspace of variables, tools, files, and sub-calls. The image explains the repo’s core inversion: the model is not chatting through a fixed loop, it is authoring the loop itself.
The runtime separates token space from variable space, so the model can keep working state outside the chat log.
Key Takeaways

The model is not just answering. It is programming

predict-rlm flips the usual agent pattern. The model gets a sandboxed Python REPL, then writes the loops, branches, and sub-calls it needs to finish the job. That means the interesting unit is no longer a single completion. It is a little program that can inspect state, recurse, and call other models from inside its own runtime.

Production focused port of RLMs that allows the LM to call its sub-lm with DSPy signatures. Define your inputs, outputs, and tools — the model handles its own control flow. Get fully interpretable trajectories and performance that scales directly with model improvements. Without context rot.

Trampoline AI, Project Developer · predict-rlm

That inversion matters because it changes where reasoning lives. Traditional agents orchestrate from the outside. Here, the orchestration is authored from the inside, in Python variables that can persist across steps and be inspected directly.

Why recursive reasoning beats a long prompt

The repo is trying to solve a simple failure mode with a blunt name: context rot. As the chat grows, the model gets more text but less reliable access to the parts that matter. `predict-rlm` sidesteps that by using the REPL as working memory. State lives in variables, not in an ever-expanding prompt.

That makes the system more like a scratchpad with execution privileges than a conversation. The model can chunk a problem, store intermediate results, and come back to them later without paying the tax of rewriting the whole situation back into tokens.

A split editorial scene contrasts a tall tower of tangled chat bubbles with a compact ledger of Python variables and nested notes. The left side looks crowded and fragile, while the right side looks organized and reusable. The image explains why recursive state in a REPL is more reliable than growing context history.
The runtime uses variable space as working memory, which is a better fit for long, messy tasks than a bloated chat log.

The runtime trick: Pyodide, Deno, and JSPI

This is the technical hinge. The sandbox is built on Deno and Pyodide, which means the model’s Python runs in a WASM environment instead of on the host. That gives isolation, but isolation alone would still make async tool use awkward. JSPI is the bridge that lets Python code look synchronous while actually waiting on JavaScript promises underneath.

JSPI is the thing that makes async sub-calls usable inside a sandboxed Python loop without forcing a chat roundtrip.

async def solve(tasks):
    results = await asyncio.gather(*[call_sub_lm(task) for task in tasks])
    shared_state["results"] = results
    return summarize(shared_state)

The point is not just safety. It is that the runtime can support fan-out. Once the sandbox can await JavaScript promises, recursive Python code can launch multiple sub-calls, collect them, and continue with the same state intact.

How PredictRLM turns schemas into behavior

`PredictRLM` is the orchestration layer that turns user intent into typed behavior. It dynamically builds DSPy signatures for action and extraction, then reconstructs nested Pydantic models from schema. In plain English, it gives the model a typed shape for what to do and what to return.

That matters because freeform agent output is expensive to clean up. Here, the runtime already knows whether it is dealing with a scalar, a list of files, or a nested object. The model can still explore, but the edges are structured.

Traditional agentpredict-rlmWhy it matters
Human-coded orchestration loopModel-written recursion in a REPLThe model can adapt its own control flow to the task.
State lives in chat historyState lives in Python variablesWorking memory becomes persistent and inspectable.
Mostly serial tool callsAsync sub-calls can fan out in parallelThe runtime can scale exploration without linear slowdown.
Freeform outputs need cleanupDSPy signatures and Pydantic schemasTyped results are easier to trust and integrate.
Failure often means lost contextFailure is contained in a sandboxThe host stays isolated from messy execution.
Best for simple step-by-step tasksBest for long, structured, recursive tasksThe control model matches the problem shape.

Skills, files, and the sandbox as a working environment

The repo gets practical fast in `rlm_skills.py` and `files.py`. Skills bundle instructions, dependencies, and tools into reusable units. Files are treated as first-class inputs and outputs, so the sandbox can mount input artifacts and sync results back out when the job is done.

That makes the runtime useful for document-heavy workflows. The model does not have to pretend a PDF is just more text. It can inspect, chunk, transform, and emit files as part of the same loop.

skills = merge_skills(base_skill, pdf_skill, pii_skill)
inputs = mount_files([invoice_pdf, supporting_docs])
result = await run_rlm(goal, skills=skills, files=inputs)

The design is small but consequential. Once files, dependencies, and tools are packaged cleanly, the sandbox stops being a demo and starts behaving like a workbench.

What it replaces, and what it does not

`predict-rlm` is not a claim that every agent framework is obsolete. It is a claim that some tasks deserve a different control model. ReAct loops, Claude Code-style orchestration, and broader frameworks like LangChain or CrewAI all optimize for a human-authored flow. This repo optimizes for a model-authored one.

PatternStrengthLimitation
ReAct-style loopSimple and familiarThe loop is still fixed by the human designer.
Claude Code-style orchestrationUseful for code-centric tasksResults often come back as text blobs that re-enter context.
LangChain or CrewAIBroad integration surfaceControl flow usually remains externally managed.
predict-rlmRecursive, typed, sandboxed executionBest when the problem is long, structured, and stateful.

The real comparison is not about features. It is about where intelligence is allowed to live. The old model says humans author the orchestration and the model fills in the steps. `predict-rlm` says the model should author more of the steps itself, inside a safer runtime.

The bigger bet: a smaller harness, a smarter model

This repo is a bet on the Bitter Lesson. Instead of building thicker orchestration, it gives the model a better sandbox and better primitives. If the base model gets better at code, the runtime gets better too.

That is why `predict-rlm` feels less like a wrapper and more like a compact distributed system for language models. It is not trying to outsmart the model. It is trying to remove the parts of the agent stack that get in the model’s way.