A recursive language model (RLM) (Zhang et al.
2025) answers questions about inputs it never sees in full.
dsprrr keeps the inputs in an R session. The model sees their names,
sizes and short previews, writes R code to inspect them, reads the
printed result, and repeats until it submits typed outputs. This page
explains that loop, how to choose where the generated code runs, and the
limits that apply. rlm_module() is experimental.
When to use RLM
| Your task | Use |
|---|---|
| Short inputs and a direct answer | module() |
| Reasoning over context you have already selected | chain_of_thought() |
| A calculation the model writes as R code | program_of_thought() |
| Calling tools or acting on outside systems |
react() or code_act()
|
| Answering from a prepared retrieval index | rag_module() |
| Exploring large or irregular R objects when you don’t know where the answer is | rlm_module() |
| Searching for the best program structure over labeled examples | flex() |
Every RLM step costs one model call and one code execution. If an R function, a query or an ordinary module already describes the path to the answer, use that instead.
Choose where the code runs
RLM executes code that a model wrote, so pick the runner first. It decides what that code can touch and how large the inputs can be.
mcp_repl_runner() |
r_code_runner(persistent = TRUE) |
|
|---|---|---|
| Isolation | Operating-system sandbox from Posit’s mcp-repl: network off, writes only to the workspace and temporary files | A separate R process with your user’s file, network and environment access |
| Use it for | Inputs you do not control, or code that must not reach your files and network | Inputs and code you trust |
| Inputs | Serialized and sent with every step, so they must be small (see limits) | Any object callr can serialize, staged once per call |
| Requires | The mcptools package and the mcp-repl executable |
The callr package |
Install the sandbox with:
Rscript -e 'install.packages("mcptools")'
uv tool install posit-mcp-repl # or: pipx install posit-mcp-replPass the runner as a zero-argument interpreter_factory.
dsprrr then creates a fresh runner for each call and shuts it down
afterwards, even when the call fails. Creating the module does not start
a runner:
library(dsprrr)
# Sandboxed: a short policy text you did not write
reviewer <- rlm_module(
"policy, question -> answer",
interpreter_factory = function() mcp_repl_runner(timeout = 30)
)
# Trusted: large data frames from your own pipeline
analyst <- rlm_module(
"sessions, changes, question -> answer, evidence",
interpreter_factory = function() {
r_code_runner(timeout = 30, persistent = TRUE)
}
)
analyst
#>
#> ── RLMModule
#> • Signature: sessions, changes, question -> answer, evidence
#> • Max iterations: 20
#> • Max LLM calls: 50
#> • Runner: fresh runner per invocation
#> • Recursive queries: "outer LM"
#> • Custom tools: 0A runner passed as runner = instead belongs to you.
dsprrr reuses it across calls, never shuts it down, and after each call
removes only the variables and inputs it staged; files or options the
generated code changed stay changed. Use such a runner for one call at a
time, within one trust boundary.
For a one-off question, rlm() builds the module and runs
it in one call, with a fresh mcp_repl_runner() unless you
pass .runner or .interpreter_factory.
One call, step by step
- dsprrr loads each input into the runner as
.context$<name>. Foranalystthese are.context$sessions,.context$changesand.context$question. - The first prompt lists each input’s class, size, declared type and a
preview of at most 1,000 characters (for a data frame, its
str()), the output fields, and the helper functions below. - The model returns its reasoning and one block of R code, such as
str(.context$sessions). The runner executes it, and variables persist to the next step. - The printed result, cut to
max_output_chars, is added to the next prompt together with every earlier step’s code and output. - The loop ends at the first valid
SUBMIT(). Ifmax_iterations(default- runs out first, dsprrr warns and a separate
extractpredictor produces the outputs from the steps so far.
- runs out first, dsprrr warns and a separate
| Helper | What it does |
|---|---|
peek(x, start, end) |
Characters start to end of a string, or
elements of a vector |
search(x, pattern) |
All matches of a regular expression in x
|
llm_query(query, context_slice) |
Asks the sub-model one question about a string and returns the answer |
llm_query_batched(queries, slices) |
Asks several independent questions; answers come back in input order |
SUBMIT(...) |
Ends the loop with the output fields |
Investigate a regression with
RLM shows a complete trajectory for a module like
analyst.
Recursive queries
llm_query() reads like a function call, but the model
call happens in your R session, not in the runner:
notes <- .context$changes$note[.context$changes$release == "2.4.0"]
suspects <- llm_query(
"Which of these change notes could lower mobile checkout conversion?",
paste(notes, collapse = "\n")
)The step stops at the call, dsprrr sends the question to the
sub-model, then reruns the step from the top with the answer filled in.
Assignments made in the step are kept once the rerun completes, and
tools run once. Anything with outside effects before the query (writing
a file, calling an API) runs again on every rerun, so keep that part
read-only. dsprrr does not pass model credentials to the runner, but
that alone does not hide them: an r_code_runner() process
inherits your environment variables.
sub_lm sets the chat for these queries, for example
sub_lm = ellmer::chat_openai(model = "gpt-6-luna") in
rlm_module(). The default, NULL, reuses the
chat that runs the main loop. Every question counts against
max_llm_calls (default 50), including each one in a batch,
and max_llm_calls = 0 removes both query helpers. In
llm_query_batched(), a provider error for one question
comes back as an "[ERROR] ..." string in its position; any
other failure ends the call.
Typed submission
The signature decides what SUBMIT() accepts. For
"logs -> error_count: int, severity: enum('low', 'high'), evidence: list[str]",
generated code submits named values:
SUBMIT(
error_count = 7L,
severity = "high",
evidence = c("AUTH-401 rate tripled at 14:02", "Token refresh retries fell to 0")
)A missing required field, an unknown field, or a value of the wrong
type does not end the loop. The error is added to the trajectory and the
model can fix its submission on the next step. Optional fields may be
left out. Outputs must use ellmer’s string, number, integer, boolean,
enum, array and object types; a type_from_schema() output
is rejected when the module is created because RLM cannot check it.
Read the result
Ask for the structured result to keep the trajectory with the answer:
result <- run(
analyst,
sessions = sessions,
changes = changes,
question = "Which cohort's conversion fell after release 2.4.0, and why?",
.llm = ellmer::chat_openai(model = "gpt-6-luna"),
.return_format = "structured"
)
result$output$answer
result$metadata$output_source # "submit", or "fallback" if no step submitted
result$metadata$repl_history # one entry per stepEach repl_history entry holds the step’s reasoning, its
code, the output the model saw, and whether it succeeded or submitted.
The metadata also counts steps (iterations), recursive
queries (llm_calls), provider calls, tokens and cost;
totals are NA when a provider cannot report every call.
Treat output_source == "fallback" as a warning that the
step budget ran out or the instructions need work.
run() stages each input as one object, whatever its
length. For many investigations, use run_dataset() with one
row per investigation and list-columns for data frames; with a factory,
each row gets its own runner.
Errors in generated code, such as a missing column, are shown to the model so it can try again. Failures of the runner itself (it cannot start, its connection breaks) end the call with an error.
Limits
| Limit | Size | Applies to | When exceeded |
|---|---|---|---|
| Output the model sees per step |
max_output_chars, default 10,000 characters |
RLM, any runner | The middle is cut; the model sees the head, the tail and how much was omitted |
| Preview of each input in the first prompt | 1,000 characters | RLM, any runner | Cut the same way |
One request to mcp-repl: the step’s code plus its serialized
.context
|
7,000 bytes, gzip-compressed when larger | RLM steps, executable Flex steps and harness sandbox actions in
mcp_repl_runner()
|
Rejected before it runs; in RLM the model sees the error |
One message back from the runner: an RLM SUBMIT(),
llm_query() or tool request, or a Flex predictor or tool
call |
3,000 bytes, encoded | RLM and executable Flex in mcp_repl_runner()
|
The step fails; in RLM the model sees the error and can retry with smaller values |
| Tool calls in one step | 1,000 | RLM, any runner | The call ends with an error |
| Sandbox code from an agent | 12,000 characters |
AutoResearch() and MetaHarness()
|
The action is rejected without running |
| Sandbox result shown to the agent | 4,000 characters |
AutoResearch() and MetaHarness()
|
Truncated and marked [truncated]
|
The request limit is the one you will meet. In RLM,
.context holds every input, and RLM’s own helper code
already takes about 4,700 of the 7,000 bytes after compression (about
5,200 when llm_query() is enabled). That leaves roughly
2,000 bytes for the compressed inputs: a few thousand characters of
ordinary text or a couple of hundred rows of a small data frame, and
much less for text that barely compresses, such as IDs or hashes. Inputs
are sent again with every step, so an input that is too large fails
every step and the call ends in fallback extraction.
r_code_runner(persistent = TRUE) has neither the request
nor the frame limit; use it for larger inputs you trust.
mcp-repl also moves very large printed output into a file preview. RLM treats that as a failed step rather than reading a partial result, so generated code should print summaries, not whole objects.
Privacy and caching
The runner boundary does not keep data away from the model provider.
The first prompt carries each input’s preview: a string of up to 1,000
characters in full, and the start and end of a longer one. Anything
printed, passed to llm_query(), or submitted also reaches
the provider. Tools given through tools = run in your R
session with your permissions, outside any sandbox.
The main-loop and fallback calls use dsprrr’s response cache, which
can include a disk cache. Pass .cache = FALSE to
run() for sensitive work. Recursive queries and code
execution are never cached.
Relationship to DSPy
dsprrr follows the contract of DSPy’s dspy.RLM, with R
instead of Python in the interpreter and ellmer chats for the models.
The two predictors inside an RLM, generate_action (writes
each step) and extract (the fallback), are ordinary modules
that GEPA() can tune. dsprrr
for DSPy users lists the differences.
For a worked example on a fixed dataset, with an independent check of the answer, see Investigate a regression with RLM.