AutoResearch: let an agent run optimization experiments
Source:R/teleprompter-harness.R
AutoResearch.RdAutoResearch() hands the search to a research agent. In a loop, the agent
forms a hypothesis, may test ideas in sandboxed R code, proposes an edit to
the program, and sees how the edit scores; it keeps or reverts edits and
decides when to stop. dsprrr keeps control of budgets, evaluation,
checkpoints and the choice of the final program.
Usage
AutoResearch(
metric = NULL,
metric_threshold = NULL,
max_errors = 5L,
max_iterations = 20L,
patience = 6L,
target_score = NULL,
max_context_examples = 20L,
max_feedback_examples = 8L,
max_agent_steps = 4L,
sandbox = TRUE,
seed = NULL,
log_dir = NULL,
verbose = TRUE
)Arguments
- metric
A metric function (required) used to evaluate candidates.
- metric_threshold
Accepted for consistency with the other optimizers (see
Teleprompter()); not used.- max_errors
Integer; stop after this many consecutive failed evaluations when
compile()gets nocontrol(default5L).- max_iterations
Integer maximum number of evaluated experiments after the baseline (default
20L).- patience
Integer; stop after this many evaluated experiments without improvement (default
6L).- target_score
Optional score at which the search stops.
- max_context_examples
Integer maximum number of training rows shown to the agent (default
20L).- max_feedback_examples
Integer maximum number of failed rows returned after each evaluation (default
8L).- max_agent_steps
Integer maximum number of consecutive sandbox or invalid actions before the agent must submit a candidate (default
4L).- sandbox
If
TRUE(the default),compile()requires arunnerthat advertises an OS sandbox. IfFALSE, the agent cannot run code andrunneris ignored.- seed
Optional whole-number random seed.
- log_dir
Directory for a durable TrialLog, or
NULL(the default).- verbose
Whether to report progress (default
TRUE).
Value
An AutoResearch object to pass to compile().
Details
AutoResearch() is inspired by Andrej Karpathy's
autoresearch and the
AutoResearch engine in the
GEPA optimize-anything project. It is an
R implementation for dsprrr programs, not a port of either command-line
tool.
A candidate is a validated snapshot of every optimizable module in the
program, so one experiment can change instructions and templates across
several pipeline steps at once. The agent can branch from any earlier
candidate and sees per-example feedback. Candidates are always evaluated in
the host R process. Only the agent's exploratory R code goes to runner,
which must advertise an operating-system sandbox, such as
mcp_repl_runner(), unless sandbox = FALSE.
Compilation arguments
Besides the standard compile() arguments, this optimizer accepts
.agent_llm (the agent's Chat; defaults to .llm), runner (for sandboxed
analysis), control (an optimizer_control() object for budgets and
checkpoints) and objective (a text description of what to optimize for).
Other named arguments, such as .cache, are passed to candidate
evaluation.
Examples
research <- AutoResearch(
metric = metric_exact_match(field = "answer"),
max_iterations = 12L
)
research
#>
#> ── AutoResearch Teleprompter
#> Max experiments: 12
#> Patience: 6
#> OS sandbox required: TRUE
if (FALSE) { # \dontrun{
compiled <- compile(
program,
research,
trainset,
valset = valset,
.llm = ellmer::chat_openai(model = "gpt-6-luna"),
.agent_llm = ellmer::chat_anthropic(model = "claude-sonnet-4-5"),
runner = mcp_repl_runner(),
control = optimizer_control(
max_trials = 13L,
max_cost = 5,
checkpoint_path = file.path(tempdir(), "autoresearch.rds")
)
)
} # }