Skip to contents

AutoResearch() hands the search to a research agent. In a loop, the agent forms a hypothesis, may test ideas in sandboxed R code, proposes an edit to the program, and sees how the edit scores; it keeps or reverts edits and decides when to stop. dsprrr keeps control of budgets, evaluation, checkpoints and the choice of the final program.

Usage

AutoResearch(
  metric = NULL,
  metric_threshold = NULL,
  max_errors = 5L,
  max_iterations = 20L,
  patience = 6L,
  target_score = NULL,
  max_context_examples = 20L,
  max_feedback_examples = 8L,
  max_agent_steps = 4L,
  sandbox = TRUE,
  seed = NULL,
  log_dir = NULL,
  verbose = TRUE
)

Arguments

metric

A metric function (required) used to evaluate candidates.

metric_threshold

Accepted for consistency with the other optimizers (see Teleprompter()); not used.

max_errors

Integer; stop after this many consecutive failed evaluations when compile() gets no control (default 5L).

max_iterations

Integer maximum number of evaluated experiments after the baseline (default 20L).

patience

Integer; stop after this many evaluated experiments without improvement (default 6L).

target_score

Optional score at which the search stops.

max_context_examples

Integer maximum number of training rows shown to the agent (default 20L).

max_feedback_examples

Integer maximum number of failed rows returned after each evaluation (default 8L).

max_agent_steps

Integer maximum number of consecutive sandbox or invalid actions before the agent must submit a candidate (default 4L).

sandbox

If TRUE (the default), compile() requires a runner that advertises an OS sandbox. If FALSE, the agent cannot run code and runner is ignored.

seed

Optional whole-number random seed.

log_dir

Directory for a durable TrialLog, or NULL (the default).

verbose

Whether to report progress (default TRUE).

Value

An AutoResearch object to pass to compile().

Details

AutoResearch() is inspired by Andrej Karpathy's autoresearch and the AutoResearch engine in the GEPA optimize-anything project. It is an R implementation for dsprrr programs, not a port of either command-line tool.

A candidate is a validated snapshot of every optimizable module in the program, so one experiment can change instructions and templates across several pipeline steps at once. The agent can branch from any earlier candidate and sees per-example feedback. Candidates are always evaluated in the host R process. Only the agent's exploratory R code goes to runner, which must advertise an operating-system sandbox, such as mcp_repl_runner(), unless sandbox = FALSE.

Compilation arguments

Besides the standard compile() arguments, this optimizer accepts .agent_llm (the agent's Chat; defaults to .llm), runner (for sandboxed analysis), control (an optimizer_control() object for budgets and checkpoints) and objective (a text description of what to optimize for). Other named arguments, such as .cache, are passed to candidate evaluation.

Examples

research <- AutoResearch(
  metric = metric_exact_match(field = "answer"),
  max_iterations = 12L
)
research
#> 
#> ── AutoResearch Teleprompter 
#> Max experiments: 12
#> Patience: 6
#> OS sandbox required: TRUE

if (FALSE) { # \dontrun{
compiled <- compile(
  program,
  research,
  trainset,
  valset = valset,
  .llm = ellmer::chat_openai(model = "gpt-6-luna"),
  .agent_llm = ellmer::chat_anthropic(model = "claude-sonnet-4-5"),
  runner = mcp_repl_runner(),
  control = optimizer_control(
    max_trials = 13L,
    max_cost = 5,
    checkpoint_path = file.path(tempdir(), "autoresearch.rds")
  )
)
} # }