Skip to contents

This recipe gives an agent a dataset and R, and lets it decide what to compute next from what it has found so far. It suits open-ended exploration. When the steps are known in advance, a fixed extraction pipeline is cheaper and easier to check.

The agent runs R code with your user account’s access to files and the network. That’s reasonable for your own data and prompts. For untrusted input, run the code in a sandbox instead; see Running R code.

Set up the agent

The agent gets the data-reading tools, run_r_code, and the bundled data_analysis skill, which adds two summary tools and some analysis guidance. The skill needs dplyr and ggplot2.

library(deputy)

agent <- Agent$new(
  chat = ellmer::chat("anthropic/claude-sonnet-5"),
  tools = c(tools_data(), list(tool_run_r_code)),
  permissions = Permissions(r_code = TRUE, file_write = FALSE),
  usage_limits = UsageLimits(max_requests = 15, max_cost_usd = 0.50),
  system_prompt = "You are a data scientist. When you explore a dataset:
    1. Start with its structure and summary statistics.
    2. Check for missing values and data quality problems.
    3. Look at the distributions of the key variables.
    4. Look for relationships between variables.
    5. Summarise what you found.
    Use R code for every calculation, and say why you run each step."
)

agent$load_skill(system.file("skills", "data_analysis", package = "deputy"))

The skill’s tools, eda_summary and describe_column, give the model quick overviews, so it can save run_r_code for the questions they don’t answer.

Run the analysis

result <- agent$run_sync(
  "Explore the airquality dataset that comes with R. Tell me about data
   quality problems, the distributions of the variables, and any patterns."
)

cat(result$response)

airquality has missing values in Ozone and Solar.R, and daily readings over five months, so a good answer will mention both. Each run can take a different path, because the model chooses its next step from the results so far.

The run stops after 15 model requests or $0.50, whichever comes first, and `result$stop_reason` says which.

Watch it work

run() returns a generator of events, so you can print progress as the agent works:

events <- agent$run("Summarise the mtcars dataset.")

repeat {
  event <- events()
  if (coro::is_exhausted(event)) {
    break
  }
  switch(event$type,
    text = cat(event$text),
    tool_start = cli::cli_alert_info("Calling {event$tool_name}"),
    tool_end = cli::cli_alert_success("{event$tool_name} done"),
    stop = cli::cli_alert("Stopped: {event$reason}")
  )
}

For a simpler log, add hook_log_tools() before calling run_sync().

Check the result

result_is_success(result)
result$stop_reason

result_n_turns(result)
length(result_tool_calls(result))

result$duration
result$usage

result_tool_calls() also shows the code the model ran, in each call’s tool_input, which is the quickest way to check how it reached a number.

Next steps