This recipe gives an agent a dataset and R, and lets it decide what to compute next from what it has found so far. It suits open-ended exploration. When the steps are known in advance, a fixed extraction pipeline is cheaper and easier to check.
The agent runs R code with your user account’s access to files and the network. That’s reasonable for your own data and prompts. For untrusted input, run the code in a sandbox instead; see Running R code.
Set up the agent
The agent gets the data-reading tools, run_r_code, and
the bundled data_analysis skill, which adds two summary
tools and some analysis guidance. The skill needs dplyr and ggplot2.
library(deputy)
agent <- Agent$new(
chat = ellmer::chat("anthropic/claude-sonnet-5"),
tools = c(tools_data(), list(tool_run_r_code)),
permissions = Permissions(r_code = TRUE, file_write = FALSE),
usage_limits = UsageLimits(max_requests = 15, max_cost_usd = 0.50),
system_prompt = "You are a data scientist. When you explore a dataset:
1. Start with its structure and summary statistics.
2. Check for missing values and data quality problems.
3. Look at the distributions of the key variables.
4. Look for relationships between variables.
5. Summarise what you found.
Use R code for every calculation, and say why you run each step."
)
agent$load_skill(system.file("skills", "data_analysis", package = "deputy"))The skill’s tools, eda_summary and
describe_column, give the model quick overviews, so it can
save run_r_code for the questions they don’t answer.
Run the analysis
result <- agent$run_sync(
"Explore the airquality dataset that comes with R. Tell me about data
quality problems, the distributions of the variables, and any patterns."
)
cat(result$response)airquality has missing values in Ozone and
Solar.R, and daily readings over five months, so a good
answer will mention both. Each run can take a different path, because
the model chooses its next step from the results so far.
The run stops after 15 model requests or $0.50, whichever comes first, and `result$stop_reason` says which.
Watch it work
run() returns a generator of events, so you can print
progress as the agent works:
events <- agent$run("Summarise the mtcars dataset.")
repeat {
event <- events()
if (coro::is_exhausted(event)) {
break
}
switch(event$type,
text = cat(event$text),
tool_start = cli::cli_alert_info("Calling {event$tool_name}"),
tool_end = cli::cli_alert_success("{event$tool_name} done"),
stop = cli::cli_alert("Stopped: {event$reason}")
)
}For a simpler log, add hook_log_tools() before calling
run_sync().
Check the result
result_is_success(result)
result$stop_reason
result_n_turns(result)
length(result_tool_calls(result))
result$duration
result$usageresult_tool_calls() also shows the code the model ran,
in each call’s tool_input, which is the quickest way to
check how it reached a number.
Next steps
- Tools: skills and custom tools.
- Running R code: keep R state between calls, or sandbox the code.
- Extraction pipeline: fixed steps with structured output.