Skip to contents

Deputy leaves provider requests to ellmer: connections, retries, timeouts and model parameters. It adds the things an agent needs on top of that: a second provider to try when the first is down, a record of every request, and OpenTelemetry spans that tie model calls to the run that made them.

Fall back to another provider

fallback_chats lists chats to try, in order, when a request fails before the model has sent anything back:

agent <- Agent$new(
  ellmer::chat("openai/gpt-6-luna"),
  fallback_chats = list(ellmer::chat("anthropic/claude-sonnet-5")),
  permissions = permissions_readonly(),
  usage_limits = UsageLimits(max_requests = 5)
)
result <- agent$run_sync("Summarise the supplied evidence.")

Deputy switches for connection errors and HTTP 408, 429, 500, 502, 503 and 504, after ellmer’s own retries. Authentication errors, errors in your callbacks and validation failures stop the run instead. So does a failure after the model has started streaming text or asked for a tool: switching then could repeat the tool’s effects, so Deputy keeps the partial result and reports the error.

Each fallback chat is copied, and the copy gets the agent’s conversation, system prompt and tools; its model, credentials and settings stay as you configured them. Pass chats with no history or tools. Once Deputy switches, it stays on the new chat for later runs, moving further down the list if that one fails too. It never picks a provider or model you didn’t list, and it doesn’t adapt content that a fallback provider can’t accept.

Automatic compaction has its own fallback list, ContextPolicy(summary_fallback_chats = ); see Conversations and context. A lead agent’s fallbacks apply to the lead only; subagents use the provider the lead is currently on.

Retries and timeouts

ellmer retries failed HTTP requests itself, up to getOption("ellmer_max_tries") attempts in total (3 by default), and gives up on a request after getOption("ellmer_timeout_s") seconds (300 by default). Set these options once when your application starts. Deputy doesn’t add a retry loop of its own, so three HTTP attempts count as one request toward max_requests, and a failed request counts too.

The asynchronous structured-output path, chat_structured_async(), goes through httr2’s req_perform_promise(), which doesn’t retry.

Each model request adds "request_start" and "request_end" events to the run, or "request_error" with the original condition when it fails, and a switch adds a "fallback" event. Find them in result$events. They may contain provider error details, so treat them as private.

Tracing with OpenTelemetry

ellmer already records OpenTelemetry spans for model calls, HTTP requests and tool execution. Deputy adds a deputy.run span around each run, with attributes such as deputy.run.id and deputy.agent.id, and events for permission decisions, hooks, fallbacks, compaction, checkpoints, delegation and the stop. ellmer’s gen_ai.conversation.id is set to the agent’s session ID, and spans from subagents stay connected to the run that delegated to them.

Configure an OpenTelemetry exporter before loading ellmer and Deputy. To try it without an exporter, record spans in memory with otelsdk:

# In a fresh R session:
record <- otelsdk::with_otel_record({
  library(deputy)
  agent <- Agent$new(ellmer::chat("openai/gpt-6-luna"))
  agent$run_sync("Summarise this text in one sentence: ...")
})
record$value$run_id
record$traces

Deputy’s spans carry identifiers and decisions only. They never include prompts, tool arguments or results, validation feedback, error messages, file paths or your run_context, even when ellmer’s content capture is on. That capture is controlled by ellmer: set OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true before loading ellmer to include messages in ellmer’s spans. ellmer’s spans can also include endpoint URLs and tool descriptions, so apply your exporter’s privacy settings.

Evaluate an agent

An evaluation runs a fixed set of cases, scores each answer, and records what each run cost. AgentResult has what you need to record: stop_reason, usage, duration, the run ID and the events. Put a case ID in run_context to connect a result, and its trace, back to its case:

result <- agent$run_sync(
  case$prompt,
  run_context = list(evaluation = list(case_id = case$id))
)

The bundled 10-evaluation.R script is a small complete example: it runs a two-case dataset, scores the answers with an exact-match rule, and returns a data frame with the case, run and session IDs, status, requests, tokens, cost, duration and any error class:

source(system.file("examples", "standalone", "10-evaluation.R", package = "deputy"))
evaluation

Deputy doesn’t have its own dataset format or scoring framework; replace the cases and the scoring rule with your own.