Skip to contents

A subagent starts every task with an empty conversation. Sometimes you want the opposite: a specialist that remembers what it did last time, so a follow-up question can build on its earlier work. Deputy calls these retained specialists. Any agent can own them, and you can let the model delegate to them, follow up with them yourself, or wire several into a graph where specialists delegate to each other.

Retain an agent

retain_agent() hands a fully configured agent to an owner agent. The owner then runs it, keeping its conversation between tasks:

library(deputy)

owner <- Agent$new(ellmer::chat("openai/gpt-6-luna"))

analyst <- Agent$new(
  chat = ellmer::chat("anthropic/claude-sonnet-5"),
  tools = tools_data(),
  permissions = permissions_readonly(),
  system_prompt = "You analyse data files and explain your reasoning."
)

handle <- owner$retain_agent(
  analyst,
  usage_limits = UsageLimits(max_requests = 20)
)

first <- owner$continue_agent(
  handle,
  "Summarise sales.csv.",
  UsageLimits(max_requests = 4)
)
second <- owner$continue_agent(
  handle,
  "Which region drove the change you found?",
  UsageLimits(max_requests = 4)
)

Each call to continue_agent() sends a new task to the same conversation, with its own allocation. The usage_limits given to retain_agent() cap the total across all calls, and max_runs (32 by default) caps how many there can be. continue_agent_async() returns a promise instead.

While retained, the specialist belongs to its owner. You can’t run it directly or change its prompt, model or tools until you release it, and it runs one task at a time: a second call while one is running fails straight away. owner$cancel_agent(handle) stops the current task and keeps the history so far; owner$release_agent(handle) ends the arrangement. The specialist must be a plain Agent, without an approval directory or fallback chats.

The specialist’s tools and any connections they hold remain yours to manage. Handles and retained conversations live in memory and don’t survive an R restart. To keep a record, export it with owner$export_subagents() (see Briefing and inspecting subagents) before releasing.

Let the model delegate to a specialist

delegation_tool() turns a handle into a tool for the owner’s model. The model supplies only the task; the specialist, its budget and everything else are fixed by your code:

owner$register_tool(delegation_tool(
  owner,
  handle,
  name = "ask_analyst",
  description = "Ask the data analyst a question about the sales files.",
  usage_limits = UsageLimits(max_requests = 4)
))

owner$run_sync("Find out what changed in sales this quarter and why.")

Every call continues the same conversation, whether the model or your code makes it. The tool returns the same compact outcome as ordinary delegation.

The specialist keeps its own permissions, and each of its tool calls must also pass the owner’s current permissions. An owner in read-only or plan mode can still use the tool, and the specialist is then held to that mode too.

Start from an ellmer chat

If you already have configured ellmer chats, perhaps with their own tools and history, adopt_chat() retains a copy of one without building an Agent by hand:

handle <- adopt_chat(
  analyst_chat,
  owner,
  permissions = Permissions(
    mode = "readonly",
    file_write = FALSE,
    tool_allowlist = names(analyst_chat$get_tools())
  ),
  usage_limits = UsageLimits(max_requests = 8),
  history = "retain",
  callbacks = "replace"
)

The copy keeps the chat’s provider, model, system prompt and tools, and the original chat is left alone. history = "retain" keeps its turns and "fresh" starts empty. callbacks = "replace" acknowledges that Deputy replaces the chat’s tool callbacks with its own. The tools still share whatever their closures hold with the original chat. Each call is also limited by the owner’s current permissions.

The curated-chats example app retains two specialists this way and lets you follow up with one of them. It runs against a local test server, so it needs no API key:

shiny::runApp(system.file("examples", "curated-chats", package = "deputy"))

A graph of specialists

retain_agent_graph() retains several agents at once and declares which may delegate to which. Each route becomes a tool on the calling agent, with a fixed target and allocation:

handles <- root$retain_agent_graph(
  agents = list(analyst = analyst, reviewer = reviewer),
  routes = list(
    root = list(
      analyze = list(
        target = "analyst",
        description = "Analyse the evidence.",
        usage_limits = UsageLimits(max_requests = 4)
      )
    ),
    analyst = list(
      review = list(
        target = "reviewer",
        description = "Check the analysis.",
        usage_limits = UsageLimits(max_requests = 2)
      )
    )
  ),
  usage_limits = UsageLimits(max_requests = 12, max_tool_calls = 8),
  max_depth = 2,
  max_delegations = 8,
  max_concurrency = 2
)

result <- root$run_sync("Analyse the evidence and have the analysis reviewed.")
root$delegation_graph_usage()

Here the root may ask the analyst, and the analyst may ask the reviewer. The graph’s usage_limits cover every run in the graph until you release it, including your own follow-ups with continue_agent(handles$analyst, ...). max_depth counts delegation steps from the root, max_delegations caps the total number of delegations, and max_concurrency caps how many run at once. An agent waiting for its own delegate counts toward concurrency, so this three-level chain needs two slots. Routes can form a cycle, but a specialist that is already busy can’t be called again.

Token and cost limits are checked when responses arrive, so work already in flight can overshoot them; request limits are checked before each request.

Permissions and tool hooks apply all the way down: a call made by the reviewer must also pass the analyst’s and the root’s permissions and PreToolUse hooks. A root in read-only or plan mode can still use its routes, and that mode then applies to the whole graph. Each specialist keeps its own provider, prompt and tools.

The root sees the whole graph through the same methods as a lead agent: list_subagents(), inspect_subagents(), observe_subagents() and interrupt_subagent(), described in Briefing and inspecting subagents. Interrupting one specialist also stops the delegations it started. When the graph is idle, root$release_agent_graph() removes the routes and handles.

The recursive-agents example app runs a three-level graph against a local test server:

shiny::runApp(system.file("examples", "recursive-agents", package = "deputy"))

Start a specialist from part of a conversation

Sometimes a specialist needs to see part of an existing conversation, for example the messages about one dataset. ContextFork() describes the turns to copy and where they came from, and fork_agent() retains a new specialist that starts with them:

fork <- ContextFork(
  owner_id = "user-1",
  conversation_id = "chat-7",
  branch_id = "main",
  revision = "rev-42",
  fork_point = 12,
  view = "context",
  turns = selected_turns,
  max_bytes = 1024 * 1024
)

handle <- fork_agent(
  owner,
  specialist,
  fork,
  authorize = check_user_can_read,
  usage_limits = UsageLimits(max_requests = 4),
  max_runs = 2
)

owner$continue_agent(
  handle,
  "Review the selected evidence.",
  UsageLimits(max_requests = 2)
)

Your application chooses the turns, from the complete conversation (view = "transcript") or from what the model currently sees (view = "context"), and describes where they came from. specialist must be a new Agent with an empty conversation, and a selection larger than max_bytes is an error. authorize is your function: it receives the fork’s description, checks that the current user may still read that conversation, and returns its owner_id, conversation_id, branch_id and revision. Deputy calls it again before every continuation and stops if the answer changes.

The copied turns are history only. They keep text, images, documents and completed tool calls, but not tool bindings or anything provider-specific, and unfinished tool calls become plain text. The specialist uses its own prompt, tools and permissions, and after the fork the two conversations are independent. Deputy doesn’t create or store branches; that’s up to your application.