deputy (development version)
ContextPolicy()gainsestimator. With the default"auto", automatic compaction now runs when the provider cannot count tokens: it adds a conservative estimate of later content to the usage reported for the latest response. Previously such providers, including gateways whose token-counting route returns HTTP 404, never compacted."provider"keeps the old behaviour. A token-counting endpoint that returns HTTP 404, 405 or 501 is asked only once per provider class and base URL (#213).New
LeadAgent$set_delegation_sources()replaces the sources a subagent can be given as evidence, so sources that change during a conversation (a drawing revised, a document added) can be offered. The scope can change only between runs; earlier subagent records belong to the old conversation, so the change is refused while they exist unlessclear_records = TRUEdiscards them. An evidence reference that doesn’t match now lists the sources available (and a stale revision names the current one) instead of “Requested evidence is unavailable.”RSession,tool_run_r_codeandtool_run_bashno longer pass the host’s environment variables to the process that runs model-written code, and R there no longer reads~/.Renvironor a project.Renviron. The process gets the variables that locate programs, libraries, locales and temporary files. Name others, such as proxy settings, with the newenvargument toRSession$new(),tools_code(),tools_preset()andtools_all(), or the command-line app’s--code-env;env = "inherit"passes everything, as before. This keeps credentials out of what the code is given, not out of its reach: code running as the same user can still read the starting environment of the R process through the operating system, and any file that user can read (#224).RSession$new()gainslibpath, the library directories the R process loads packages from, in order, so a host can give its worker a library of its own ahead of the site library while keeping it off its own search path.New
Agent$set_chat()replaces the Chat an agent sends requests to, so a host can continue a conversation with a model from another provider. The conversation, system prompt and tools move to the new Chat, reasoning content is dropped from the history, and the agent’s permissions, hooks, tool observers and usage limits keep applying.New
Agent$set_context_policy()replaces theContextPolicybetween runs, for example to compact at a different size after changing model. The policy must keep the sameoffload_dir.LeadAgent$register_sub_agent()gainsreplace. Withreplace = TRUEit replaces a registered definition with the same name and updates the lead’s prompt; delegations already running keep the definition they started with.$stream()and$stream_async()reset a cancelled stream controller when a new run starts, as ellmer does. A host such as shinychat that reuses its controller no longer sees the previous run’s cancellation when the next run starts.A write restricted to a
file_writedirectory is checked again just beforewrite_file,edit_fileormulti_editruns. If a file or symbolic link in the path changed after the permission check so that the path now resolves outside the directory, the call is refused instead of writing there. A write that aPermissionRequesthook allowed outside the directory is not checked again (#219).File writes restricted to a directory no longer refuse names that merely contain two dots, such as
notes..v2.md. Only a..path segment counts as path traversal (#219).Native tool names no longer grant native permission treatment to other tools (#216). A host, skill or package tool named
read_file,ask_user,web_fetchor any other native name (including variants such as"Read-File") is now checked like any other custom tool: missing annotations take the conservative defaults, readonly mode does not treat it as a known read, and plan mode’s prompt-tool shortcut applies only to Deputy’s ownask_user. The name’s restrictions, such as readonly’s denial of write tools and write-path limits, still apply. A host-chosen prompt tool name keeps its shortcut, and directpermissions_check()calls, including ones built fromtool_metadata(), keep name-based classification.edit_fileandmulti_editnow change only the replaced text. They previously rewrote the whole file throughreadLines()/writeLines(), converting CRLF line endings to LF and adding a final newline. Matching still reads every CRLF as LF, so multi-line edits written with"\n"apply in Windows and mixed files; new lines take the edited line’s ending. Files are compared and written as bytes, so non-UTF-8 content is kept (#217).A
can_use_toolpermission callback can no longer allow a call that the rest of the policy denies. In standard mode its allow used to skip the capability checks, so a callback that allowed everything it didn’t block let the model run R code withr_code = FALSEor write outside thefile_writedirectory. Plan and full modes ignored the callback. It is now called in every mode, for each call the policy allows, and can deny the call or pause it for approval. To allow a call the policy denies, use aPermissionRequesthook.Read-only and plan policies now allow an agent’s own delegation tools: a
LeadAgent’sdelegate_to_agent, and thedelegation_tool()andretain_agent_graph()route tools that call its retained agents. An agent created withpermissions_readonly()orpermissions_plan()can therefore delegate, and each tool call its subagents and retained agents make is still checked against its policy, so they can’t write or run code either. A custom tool doesn’t qualify by using one of these names (#227).hook_log_tools(),hook_block_dangerous_bash()andhook_limit_file_writes()now returnNULLwhen they don’t deny a call, so hooks added after them for the same event still run.Agent$usage()andAgent$cost()now include turns that compaction removed from the model’s context, andAgent$usage()$tool_callscounts the tool calls the model asked for instead of always being 0.The command-line app now reports how many model requests a run made, instead of “NA turn(s)”.
edit_fileandmulti_editnow change only the text they replace, instead of rewriting the file, which converted CRLF to LF and added a final newline. Non-UTF-8 bytes are kept too. Search text written with"\n"still matches CRLF and mixed files, and new lines take the ending of the line they replace (#217).run_bashnow treats a non-zero exit status, including “command not found”, as an error and tells the model the status. Standard error is returned after a[stderr]line instead of being discarded (#215).tools_interactive()handlers can now answer without blocking, for apps such as Shiny. A handler may return a promise for the answers, orAskUserDeferred()to show the questions and have the model end its turn so the answers arrive as the user’s next message.AskUserDeferred(extra = )attaches display data to the tool result. Subagents accept promises but not deferral (#215).New
Agent$microcompact()replaces old tool results in the model’s context with a short marker, like Posit Assistant’s/microcompact. Results in the lastkeep_lastturns, or from tools inkeep_tools, are kept. It makes no model call, keeps any earlier compaction summary, and leaves the original results in$get_turns(),$last_turn()and saved sessions (#209).agent_definition(model = )can now be a bare model id such as"gpt-6-luna", which runs the subagent with that model on the lead’s provider, endpoint and credentials, including gateway clients built outside ellmer."inherit"and"provider/model"work as before; ids containing/still need the"provider/model"form (#210).New
TrustedResults()supports the trusted mini-agent pattern from Will Landau and Sam Parmar’s Trusted Mini-Agents: pass it toAgent$new(trusted_results = )to name the one local tool that may produce each kind of result. That tool’s return value reaches your app unchanged, as a"trusted_result"event (seeresult_trusted_results()) and through an optionalon_resultcallback, before the model sees it; withmodel_receipt = TRUEthe model gets only a receipt. Registration fails if another tool could produce results: code execution and delegation tools always, and tools that may write or reach the open world unless listed inexempt_tools(#197).Permission callbacks and
PreToolUsehooks now get the tool’s argument types incontext$tool_arguments. Newtool_input_review()turns a proposed tool input into a table of each field’s type, description and value (#197).LeadAgent$new(trusted_results = )applies the policy to all subagents: each definition’s tools must pass the same check, a designated tool must be the same tool everywhere, and subagents’ trusted results reach the lead’son_resultand run events (#197).New
approval_review_ui()andapproval_review_server()are a Shiny module for reviewing a pending durable approval. They show each argument’s type, description and value, let the reviewer edit simple fields, and approve or deny. Passdecideto carry out the decision elsewhere, such as in a background process (#197).New
inst/examples/trusted-results/app, adapted from Landau and Parmar’s R template and weather example, shows chat, input review and results side by side. Only the forecast tool can fill the results panel (#197).New
vignette("trusted-mini-agents")explains Landau and Parmar’s trusted mini-agent pattern and how Deputy enforces each of its rules (#197).New
mcp_console_connection()connects an agent to an MCP Console 0.0.4 server, a sandboxed workbench that keeps R, Python and DuckDB SQL state for a conversation, andmcp_console_control()interrupts or restarts it (#190). Deputy refuses--no-sandbox, options that widen file access, add a proxy or select a remote target, and an unreviewed.agents/console/config.yaml, and closes the connection unless the server reports its sandbox with restricted networking.sendcounts as shell code execution and needs thebashandwebpermissions; installing dependencies runs outside the sandbox and also needsdependencies = "allow"andinstall_packages = TRUE.timeout_msis capped at 2500 ms, so long cells return a running marker and the model polls. MCP Console records every call and result, unredacted, under.agents/console/sessions/in the agent’s working directory.Deputy temporarily requires coro >= 1.1.0.9000 from GitHub: CRAN coro 1.1.0 recompiles every generator, costing 0.3 to 0.7 seconds of CPU per model request. Deputy will return to CRAN coro once the fix is released (#192).
McpConnectionnow checks the JSON-RPC id of every MCP stdio reply. The supported mcptools releases wait about 4 seconds for a reply, then take the next output line without checking its id, so a slow reply could become the answer to the next call. A missing or mismatched reply now stops the server, losing its session state, and signals adeputy_mcp_desynchronizederror;tools_mcp()tools, which can’t see reply ids, stop the server on the first lost reply.mcp_repl_connection()andtools_mcp_repl()cap mcp-repl’stimeout_msat 3000 ms, also when it is omitted, so long cells return a busy result and the model polls (#196).McpConnection,tools_mcp()andmcp_repl_connection()now support CRAN mcptools 1.0.3 as well as 1.0.2. Other versions are refused with an error that lists the supported ones (#195).Model requests and tool calls use less CPU (#185).
New
job_create(),job_run(),job_read()andjob_cancel()save an agent task to disk for your own scheduler to run later, in any R process. A job keeps the definition and context revisions you supply, its budgets, completed tool effects and any pending approval. A job interrupted mid-run is marked indeterminate instead of retried, so side effects aren’t repeated;job_cancel()asks a running job to stop at its next checkpoint (#42).New
ContextFork()andfork_agent()start a retained specialist from a copy of turns you select from another conversation, either its transcript or its current model context. Each continuation rechecks that the source may still be read. The copy is inert: tool bindings and private provider data are dropped, and current permissions apply (#62).Subagent inspection and observation now keep
difftimevalues, their units and missing values, including in data-frame columns. The subagent chat panel shows content nested in tool results, labels where it came from, notes anything left out, and keeps the original for replay (#166, #167).New
Agent$retain_agent_graph()sets up retained specialists that delegate to each other along routes you declare. The graph shares one lifetime budget with limits on depth, delegations and concurrency; cancelling a run cancels everything beneath it; and the root controls who may view each specialist’s conversation. Seeinst/examples/recursive-agents/(#169).New
adopt_chat()turns a configured ellmer chat into a retained specialist, copying its provider, system prompt, tools and optionally its history, and replacing its callbacks so the owning agent’s permissions and hooks apply.delegation_tool()gives the owning agent a tool for sending it tasks (#168).New
Agent$retain_agent()keeps a specialist agent and its conversation for repeated use: continue it with$continue_agent()or$continue_agent_async(), stop it with$cancel_agent()and free it with$release_agent(). Its budget is cumulative, the handle works only with the agent that retained it, and continuing it while it runs is an error. PlainAgents now have the subagent inspection and observation methods too (#152).New
inst/examples/trusted-mini-agent/app: a subagent proposes inputs for a small plant-weight analysis, a person reviews them, and only a designated R tool produces the result (#154).New
subagent_chat_ui()andsubagent_chat_server()add a read-only Shiny panel showing each subagent’s live activity and conversation, with shinychat tool cards and attachments. It can replay saved history and shows a cancel button if you supplyon_cancel. Seeinst/examples/subagent-chats/(#158).New
LeadAgent$observe_subagents()returns a subscription for following subagent activity: a snapshot, then new events each time you poll, with access checked on every read. Events dropped from the limited buffer, or too large to keep, are reported as gaps. Closing a subscription doesn’t stop anything;$interrupt_subagent()cancels a subagent (#157).Delegation now returns a
DelegationOutcome, keeping the subagent’s reply separate from its history.DelegationDisclosure()decides who may inspect subagents with$inspect_subagents(),$read_subagent_result()and$export_subagents(), which return redacted views.delegation_history()replays exported history without running tools or changing the lead’s context. Long replies are offloaded like large tool results (#153).New
DelegationPolicy()controls subagents’ tools and human input. Tools can be shared, held exclusively (overlapping use is an error), or built per delegation by a factory returningDelegationResources(), which are cleaned up when the subagent finishes or fails to start. A subagent usingask_userneeds a handler from the policy, and durable approvals in subagents are an error for now. Each delegation’s manifest records the policy (#151).New
DelegationInput()describes a delegated task as a brief (task, constraints, evidence, deliverable and stop conditions); plain strings still work. Evidence names exact revisions of records passed toLeadAgent$new(delegation_sources = )and is checked before any request; each delegation’sDelegationManifestrecords what the subagent received, and text in a brief can’t pose as the evidence list. Directly invoked subagents now stop when the lead is interrupted, parallel delegation enforces the lead’s token limit, and a malformed brief no longer loses the other tasks in a batch (#149).LeadAgent$list_subagents()now lists queued and running delegations in the order they were accepted, with exact stop reasons and any hook error; stopped runs no longer appear completed. Interrupting the lead interrupts its running subagents, and queued parallel tasks stay listed after cancellation (#150).Agent$chat(),$stream()and their async versions accept dynamic dots (!!!), so shinychat can pass its input straight through. The Shiny chat example now shows compaction progress and the summary.Agent$get_turns()andAgent$turns()now return the whole conversation after compaction, so shinychat history keeps every reply; newAgent$get_context_turns()returns what the model receives. Saved sessions (schema 3) store both, and sessions saved with earlier development schemas can’t be loaded (#146).RSession$new(agent, tools = )lets model-written R code call selected registered tools astools$name(...), through the agent’s permissions, hooks and limits, and keep the results as R data. The tools must useconvert = FALSEand can’t pause for durable approval (#186).New
RSessiongives a conversation a persistent R process forrun_r_code: variables persist between calls, output and plots come back in order, and calls queue.$cancel()or a timeout discards the variables, and the next call says so. The code runs with your account’s access and is not sandboxed. NewContextPolicy()argumentsmax_tool_result_imagesandmax_tool_result_image_byteslimit images in tool results separately from text; the full result stays available for display (#143).New
mcp_repl_connection()gives an agent its own sandboxed mcp-repl session;mcp_repl_control()interrupts or resets it and reports the outcome. Interrupting the agent cancels its active MCP connections. Plots and output previews come back as ellmer content (#69).New
McpConnectionconnects one agent to an MCP server through a client in a separate R process. You can page through the server’s tools, resources and prompts, and must allow each one before use. Calls return promises; after a cancellation, timeout orclose()the connection’s tools stop working. It supports mcptools 1.0.2 and is a stopgap until mcptools has a public client API (#48, #99).New durable approvals: in an agent with an
approval_dir, a permission callback can returnPermissionResultPending()to pause a tool call and save it to disk for a person to decide later, even in another R process.approval_read()shows the pending call, andAgent$resume_approval()approves it (optionally with edited input or a higher budget) or denies it, checking both the saved and the current permissions. The tool must useconvert = FALSE. Completed tool effects are logged, a call interrupted mid-run can’t be resumed and must be checked by hand, and your app tracks which conversation each approval belongs to (#43).Hook and permission result constructors, such as
HookResultPreToolUse()andPermissionResultAllow(), now return read-only S7 objects. Check their class withS7::S7_inherits()(for example againstHookResultorPermissionResult) and useS7::props()for a plain list.continueandinterruptmust be a singleTRUEorFALSE, and text fields must be strings;suppress_outputis still coerced withisTRUE()(#129).Skill()now returns a read-only S7 object.skill$check_requirements()is replaced byskill_check_requirements(skill); to change a skill, build a new one fromS7::props(skill).skill_create()andskill_load()are unchanged (#127).agent_definition()andAgentDefinition()are now the same constructor and return a read-only S7 object; build a changed definition fromS7::props(). Names, YAML files and delegated limits work as before (#125).AgentUsage()andUsageLimits()now return read-only S7 objects.$still reads fields; useS7::props()instead of[[orunclass()for a plain list (#121).AgentResult()andPermissions()now return read-only S7 objects, and their$new()and methods are removed. Inspect results withresult_n_turns(),result_tool_calls(),result_tool_results(),result_text_chunks()andresult_is_success(), and test a policy withpermissions_check().$still reads fields.Permissions()now rejects malformed flags and callbacks (#117).AgentEvent()andHookMatcher()now return read-only S7 objects. Create matchers withHookMatcher(...)instead ofHookMatcher$new(), and test a tool name withhook_matches(). Check an event’s type withevent$typeinstead of its S3 class;event$dataholds the payload, and$still reads its fields directly (#59).ContextPolicy()andDeputyCompaction()now return read-only S7 objects. Read fields with$, or useS7::props()for a plain list (#123).print()methods now format with cli, wrap to the console width, show braces in your values literally, and write to stdout socapture.output()works (#58).Compaction summaries now include tool results, formatted by ellmer so structured results keep their field names. Previously a source returned by a tool could be missing from the summarizer’s input.
New
inst/examples/history-recovery/experiment tests whether a compacted agent answers better when it can also search and read the earlier conversation, including after a later instruction replaces an earlier rule (#112).Automatic compaction now runs within the current run: it shares the run’s usage limits, stops when the run is interrupted, and can happen between tool rounds without losing completed tool calls or usage. New
ContextPolicy(summary_fallback_chats = )lists chats to try in order when the summary request fails with a transient error, separately from the agent’sfallback_chats. A failed or interrupted compaction leaves the context unchanged, an accepted summary survives a later task fallback, and each summary attempt is recorded (#111).Deputy now requires ellmer 0.5.0 or later.
Structured output now uses ellmer types. Pass
typeto$run(),$run_sync()or$run_async()and the agent finishes its tool calls, then returns structured data within the same budget. An optionalvalidatefunction can ask for up tomax_correctionscorrections, and each attempt is recorded. Theoutput_formatargument, its JSON parsing and validation helpers, and the jsonvalidate dependency are removed.Agent$new()andLeadAgent$new()acceptfallback_chats, tried in order when a request fails with a transient error before any response or tool request. A request isn’t retried once output has arrived or a tool has run. Failed requests still count toward usage, and unknown costs stay unknown.With otel installed and a tracer configured, each run gets a
deputy.runspan around ellmer’s spans, including for async runs and subagents, with events for permission decisions, hooks, compaction, fallbacks and delegation. These spans contain no prompts, tool arguments or results; ellmer’s message capture stays opt-in. The standalone10-evaluation.Rexample joins evaluation cases to run IDs.After a run stops with a limit error,
Agent$last_run()still returns its result.The
deputycommand-line tool now defaults to OpenAI’sgpt-5.6-luna. Models you choose, other providers’ defaults and chats you supply are unchanged (#106).New standalone example
09-debate.R: two subagents argue for and against a question in parallel, and a moderator weighs their arguments with the bundleddebateskill. If either side fails, the script stops before the moderator (#40).New
LeadAgent$parallel_delegate()and$parallel_delegate_async()ask several tool-free subagents for one reply each, in parallel, at mostmax_activeat a time. Results come back by name, including partial results when some fail. Each request is reserved from the lead’s budget in advance, failed requests count toward limits, and batches can be cancelled. Each subagent starts from a fresh chat without the lead’s history, tools or callbacks (#39).New
tool_metadata()reports each tool’s origin, declared and missing annotations, and the defaults used, including after cloning and delegation. MCP tools now keep the server’s annotations, connect only to the servers you name, and stop working after their connection reconnects. An MCP tool named like a built-in tool gets none of its privileges, and its path arguments are not rewritten. Subagents take their tools from their definition instead of inheriting the lead’s (#50).Registering a tool whose name is taken is now an error unless you pass
replace = TRUE; duplicate names within a batch are always an error. Each batch is checked in full before any tool is added, fromAgent$new(),$register_tools(),$set_tools()or a skill. A custom tool without annotations is treated as possibly writing, destructive and reaching outside the workspace (#49).New
agent_definition_read(),agent_definition_write()andagent_definitions()read and write agent definitions as YAML files, by default in.deputy/agents/. Tools and skills are looked up by name in registries you supply; reading a file never runs code (#41).A missing suggested package now triggers the standard install prompt, naming the feature that needs it. Loading a skill with YAML front matter now requires yaml instead of silently dropping the metadata (#57).
Agentcan now stand in for an ellmer chat:$chat(),$chat_async(),$stream()and$stream_async()work like ellmer’s, and they,$run_sync()and$run_async()all apply the agent’s permissions, hooks, limits, checkpoints and usage tracking. shinychat can useagent$stream_async()directly, attachments included.run_shiny()and public access to the wrapped chat are removed, andLeadAgentdelegation no longer blocks the R process.New
ContextPolicy()compacts the conversation automatically before a request when it gets too long, reporting whether it used the model, aPreCompacthook or the text fallback, and its own usage. It also stores large tool results on disk behind a short reference the model can read in chunks;PostToolUsehooks still see the full result, and a relativeoffload_diris resolved when the policy is created. Saved sessions keep the summary and offloaded results, and loading a session replaces the agent’s offloaded results so they can’t leak into later saves.$set_turns()clears the summary but keeps the rest of the system prompt, andLeadAgentpasses the policy to its subagents.File and code tools now run in the agent’s
working_dirwithout changing R’s working directory.Concurrent delegations from a
LeadAgentnow share its remaining usage limits instead of each receiving the full balance. A clonedLeadAgentnow delegates with the clone’s own subagents, hooks and run history.The
deputycommand-line tool is now a Rapp 0.4 executable: run it once withrxor install a launcher withir tool install. Its task and interactive modes now stream correctly, report tool failures and can resume saved sessions.Agent$new()and every run method acceptrun_context, a JSON-compatible list of your own identifiers, such as a user or conversation ID, which is attached to results, hook contexts, saved sessions and subagents. ID fields set when the agent was created can’t be changed per run. Tool events and subagent results also carry agent, run, parent, tool call and delegation IDs, and generating them no longer advances R’s random number generator.The public API is now smaller, centred on agents, tools, permissions, hooks, skills, delegation and usage. Removed: the pre-release Agent SDK and Claude compatibility functions, the Claude settings loader, automatic session stores, vendor tool and permission aliases, deprecated run arguments, the todo tools and redundant convenience exports. Each agent has a stable
session_id, and conversations are saved and loaded only withAgent$save_session()andAgent$load_session(); sessions and file checkpoints from earlier versions can’t be loaded.Agent$provider()no longer fails with “Can’t get S7 properties with$” on current ellmer, which moved the model fromProviderto a newModelclass.Agent$set_permission_mode()can now only keep or narrow the permissions the agent was created with. Subagents follow the same rule and keep all the lead’s restrictions: capability flags, tool allow and deny lists, permission callback and write directory.agent_definition()now validates its fields and lowercases names.LeadAgentrejects duplicate names, and$sub_agent_defsis a read-only copy: add definitions with$register_sub_agent()so delegation and the lead’s prompt stay in sync (#79).Agentnow rejects a tool request with a missing or malformed tool name before it reaches usage accounting, permissions, hooks or the tool (#26).hook_limit_file_writes()now checks paths the same way permissions do, blocks escapes through symlinks and through paths that only share a prefix (/data2when/datais allowed), and covers every built-in file-writing tool (#75).HookMatcher()callbacks now run in your R process by default; a positivetimeoutruns them in a subprocess and reports their full error messages.HookMatcher()also validatestimeout, and rejects callbacks that can’t accept the event’s arguments and patterns that aren’t valid regular expressions (#35, #36, #74).skill_check_requirements()now treats malformed or unknown provider names as not matching, while a skill and chat that name the same generic provider still match (#30).Agent$compact()now summarizes with a copy of the agent’s own chat, with tools and callbacks removed. Previously, providers other than OpenAI, Anthropic and Google fell back toellmer::chat_openai("gpt-4o-mini"), which sent the conversation to OpenAI wheneverOPENAI_API_KEYwas set.run_r_codeandrun_bashnow tell the model when a command timed out or its subprocess failed, instead of failing internally or reporting a generic “Command failed” (#27).Agent$cost()now returnsNAwhen the cost of any request is unknown; itscompleteandmissingfields tell a real zero from an unknown total. A run with a cost limit stops with reason"cost_unavailable"when its cost can’t be known, rather than enforcing an understated total (#29).A run now stops with reason
"tool_loop"after three consecutive identical tool calls that return the same result. Small differences in the model’s surrounding text don’t reset the count, but a changed result does, so polling still works (#34).permissions_standard(), andPermissions()policies that don’t setr_code, no longer allow running R code, andtools_preset("standard")no longer includesrun_r_code. The built-in R and shell tools run code with your account’s access and are not sandboxed. For sandboxed R, newtools_mcp_repl()loads mcp-repl only with theread-onlyorworkspace-writesandbox, and refuses missing, inherited, external or unrestricted sandbox settings (#32).tools_web()now behaves as documented. Provider-run web search and fetch tools are checked against thewebpermission at registration, because Deputy can’t see their individual calls (so acan_use_toolcallback can’t approve them either); other provider-side tools are refused. Narrowing permissions to remove web access removes these tools.tools_interactive()now takes acallbackandcontextfor itsask_usertool, so concurrent agents, such as one per Shiny session, each get their own handler. Without a handler,ask_usersignalsdeputy_human_input_unavailable.set_ask_user_callback()remains as a process-wide fallback for single-agent scripts (#76).
deputy 0.0.0.9000
- Streaming now emits
tool_start,tool_end,usageandfile_checkpointevents. Each run has a stable ID,Agent$interrupt()asks a running agent to stop, andAgentResultincludes the run’s usage. - New
UsageLimits()andAgentUsage()track and limit requests, tool calls, input, output and total tokens, and estimated cost for each run. Subagents inherit the lead’s remaining limits, and their usage is added to the lead’s. - New file checkpoints: Deputy’s write and edit tools save each file’s exact bytes before changing it, so the agent can rewind them. The lead and its subagents share one checkpoint journal per workspace, which has a size limit and is kept in saved sessions.
- Stricter permission checks: every file-writing tool respects the allowed write directory, and
"readonly"mode denies unknown, writing, destructive and disallowed open-world tools. Loading a saved conversation keeps the agent’s own permissions, workspace and tools. - When a
PostToolUsehook replaces or suppresses a tool’s output, streamed tool events show the change too.PermissionResultDeny(interrupt = TRUE)now stops an active stream. - New
run_shiny()runs an agent for Shiny with the agent’s limits, file roots and checkpoints. It starts lazily, can be cancelled, and recovers from incomplete tool calls. While a run is active, loading a session, rewinding files or compacting is an error. - MCP status reporting and more detailed subagent run metadata.
- Initial development version.
- Core
Agentclass with streamingrun()and blockingrun_sync()methods. - Built-in tools:
tool_read_file,tool_write_file,tool_list_files,tool_run_r_code,tool_run_bash,tool_read_csv. - Tool bundles:
tools_file(),tools_code(),tools_data(),tools_all(). - Permissions with
permissions_readonly(),permissions_standard(),permissions_full()and customPermissions. - Hooks with
HookMatcherfor thePreToolUse,PostToolUse,Stop,UserPromptSubmitandPreCompactevents. - Delegation with
agent_definition()andLeadAgent. - Skills with
skill_load(),skill_create()andSkill. - Saving and loading conversations with
Agent$save_session()andAgent$load_session(). - Works with any chat provider that ellmer supports.