Workspace tools
There is no read_file here, and no edit_file. Reading and editing want opposite things: reading wants the exact bytes, lazily, addressable by line so the model can grep a 50 MB log without pulling it into the prompt, while editing wants a mutable copy, because producing new bytes is the operation. One verb serving both cannot be lazy and mutable at once.
So open_file observes and stage_file authors, and the return type tells you which you are holding: a queryable artifact, or an in-memory Media that is not on disk until you save it.
There is no dedicated git tool either, and that is not an omission: run_shell_command is the git tool, with the real binary and its real output. See run_shell_command.
Register the tools
createSandboxTools requires the real filesystem, path translator, handle, write root, trust tier, and gate. The backend types are deliberately explicit: do not replace them with a permissive any.
import { createSandboxTools } from '@nhtio/adk/batteries/sandbox'
import type { SandboxToolsOptions } from '@nhtio/adk/batteries/sandbox/tools'
const options: SandboxToolsOptions = {
handle: sandbox,
fileSystem: workspaceFileSystem,
pathTranslator,
writeRoot: '/workspace/out',
trustTier: 'untrusted',
gate: async (_ctx, call) => {
const approved = await approvalService.ask(call.tool, call.args)
return approved ? { approved: true } : { approved: false, note: 'operator declined' }
},
search: searchBackend,
}
const tools = await createSandboxTools(options)
registry.register(tools)The factory includes open_file, open_json_file, and open_markdown_file, plus stage_file, save_media, list_directory, search_files, and find_files. registeredTools can be supplied when descriptions need to reflect a subset, but it does not make an unregistered capability appear.
Gates are mandatory, including reads
The obvious objection is that gating a read is paranoid — a read changes nothing, so why suspend a turn for one? The objection has the threat model backwards.
A bad write can be reverted. A read cannot. The moment .env or id_rsa is read, the secret is in the turn, the transcript, and the provider's logs, and no amount of care afterwards puts it back. That is the asymmetry: writes are recoverable, reads are terminal. And search_files is worse than a read — it is a secret-discovery primitive, because a model-supplied regex finds credentials without knowing where they live. A gate on writes and not reads protects the recoverable half.
So every tool here gates before it translates or opens a path, and no reference gate ships — a default gets adopted unread as "the safe one", carrying an arbitrary assumption about which subtrees are boring.
A blanket gate is a decision
gate: async () => ({ approved: true }) on the read tools is legal and supported. It is also this:
You have just handed the model an unsandboxed shell. That was a choice. Own it.
Every file the policy permits is now one tool call from the transcript.
A gate is a real suspension — the turn stops until a decider answers, so a headless harness without one hangs indefinitely. Wire a decider, or wire an explicit auto-verdict knowing that is what it is.
Denials throw a narrated E_SANDBOX_REFUSED; operational failures throw a narrated E_SANDBOX_*. Both arrive inline on every backend, which is the entire reason they throw rather than return — the mechanics are on the index page.
The query path: open_file → artifact_grep
open_file streams the file into the turn as a spooled artifact and hands back a handle, not text. That is uniform inline: false behaviour across all six adapters, and the handle is the point rather than a limitation: a 5 GB file opens with bounded memory, and the model pages through it with the artifact_* tools instead of spending its context on bytes it will not read.
// Model/tool calls, in order:
const opened = await dispatch('open_file', { path: 'src/config.ts' })
// `opened` is an artifact handle, not source text.
const matches = await dispatch('artifact_grep', {
callId: opened.callId,
pattern: 'timeout',
maxResults: 20,
})The actual artifact tool is forged by core from the live result; the workspace battery does not mint a private read_file or artifact_grep. If an adapter prints "[object Object]", a handler returned something its adapter could not wrap; return a supported string/bytes/Media/artifact result instead.
Searching: what the arguments do, and the one that is refused
search_files searches file contents; find_files matches names. They share a bounding pair and diverge on the flags that only mean something when there is a pattern to match:
| Argument | search_files | find_files | Effect |
|---|---|---|---|
pattern / glob | pattern, required | glob, required | what to match |
limit | optional | optional | integer >= 1; truncates results. Omitted = unbounded. |
max_depth | optional | optional | traversal boundary. Omitted = unbounded (no --max-depth flag). |
ignore_case | yes | — | case-insensitive matching |
literal | yes | — | treat the pattern as a fixed string, not a regex |
glob / iglob | yes | iglob only | narrow which files are searched |
hidden | yes | yes | include dot-prefixed entries |
no_ignore | yes | yes | ignore .gitignore and friends |
follow | rejected at validation unless declared | rejected at validation unless declared | see below |
ignore_case and literal are absent from find_files' schema rather than accepted and ignored: they are no-ops for a filename walk, and a silently-ignored argument is worse than a rejected one.
limit bounds what reaches the model, not what the adapter holds. It is post-collection truncation — rg --files on a hostile tree still buffers fully — so it is a context-window control, not a memory guarantee. Exceeding it is the ordinary outcome: you get the bounded results plus a result-limited note naming the value to raise. It is not an error, and a tool that turned it into one would make broad searches unusable.
Omitting limit or max_depth means unbounded — there is no default cap. This is deliberate and it is the contract (issue #48): a deployment that needs a bound names one; the framework never picks one on the model's behalf, because Number.MAX_SAFE_INTEGER is a disguised cap and an undocumented default is an invisible one. Omitted max_depth emits no --max-depth flag at all (rg's own unbounded behaviour), and omitted limit yields every match followed by a terminal { kind: 'done', complete: true }. An EXPLICIT limit must still be an integer >= 1 — 0, negatives, non-integers, Infinity and NaN all reject at validation, so a caller cannot name a bound and silently get an unbounded search.
An unbounded search buffers its whole output
Because limit is applied after the adapter has collected rg's stdout, an unbounded search on a very large tree uses memory in proportion to its output — the results are held before the first frame is yielded, and an explicit limit truncates only afterwards. Pass a limit or max_depth wherever that matters (an untrusted or unknown tree, a long-lived process). Omission is the right default for interactive, bounded lookups; it is not a memory guarantee.
The same frame means opposite things to the two operations, and that is the point.list_directory has no limit — its only boundary is max_depth — so an over-limit frame from a list backend is impossible by construction and is treated as a protocol violation: a narrated io-failure naming the backend, rather than being silently reported as depth-limited, which would give the reader wrong advice. search_files and find_files take an optional limit, so for them the identical frame is the ordinary truncation described above — and when limit was omitted the frame is never emitted. Custom SandboxSearch adapters MUST honour omission the same way: no result cap when limit is absent, no depth flag when maxDepth is absent, and a terminal complete: true when an unbounded scan finishes. An adapter that cannot honour omission (a backend with a hard internal cap) must refuse unbounded calls rather than silently imposing its own bound.
The rule is therefore per operation, not per frame. Done is one union shared by both, which is the only reason a bound: 'limit' frame is even expressible on a list. Copying the protocol-violation throw into the search path is a live hazard rather than a hypothetical one — it makes every broad find_files query fail on its own limit, which is its normal successful outcome.
follow is rejected unless your adapter declares it can contain symlinked descendants
This looks like an inconsistency. It is a deliberate split, and the reason is the BYO seam.
rg --follow traverses symlinked descendants, which never pass through the path translator — so with an unbounded backend it is an unbounded read of wherever those links point. Whether the enforcer contains them has not been audited, and the bundled createRipgrepSearch therefore throws, naming the pending audit rather than pretending the flag works.
The tools layer cannot inspect what is behind SandboxSearch, so the adapter declares it:
const mySearch: SandboxSearch = {
supportsFollow: true, // only if you have actually verified containment
async *searchContent(o) { /* ... */ },
async *findPaths(o) { /* ... */ },
}Undeclared (the default, and the bundled adapter's position), the forged schemas reject follow: true at validation — no option is advertised that always fails. Declared, the full option is restored. The rule is built with .valid(false) on a copy of the permissive schema, so narrowing the bundled case does not disable the flag for a deployment that has earned it.
The bundled createRipgrepSearch additionally throws if it somehow receives follow: true, naming the pending audit — defence in depth, since a caller may reach the adapter without going through the forged tools at all.
In the source the rule is a named binding, followRejectedUnlessDeclared, rather than an inline validator: the schema entry reads follow: followRejectedUnlessDeclared, which says at the use site what the option actually does.
The mutation path: stage → patch → save
stage_file returns a reader-backed Media and writes nothing. A media verb produces another media result, still in memory. save_media is the step that touches the disk: it resolves the media_id, checks the explicit write root, refuses symlinked components, and writes.
The unsaved-mutation trap
A model that stages a file, patches it, and stops has changed nothing on disk. git diff shows clean. This is the single most likely authoring mistake in the battery, which is why both tool descriptions lead with it — stage_file says the edit is in memory, and save_media says it is the step that makes it real.
const staged = await dispatch('stage_file', { path: 'src/config.ts' })
const patched = await dispatch('apply_patch', {
media_id: staged.id,
patch: { find: 'timeout: 30', replace: 'timeout: 60' },
})
await dispatch('save_media', {
media_id: patched.id,
path: 'out/config.ts',
})apply_patch is a media verb supplied by the surrounding media/tool assembly, not one of the eight workspace tools returned by createSandboxTools; use the verb your media registry actually registers. The important invariant is that save_media is the commit. A staged Media has no serialisable reader descriptor. Encoding it before saving throws E_READER_NOT_DESCRIBABLE.
Policy and residuals
These tools are not OS-enforced. run_shell_command and the search tools spawn binaries this library did not write, so they get the OS boundary. These tools run library code, so there is no untrusted binary between the policy check and the open(), and the in-process evaluator is the boundary — the same rules, read from SRT's own derived config.
That decides how much to trust each path. A denyRead on a credential file stops open_file because the evaluator checks it, and stops run_shell_command because the kernel does. The first is defeated by a bug in the evaluator; the second is not.
What that leaves unmitigated:
- TOCTOU. A symlink swapped between the check and the open wins. There is no OS backstop on this path, and spawning a helper would not help — the helper races identically.
- Evaluator bugs. See above. The compensating controls are the mandatory gate and a narrow
writeRoot, not a claim of correctness. - Clobbering. A permitted write can overwrite a build output or a script the agent later runs, and nothing can distinguish that from intended work.
list_directory's existence-hiding covers the result channel only. A denied entry is omitted silently and changes neither the paths nor their order — but it still costs areaddirand astat, so at scale it shifts operator-visible latency, and a racing descendant can still turn a clean listing into anio-failure.
Which is why a deployment holding secrets keeps them outside the sandbox root, rather than trusting denyRead to hide them.
Troubleshooting
| Symptom | Meaning and fix |
|---|---|
E_READER_NOT_DESCRIBABLE on encode() | Save the staged media first; a staged reader is intentionally not a durable descriptor. |
| Artifact handle where text was expected | Expected inline: false; call artifact_grep or another artifact query tool. |
| Hung turn | A mandatory gate has no decider. |
"[object Object]" | The adapter could not wrap the handler's returned object; use the declared return shapes. |
| Search finds a secret | The read/search gate was too broad. Gates are required precisely because discovery is exfiltration. |