Skip to content
10 min read · 2,054 words

Workspace tools ​

There is no read_file here, and no edit_file. Reading and editing want opposite things: reading wants the exact bytes, lazily, addressable by line so the model can grep a 50 MB log without pulling it into the prompt, while editing wants a mutable copy, because producing new bytes is the operation. One verb serving both cannot be lazy and mutable at once.

So open_file observes and stage_file authors, and the return type tells you which you are holding: a queryable artifact, or an in-memory Media that is not on disk until you save it.

There is no dedicated git tool either, and that is not an omission: run_shell_command is the git tool, with the real binary and its real output. See run_shell_command.

Register the tools ​

createSandboxTools requires the real filesystem, path translator, handle, write root, trust tier, and gate. The backend types are deliberately explicit: do not replace them with a permissive any.

ts
import { createSandboxTools } from '@nhtio/adk/batteries/sandbox'
import type { SandboxToolsOptions } from '@nhtio/adk/batteries/sandbox/tools'

const options: SandboxToolsOptions = {
  handle: sandbox,
  fileSystem: workspaceFileSystem,
  pathTranslator,
  writeRoot: '/workspace/out',
  trustTier: 'untrusted',
  gate: async (_ctx, call) => {
    const approved = await approvalService.ask(call.tool, call.args)
    return approved ? { approved: true } : { approved: false, note: 'operator declined' }
  },
  search: searchBackend,
}
const tools = await createSandboxTools(options)
registry.register(tools)

The factory includes open_file, open_json_file, and open_markdown_file, plus stage_file, save_media, list_directory, search_files, and find_files. registeredTools can be supplied when descriptions need to reflect a subset, but it does not make an unregistered capability appear.

Gates are mandatory, including reads ​

The obvious objection is that gating a read is paranoid — a read changes nothing, so why suspend a turn for one? The objection has the threat model backwards.

A bad write can be reverted. A read cannot. The moment .env or id_rsa is read, the secret is in the turn, the transcript, and the provider's logs, and no amount of care afterwards puts it back. That is the asymmetry: writes are recoverable, reads are terminal. And search_files is worse than a read — it is a secret-discovery primitive, because a model-supplied regex finds credentials without knowing where they live. A gate on writes and not reads protects the recoverable half.

So every tool here gates before it translates or opens a path, and no reference gate ships — a default gets adopted unread as "the safe one", carrying an arbitrary assumption about which subtrees are boring.

A blanket gate is a decision

gate: async () => ({ approved: true }) on the read tools is legal and supported. It is also this:

You have just handed the model an unsandboxed shell. That was a choice. Own it.

Every file the policy permits is now one tool call from the transcript.

A gate is a real suspension — the turn stops until a decider answers, so a headless harness without one hangs indefinitely. Wire a decider, or wire an explicit auto-verdict knowing that is what it is.

Denials throw a narrated E_SANDBOX_REFUSED; operational failures throw a narrated E_SANDBOX_*. Both arrive inline on every backend, which is the entire reason they throw rather than return — the mechanics are on the index page.

The query path: open_file → artifact_grep ​

open_file streams the file into the turn as a spooled artifact and hands back a handle, not text. That is uniform inline: false behaviour across all six adapters, and the handle is the point rather than a limitation: a 5 GB file opens with bounded memory, and the model pages through it with the artifact_* tools instead of spending its context on bytes it will not read.

ts
// Model/tool calls, in order:
const opened = await dispatch('open_file', { path: 'src/config.ts' })
// `opened` is an artifact handle, not source text.
const matches = await dispatch('artifact_grep', {
  callId: opened.callId,
  pattern: 'timeout',
  maxResults: 20,
})

The actual artifact tool is forged by core from the live result; the workspace battery does not mint a private read_file or artifact_grep. If an adapter prints "[object Object]", a handler returned something its adapter could not wrap; return a supported string/bytes/Media/artifact result instead.

Searching: what the arguments do, and the one that is refused ​

search_files searches file contents; find_files matches names. They share a bounding pair and diverge on the flags that only mean something when there is a pattern to match:

Argumentsearch_filesfind_filesEffect
pattern / globpattern, requiredglob, requiredwhat to match
limitoptionaloptionalinteger >= 1; truncates results. Omitted = unbounded.
max_depthoptionaloptionaltraversal boundary. Omitted = unbounded (no --max-depth flag).
ignore_caseyes—case-insensitive matching
literalyes—treat the pattern as a fixed string, not a regex
glob / iglobyesiglob onlynarrow which files are searched
hiddenyesyesinclude dot-prefixed entries
no_ignoreyesyesignore .gitignore and friends
followrejected at validation unless declaredrejected at validation unless declaredsee below

ignore_case and literal are absent from find_files' schema rather than accepted and ignored: they are no-ops for a filename walk, and a silently-ignored argument is worse than a rejected one.

limit bounds what reaches the model, not what the adapter holds. It is post-collection truncation — rg --files on a hostile tree still buffers fully — so it is a context-window control, not a memory guarantee. Exceeding it is the ordinary outcome: you get the bounded results plus a result-limited note naming the value to raise. It is not an error, and a tool that turned it into one would make broad searches unusable.

Omitting limit or max_depth means unbounded — there is no default cap. This is deliberate and it is the contract (issue #48): a deployment that needs a bound names one; the framework never picks one on the model's behalf, because Number.MAX_SAFE_INTEGER is a disguised cap and an undocumented default is an invisible one. Omitted max_depth emits no --max-depth flag at all (rg's own unbounded behaviour), and omitted limit yields every match followed by a terminal { kind: 'done', complete: true }. An EXPLICIT limit must still be an integer >= 1 — 0, negatives, non-integers, Infinity and NaN all reject at validation, so a caller cannot name a bound and silently get an unbounded search.

An unbounded search buffers its whole output

Because limit is applied after the adapter has collected rg's stdout, an unbounded search on a very large tree uses memory in proportion to its output — the results are held before the first frame is yielded, and an explicit limit truncates only afterwards. Pass a limit or max_depth wherever that matters (an untrusted or unknown tree, a long-lived process). Omission is the right default for interactive, bounded lookups; it is not a memory guarantee.

The same frame means opposite things to the two operations, and that is the point.list_directory has no limit — its only boundary is max_depth — so an over-limit frame from a list backend is impossible by construction and is treated as a protocol violation: a narrated io-failure naming the backend, rather than being silently reported as depth-limited, which would give the reader wrong advice. search_files and find_files take an optional limit, so for them the identical frame is the ordinary truncation described above — and when limit was omitted the frame is never emitted. Custom SandboxSearch adapters MUST honour omission the same way: no result cap when limit is absent, no depth flag when maxDepth is absent, and a terminal complete: true when an unbounded scan finishes. An adapter that cannot honour omission (a backend with a hard internal cap) must refuse unbounded calls rather than silently imposing its own bound.

The rule is therefore per operation, not per frame. Done is one union shared by both, which is the only reason a bound: 'limit' frame is even expressible on a list. Copying the protocol-violation throw into the search path is a live hazard rather than a hypothetical one — it makes every broad find_files query fail on its own limit, which is its normal successful outcome.

follow is rejected unless your adapter declares it can contain symlinked descendants

This looks like an inconsistency. It is a deliberate split, and the reason is the BYO seam.

rg --follow traverses symlinked descendants, which never pass through the path translator — so with an unbounded backend it is an unbounded read of wherever those links point. Whether the enforcer contains them has not been audited, and the bundled createRipgrepSearch therefore throws, naming the pending audit rather than pretending the flag works.

The tools layer cannot inspect what is behind SandboxSearch, so the adapter declares it:

ts
const mySearch: SandboxSearch = {
  supportsFollow: true, // only if you have actually verified containment
  async *searchContent(o) { /* ... */ },
  async *findPaths(o) { /* ... */ },
}

Undeclared (the default, and the bundled adapter's position), the forged schemas reject follow: true at validation — no option is advertised that always fails. Declared, the full option is restored. The rule is built with .valid(false) on a copy of the permissive schema, so narrowing the bundled case does not disable the flag for a deployment that has earned it.

The bundled createRipgrepSearch additionally throws if it somehow receives follow: true, naming the pending audit — defence in depth, since a caller may reach the adapter without going through the forged tools at all.

In the source the rule is a named binding, followRejectedUnlessDeclared, rather than an inline validator: the schema entry reads follow: followRejectedUnlessDeclared, which says at the use site what the option actually does.

The mutation path: stage → patch → save ​

stage_file returns a reader-backed Media and writes nothing. A media verb produces another media result, still in memory. save_media is the step that touches the disk: it resolves the media_id, checks the explicit write root, refuses symlinked components, and writes.

The unsaved-mutation trap

A model that stages a file, patches it, and stops has changed nothing on disk. git diff shows clean. This is the single most likely authoring mistake in the battery, which is why both tool descriptions lead with it — stage_file says the edit is in memory, and save_media says it is the step that makes it real.

ts
const staged = await dispatch('stage_file', { path: 'src/config.ts' })
const patched = await dispatch('apply_patch', {
  media_id: staged.id,
  patch: { find: 'timeout: 30', replace: 'timeout: 60' },
})
await dispatch('save_media', {
  media_id: patched.id,
  path: 'out/config.ts',
})

apply_patch is a media verb supplied by the surrounding media/tool assembly, not one of the eight workspace tools returned by createSandboxTools; use the verb your media registry actually registers. The important invariant is that save_media is the commit. A staged Media has no serialisable reader descriptor. Encoding it before saving throws E_READER_NOT_DESCRIBABLE.

Policy and residuals ​

These tools are not OS-enforced. run_shell_command and the search tools spawn binaries this library did not write, so they get the OS boundary. These tools run library code, so there is no untrusted binary between the policy check and the open(), and the in-process evaluator is the boundary — the same rules, read from SRT's own derived config.

That decides how much to trust each path. A denyRead on a credential file stops open_file because the evaluator checks it, and stops run_shell_command because the kernel does. The first is defeated by a bug in the evaluator; the second is not.

What that leaves unmitigated:

  • TOCTOU. A symlink swapped between the check and the open wins. There is no OS backstop on this path, and spawning a helper would not help — the helper races identically.
  • Evaluator bugs. See above. The compensating controls are the mandatory gate and a narrow writeRoot, not a claim of correctness.
  • Clobbering. A permitted write can overwrite a build output or a script the agent later runs, and nothing can distinguish that from intended work.
  • list_directory's existence-hiding covers the result channel only. A denied entry is omitted silently and changes neither the paths nor their order — but it still costs a readdir and a stat, so at scale it shifts operator-visible latency, and a racing descendant can still turn a clean listing into an io-failure.

Which is why a deployment holding secrets keeps them outside the sandbox root, rather than trusting denyRead to hide them.

Troubleshooting ​

SymptomMeaning and fix
E_READER_NOT_DESCRIBABLE on encode()Save the staged media first; a staged reader is intentionally not a durable descriptor.
Artifact handle where text was expectedExpected inline: false; call artifact_grep or another artifact query tool.
Hung turnA mandatory gate has no decider.
"[object Object]"The adapter could not wrap the handler's returned object; use the declared return shapes.
Search finds a secretThe read/search gate was too broad. Gates are required precisely because discovery is exfiltration.