Skip to content
19 min read · 3,841 words

Sandbox Battery ​

This library shipped twenty-three tool categories — vector databases, browser automation, media pipelines — and could not run ls. That gap was deliberate, not an oversight: an ungated shell is the most dangerous thing you can hand a language model. Give an agent raw shell access and one prompt injection turns your host into an exfiltration pipeline. Give it soft path-sanitisation inside your own Node process instead, and a symlink or a ../ spelling you did not think of walks straight through the guardrail. Closing that gap properly means answering two questions that only look like one: what can this process reach? and what can this code reach? This battery answers both, separately, and is honest about where each answer stops.

The thesis ​

1. Two boundaries, deliberately unbundled. What can this process reach? is answered by OS policy — Anthropic's @anthropic-ai/sandbox-runtime, which compiles your rules into a Seatbelt profile on macOS, bubblewrap+seccomp on Linux, WFP on Windows, and no container anywhere. What can this code reach? is answered by a hardened SES Compartment with capabilities you enumerate by hand. A container appears to solve both at once and solves neither cleanly: it brings a daemon, an image, and seconds of startup, while doing nothing about a snippet reaching process.env inside the sandbox it just built for you. Keeping the two apart means the kernel restricts syscalls, SES restricts evaluation, and you can reason about which one failed. Neither substitutes for the other, and a deployment that conflates them has one boundary where it believes it has two.

2. A refusal the model can read, not an errno it can only guess at. When the OS denies something, the child sees Permission denied or a connection reset — indistinguishable from a typo. So the model concludes its command was wrong and retries variations that can never work, flailing against a wall it cannot see, burning your tokens on syntax. The sandbox separately records what actually happened: tried to connect to api.example.com, which is not on the allowlist. Surfacing that is the whole value-add — the difference between a boundary and a mystery. Read the shell tool's page before you rely on it: on macOS today that record comes back empty.

3. Every tool is gated, reads included — because a read is the unrecoverable one. Gating writes and trusting reads is the intuitive split and the wrong one. A bad write can be reverted; a read of .env cannot, because by the time you notice, the secret is in the turn, the transcript, and your provider's logs. search_files is worse than a read: it is a secret-discovery primitive that finds credentials without knowing where they live. So the gate is mandatory on all nine tools and no reference implementation ships — a default gate would be copied unread as "the safe config", carrying an arbitrary assumption about which subtrees are boring.

Assemble one policy and one handle ​

The policy is intentionally boring: reads are deny-then-allow, writes are allow-only, and network is an explicit allowlist unless disabled. gitSafeDirectories must be threaded into the policy. Omitting it is the cause of Git's dubious ownership failure inside an otherwise working shell.

ts
import { createSandbox } from '@nhtio/adk/batteries/sandbox'
import { srtEnforcer } from '@nhtio/adk/batteries/sandbox/node'
import type { SandboxPolicy } from '@nhtio/adk/batteries/sandbox'

const policy: SandboxPolicy = {
  filesystem: {
    allowRead: ['/workspace'],
    allowWrite: ['/workspace/out'],
    denyRead: ['/workspace/.env', '/workspace/.git/hooks'],
    denyWrite: ['/workspace/.git/hooks'],
    gitSafeDirectories: ['/workspace'],
  },
  network: { allowedDomains: ['api.example.test'] },
}
const enforcer = await srtEnforcer({ policy })
const sandbox = await createSandbox({ policy, enforcer, strictMode: true })
// Pass `sandbox` to the tools below. Dispose it when the owning turn/session ends.
await sandbox.narrow({ ...policy, network: { allowedDomains: [] } })

createSandbox is process-global. A second handle must be no wider than the admitted baseline; subsequent SRT drift is detected, not magically prevented. Set allowUnsandboxedFallback only as a loud, audited product choice. Native Windows is not the supported SRT path; use WSL2.

Spawn admission and lifecycle ordering ​

probeSpawn: true is an opt-in construction check for the blind spot between a successful dependency check and a real child spawn. It runs one true command under the admitted policy, drains both output streams, and fails closed. The probe is off by default, and deliberately does not consult allowUnsandboxedFallback: that option does not provide an unsandboxed runner for this path.

It fails closed without asserting a cause it cannot know, so the exception tells you where to look:

What happenedWhat is thrown
The spawn threw a typed sandbox exception (E_SANDBOX_POLICY_CONFLICT, E_SANDBOX_REFUSED, …)That exception, unchanged — your policy or lifecycle is the problem, not a missing binary
The child could not be spawned, or its output could not be read (an untyped throw)E_SANDBOX_DEPENDENCY_MISSING — the liveness failure the probe exists to detect
The child ran and exited non-zeroE_SANDBOX_FAILED, carrying the exit code and stderr

That last row is why the classification is not simply "dependency missing": issue #22's own symptom, bwrap: loopback: Failed RTM_NEWADDR: Operation not permitted, arrives as a non-zero exit and is a kernel permission problem. Telling an operator to install bubblewrap there would send them the wrong way entirely.

The probe runs after policy admission, not during preflight, so an unadmitted wider policy is never used to test a child. Construction is serialized behind an establishment promise: concurrent createSandbox() calls queue until the preceding establishment (including its probe) settles. This is a lifecycle change for every consumer, not only callers that opt into probing. dispose() joins the same queue, so a slow teardown can delay a concurrent construction; in return, teardown cannot reset the owner while a construction is being established.

Gates are the security control ​

Every tool has a gate, reads included. A read of .env is unrecoverable exfiltration, and search_files is a secret-discovery primitive. Gate the operation before path translation or I/O. The gate is a real suspension: a headless harness with no decider HANGS THE TURN — indefinitely, with no timeout of its own and nothing in the logs to say why.

There is deliberately no reference gate implementation. A shipped default gets adopted unread as “the safe config.” Here is a real decider: deny by default, ask a human/policy service, and return an explicit verdict.

ts
import type { DispatchContext } from '@nhtio/adk'

type Verdict = { approved: true } | { approved: false; note: string }
const gate = async (
  _ctx: DispatchContext,
  call: { tool: string; args: unknown },
): Promise<Verdict> => {
  const approved = await approvalService.ask(call.tool, call.args) // your decider
  return approved ? { approved: true } : { approved: false, note: 'operator declined' }
}

// The blanket example is intentionally frightening and never a default:
const blanketApprove = async () => {
  // You have just handed the model an unsandboxed shell. That was a choice. Own it.
}

A gate denial is narrated as E_SANDBOX_REFUSED; operational failures throw narrated E_SANDBOX_* errors so the model receives actionable text. A command that did run returns its artifact even for non-zero exit, timeout, violation, or post-spawn I/O failure: its output is the signal. Returning an error string would turn it into a handle on most adapters.

What the model actually reads: the prefix is not ours ​

Failures throw, where the usual ADK convention is to return an Error: … string. The reason is mechanical: a returned string is spooled, and a spooled inline: false result renders as an artifact HANDLE, so the model would have to spend a whole call querying an artifact to read one line of refusal. A throw becomes inline text on every adapter instead.

The consequence is that the model does not read your narration alone. Tool.executor() wraps a handler throw in E_TOOL_DOWNSTREAM_ERROR, whose message is a fixed sentence, and the adapter concatenates that with the immediate cause. So the model reads exactly:

text
The tool handler threw an error during execution. <narration>

That prefix is core behaviour and not configurable. Two rules follow, and both are easy to get wrong:

  • Throw E_SANDBOX_* as the DIRECT cause. Only the immediate cause's message is appended, so one extra wrapping layer drops your narration out of what the model sees entirely.
  • Never prefix the narrator template. 'Sandbox refused: %s' would read as a second prefix on top of the core's. The exception templates are bare '%s' for this reason.

A deployer who reads "the model sees the narration" and then meets that sentence in a transcript would reasonably conclude something is broken. Nothing is: that is the delivered form.

Platform and path-class rules ​

Do not say merely “the sandbox.” These are distinct matching rules:

PlatformPath classRule and consequence
macOSmandatory sensitive filesMatching is case-sensitive; .BASHRC is not the mandatory deny. A case-insensitive volume can still resolve a name the profile did not cover.
Linuxmandatory sensitive filesDangerous file matching is case-insensitive, so .BASHRC is denied.
Linux.git directory suffixesMatching is mixed: a .GIT directory matches, but the literal suffix test means sub/repo/.GIT/hooks/pre-commit is PERMITTED. Do not infer safety from the directory match.
macOS and Linuxordinary read/write policy listsLists are case-sensitive. Read deny/allow and write allow/deny therefore need the exact path spelling.
Windowsordinary read/write policy lists and path classesMatching folds case. Native SRT is unsupported here; use WSL2 so the Linux rules above are the rules you test.

Linux mandatory scanning also stops at mandatoryDenySearchDepth (default 3); macOS globs match at any depth. Linux root .git/hooks and .git/config mandatory entries exist only for a real .git directory, not a worktree .git file.

Adopting a foreign sandbox: network.disabled does not mean what it looks like ​

SandboxManager is a process-global singleton, and its second initialize() is a no-op — verified directly: initialise with denyRead: ['/tmp/AAA'], then again with ['/tmp/BBB'], and the derived rules still report /tmp/AAA. So an enforcer that always initialises would, on a process where something else enabled SRT first, report the policy you requested while a different one is in force.

srtEnforcer() therefore detects before it initialises. If sandboxing is already enabled it adopts that sandbox read-only: it does not call initialize(), it derives its snapshot from the LIVE manager, and it never reset()s a manager it did not create — a reset would tear down ACEs the host application depends on, so dispose() is a reported no-op there.

Adoption is still a degraded mode, and one field in effectivePolicy() will mislead an operator who reads it as authoritative:

  • network.disabled is this battery's field, not upstream's. SRT has no such flag. It records "this handle was constructed in disabled mode" — and an adopted sandbox was never constructed here at all.
  • Therefore disabled: true can only ever appear on a handle this battery created. On an adopted handle it is always false, and that false does not mean "the foreign sandbox restricts domains". It means only "its domain lists are being compared literally". Reading it the other way inverts the conclusion.

Two consequences follow. Both are benign, and both look alarming:

  • The domain-skip branch never fires under adoption. With no baseline disabled: true there is nothing to skip, so both domain axes are always compared — strictly MORE checking, not less.
  • A foreign mode flip is not reported as a MODE change — there is no upstream mode bit to read, so the only thing a comparison could ever see is an EFFECT on the lists.

A foreign mid-session widening IS detected, by a mechanism that differs per mode:

  • Adopted ⇒ effectivePolicy() re-derives from the live manager on every call. A foreign updateConfig() that widens the domain lists between two invocations is caught by the ordinary per-invocation drift check — verified end to end: a foreign ['*'] baseline widened to ['*','bar.com'] is reported as drift.
  • Owned ⇒ the snapshot is cached. This battery's updateConfig() is the only writer, so re-deriving would buy nothing but a getFsReadConfig() call per tool invocation (and on Linux a ripgrep scan).

What remains genuinely undetectable is why a foreign policy changed — there is no upstream mode bit, so a comparison only ever sees the effect on the lists. And detection is not prevention: SRT's proxies consult policy per request, so a widening affects an already-spawned child for its whole lifetime. You learn about it on the next invocation, not before the current one finishes.

A foreign allowedDomains: ['*'] is recorded as an ordinary allow-everything list, not as evidence of disabled mode, because that is what the config actually says. Note the honest cost: '*' is a glob, and two globs compare only when lexically identical, so ['*'] → ['*','bar.com'] is reported as drift even though it widens nothing semantically. That is the conservative bias the subset rules declare — a false "not a subset" costs a reconfiguration, a false "is a subset" costs containment.

Deployment consequence: this is a process-global capability wearing a per-handle API. One policy per process is the safe shape; multi-tenant agents with different policies want separate processes.

Per-call policy semantics under SRT ​

SandboxHandle.run({ policy }) accepts a per-invocation policy. Whether a per-call GRANT (something outside the process baseline) is honoured depends on the axis, and this was measured against a real seatbelt sandbox on macOS and a real bubblewrap sandbox on Linux (SRT 0.0.70, issue #50). The two platforms agreed on every step, so this states the behaviour of the SRT enforcer as a whole:

AxisPer-call grant honoured?ScopeMeasured behaviour
Filesystem write/readYesThat child onlySRT bakes the per-call allowWrite/allowRead into the wrapped command for that one invocation (a seatbelt profile on macOS, a bwrap --bind over a read-only root on Linux). It replaces the session baseline for that child — the per-call policy IS the child's complete filesystem policy; it does not union with the baseline.
Network domainsRefused, any difference—A per-call network section must EQUAL the session's (order-insensitive set equality of the domain lists plus the disabled flag). Any difference — a wider allow-list, a NARROWER one, an added/removed deny, a different disabled, a different reasons map — throws E_SANDBOX_NETWORK_GRANT_UNSUPPORTED before the child spawns. Nothing is silently ignored.

Concretely, measured on both hosts with a baseline allowWrite: [baselineW] and allowedDomains: ['example.com']:

text
baseline child  touch baselineW/ok        → exit 0
baseline child  touch perCallW/denied     → exit 1  macOS Operation not permitted / Linux Read-only file system
per-call child  touch perCallW/ok         → exit 0   (grant honoured)
per-call child  touch baselineW/denied    → exit 1   (REPLACEMENT, not union)
baseline child  touch perCallW/after      → exit 1   (no leak)
baseline child  curl example.com          → exit 0   (baseline domain reachable, HTTP 200)
baseline child  curl example.org          → exit 56  (host outside baseline blocked, even while another child is in flight)
per-call child  grant api.github.com      → E_SANDBOX_NETWORK_GRANT_UNSUPPORTED, child never spawns
per-call child  narrow to []              → E_SANDBOX_NETWORK_GRANT_UNSUPPORTED, child never spawns
per-call child  repeat [example.com]      → exit 0   (equal section accepted; only the fs axis is per call)

So the D12 pattern — a minimal baseline with per-call granted filesystem sources — is supported, with the caveat that each per-call policy IS the complete filesystem policy for that child: repeat the baseline writes you still want, and on Linux make sure the per-call write path already exists (SRT drops absent allow paths before the --bind). Per-call network policies are refused unless they equal the session's: neither a grant NOR a narrowing is consulted by the proxy, so an unequal section would make the child run under the session list while the caller believed otherwise. Repeat the session's network section (which a filesystem-only per-call policy does by construction) and network policy stays a session property.

Why the network refusal is correct — SRT cannot represent a per-child allow-list. This is not a decision ADK could have made differently without leaving the boundary unenforced. SRT 0.0.70 starts ONE mux proxy per process, at initialize(), and its filter closes over the module-global session config. From sandbox-manager.js (v0.0.70):

text
startMuxProxyServer:  httpProxyServer = createHttpProxyServer({
    filter: (port, host, _socket, encodedCommand) =>
        filterNetworkRequest(port, host, sandboxAskCallback, encodedCommand), ...
filterNetworkRequest: async function filterNetworkRequest(port, host, sandboxAskCallback, encodedCommand) {
    ...for (const allowedDomain of config.network.allowedDomains)   // ← the SESSION config, global

The per-call customConfig.network handed to wrapWithSandboxArgv is read only to decide WHETHER the child gets restricted (its allowedDomains presence makes hasNetworkConfig true), never for WHAT is allowed — enforcement always reads the session config. SandboxManager.updateConfig() can change that list live, but it changes it for every child in the process, with no per-child key anywhere in the API (wrapWithSandboxArgv(customConfig) has no way to reach the filter; there is no per-command proxy port, and an external proxy would still be one shared filter). So a per-call NETWORK section is unenforceable in either direction: ADK refuses every difference (E_SANDBOX_NETWORK_GRANT_UNSUPPORTED) rather than spawning a child whose effective policy is not the one it was handed, and the error message names the fix — establish a session whose policy includes the domains. Concurrent children are safe by that same design: A under the baseline reaches an allowed domain while B's diverging attempt is refused at the same moment, and neither call mutates the session the other is using.

The per-child route was investigated and MEASURED dead (round 3). SRT does expose two embedder hooks that see per-connection data, and both were tested against the real proxy rather than dismissed by reading:

  • SandboxManager.initialize(config, sandboxAskCallback) — the ask callback fires for a host NOT on the session allow-list, but receives only {host, port}; the child identity the proxy actually computed (encodedCommand, from the proxy username) is passed to the FILTER but never forwarded to the callback. Measured: the callback's argument keys are exactly ["host","port"]. Worse, ADK's mapPolicy always sets strictAllowlist: true, and SRT then never consults the callback at all (measured: 0 invocations, HTTP 403) — the allow-list is deterministic. And even if it were consulted, the identity is encodeSandboxedCommand(command) = base64 of the command's first 100 characters: two different commands sharing a 100-char prefix are the SAME identity (measured: same_first_100: true, encoded_equal: true), so a lookup keyed on it would collide.
  • network.filterRequest — receives the parsed request (including proxy-authorization) and runs only AFTER the session allow-list has already admitted the host, and is deny-only. Measured: a callback returning deny 403s an otherwise-allowed host; it cannot GRANT a host the session list excludes.

Forgery seals it: the proxy username is client-controlled inside the sandbox, and the proxy token is in that child's own environment. Measured: a child presenting a DIFFERENT child's username (srt.<victim>) with the valid token is accepted, and the denial is filed under the victim's identity (attributedToVictim: true). SRT's own comment at sandbox-utils.js:703 says the suffix "can only misattribute a denial", which is exactly right — it cannot authenticate — so it cannot be the basis for a per-child GRANT either. With identity unavailable, forgeable, and collidable, and the allow-list consulted before any embedder callback, there is no route to a per-child network allow-list that is both unforgeable AND collision-free — and any fix would require patching SRT or mutating the process-global per spawn, both excluded. The typed refusal stands.

Measured on macOS/seatbelt and Linux/bwrap (SRT 0.0.70), with identical outcomes. The filesystem replacement and the proxy-filter behaviour are upstream code paths shared by both backends; both were run against real sandboxes, so the behaviour above is verified for each rather than inferred for one.

handle.narrow() is a separate, unrelated mechanism: it throws E_SANDBOX_NARROWING_UNSUPPORTED with the SRT enforcer because SRT exposes no in-place narrowing, and that is unchanged by the above. Use a per-call run({ policy }) for one-child grants instead.

What this is not ​

Not a guarantee that arbitrary code is safe

This is a policy boundary, and a policy boundary is only as good as the policy plus the code enforcing it. It does not make the model trustworthy, the filesystem race-free, or JavaScript magically native. The file tools run in-process library code, so a denial there holds because the evaluator checks it — TOCTOU between the check and the open() is a real residual, and spawning a helper process would not fix it, because the helper races identically. Only the shell and search paths get kernel enforcement.

Not verified by a green pipeline

The suites that prove the OS boundary actually works are gated behind TEST_SANDBOX_LIVE, and they skipIf themselves into a passing report when it is unset. So a green CI run is not weak evidence about enforcement — it is no evidence. Run them yourself, on both macOS and Linux, and record the SRT version you ran against:

sh
TEST_SANDBOX_LIVE=1 pnpm vitest run --project node tests/live/

There is now one CI job that does set the flag, in docker-in-docker. It is non-blocking and proves less than it sounds: that runner's daemon forbids unprivileged user namespaces, so bubblewrap cannot start there at all. It exercises the path and reports what it found; it does not certify enforcement. Your own run, on your own kernel, is still the evidence.

Not a stable dependency contract

Node's enforcement rides on @anthropic-ai/sandbox-runtime, a Beta Research Preview whose own notes say the APIs "may evolve". Every enumerated list in this battery is a snapshot of one version's source — not its README, which is wrong about ripgrep on macOS (it documents rg as a dependency there; the code gates that check inside if (platform === 'linux')). Treat an SRT version bump as a code change requiring a source re-read, not a routine dependency bump.

Opt-in, like every other battery here

Nothing in the ADK requires the sandbox. It exists for the case where you want to hand a model real execution and keep a boundary you can reason about — reach for it then, ignore it otherwise.

And what was deliberately not built ​

  • No reference gate. A safe-looking default gets adopted unread as "the safe config", carrying an arbitrary assumption about which subtrees are boring. The signature and worked examples ship; the decider is yours.
  • No browser no-op shim for the OS layer. A shim that degrades to nothing produces code that reads as sandboxed and enforces nothing — worse than a build error. In a browser, SES is the boundary.
  • No read_file/edit_file. open_file queries, stage_file + save_media mutate explicitly. The split is the point: mutation should be impossible to do by accident.
  • No dedicated git tool. run_shell_command is the git tool — the real binary, its real output, no wrapper to keep in sync with a fast-moving CLI.
  • No byte caps on results. Truncating a shell's output can cut the one line that explained the failure, and killing a child mid-write can leave the real system half-changed. Output is spooled whole into a queryable artifact; bound it with timeout_seconds and policy, not silent truncation.

Troubleshooting: what this failure means ​

SymptomMeaning and fix
dubious ownershipgitSafeDirectories was not threaded into the policy/enforcer. Add the workspace root.
E_READER_NOT_DESCRIBABLE on encode()A staged Media was never saved. Call save_media before persisting history.
"[object Object]"A handler returned a value the adapter could not wrap. Return the supported handler shape, not an arbitrary object.
Hung turnA gate suspended and there was no headless gate decider. Supply one or fail closed.
Artifact handle where text was expectedUniform inline: false behaviour: query the artifact with artifact_grep/other artifact tools.

Browser deployment ​

CORS gates response readability, not request emission. no-cors fetches and image beacons can still send secrets. The embedder MUST provide CSP, especially connect-src, plus img-src and default-src:

http
Content-Security-Policy: default-src 'self'; connect-src 'self' https://api.example.test; img-src 'self' data:;

Keep real secrets outside the sandbox root. File tools have evaluator-level checks and retain TOCTOU and implementation-bug residuals; shell and search paths receive OS enforcement.

Where to go next ​

  • run_shell_command — the streaming shell and the git tool: how stdout and stderr are merged, why a non-zero exit is a final line rather than an error, and the measured state of violation reporting on macOS.
  • Workspace tools — the read/mutate split (open_file queries, stage_file + save_media change), per-entry authorisation in list_directory, and the search pair.
  • Media executor — wrapping an existing BinaryExecutor for media pipelines, the audit-only bypass, and the retention semantics it inherits.
  • evaluate_javascript — the SES boundary: guest-side lockdown(), enumerated capabilities, and which availability guarantees hold in which environment.
  • Assembly → Sandbox batteries — the full option and exception surface, side by side with the other battery domains.
  • Isolation battery — the related-but-different substrate: it isolates trusted code for stability, where this battery confines untrusted code for authority. Compose them for the full stack.