Reusable results
Successful direct JSON and text results larger than 10KB are stored outside the model conversation. The first response contains the originalresultId,
capabilities, and the first contiguous text page in content. pagination
contains page, totalPages, pageSizeBytes, totalBytes, and hasMore.
Follow the returned next call unchanged; the last page has next: null.
The model can use the read-only results tool with list, read, grep, or
jq. Direct read uses numbered pages; legacy offset/maxBytes calls still
return byte pages with nextOffset. Do not mix the two pagination modes.
Small jq responses retain {values, truncated}. Oversized responses, or explicit
page/pageSizeBytes requests, page the serialized {values, truncated} response.
The query expression and optional stream limit stay in every continuation, using
the same original result ID. Queries execute again on each page, within existing
time/memory/output limits; no derived result is stored. A query’s stream-limit
truncated flag is separate from page availability and survives reconstruction.
Use deterministic expressions when paging: the stored source is immutable, but
time-dependent expressions such as now can change when the query executes again.
Pages are UTF-8-safe text chunks, not necessarily valid JSON individually.
Concatenate content before parsing. The default requested size is 16,384 bytes;
JSON escaping and envelope overhead can reduce it. Actual boundaries determine
totalPages, and next preserves the effective size. Result responses remain
within the configured 20KB wire budget. Exceptionally small budgets or large query
metadata can still fail explicitly when the envelope itself cannot fit.
Structured command results also support grep on their searchable output text;
list returns reference metadata only.
References are immutable and scoped to one session and agent. Reading one never
replays its tool or side effect. Refresh the original tool only when current
external state is required or an intervening mutation may have invalidated the
stored snapshot.
Provider-native and binary results that use native model-output conversion stay
on their direct delivery path and are not replaced with reusable handles.
Code Mode
Every normal agent run receivescode_exec automatically. It executes a
TypeScript function body in an isolated QuickJS-WASM heap so the model can use
ordinary code for loops, filtering, joins, branching, batching, aggregation,
and bounded parallel tool calls.
.agentuse configuration. Eligible JSON
tools are normally hidden from the top-level tool list and called through
code_exec, including when the program needs only one tool call. Each nested
call uses the same input object as the direct tool and passes through the
runtime dispatcher, including schema validation, plugin policy, mock isolation,
effect WAL, and session tracing.
Before any nested tool starts, the runtime strictly type-checks the program
against declarations generated from the effective tool schemas. AgentUse-owned
tools can publish trusted output contracts, including the built-in store and Bash tools, so
misspelled inputs, unhandled result variants, and invalid result paths fail
before an effect occurs. External provider and MCP output schemas remain
untrusted and appear as -> ?. Inspect an unknown result first, then narrow it
with runtime checks before using its fields in a later program. Completed
JSON-serializable nested calls use the same immutable, same-session result
references, including results too large to read directly into Code Mode:
nextOffset until it is null. Pages preserve UTF-8 characters
and share the direct results tool’s response limits. JSON pages need not parse
individually. The one-argument results.read(resultId) still returns the original
bounded value. Both forms count toward Code Mode’s result-read budget.
Numbered pages are also available through results.page:
next back unchanged; it is null on the last page. When numbered page fields are present they take precedence over the legacy offset and maxBytes options, which are still supported.
Bash returns { output, stdout?, stderr?, metadata? }. Completed processes have
separate stdout and stderr strings, including empty strings for empty streams.
Parse stdout for JSON or other structured command output after checking
metadata.exitCode, timedOut, aborted, and truncated. The compatible output
field remains combined display text with stderr and runtime diagnostics; it is
not a machine-readable stdout field. Refusals, startup errors, and historical
stored results may lack stream fields. Refusal errors remain JSON text in
output; there is no value field. Search declarations retain
parameter descriptions and small integer ranges, including context_lines
from 0 to 5 for filesystem search.
Preflight reports up to eight diagnostics, bounded to 6,000 characters, with
source locations, TypeScript error codes, and relevant signatures or repair
hints. Annotate helper parameters, such as function summarize(r: unknown),
and narrow unknown values before reading their fields. A failed preflight
executes no nested tools.
Use these references for immediate retries and continued analysis instead of
repeating the same tool call. Refresh the tool when current external state is
required or an intervening mutation may have invalidated the earlier result.
References never cross session or subagent boundaries and do not replay tools
or side effects. results.grep() performs bounded, case-insensitive literal
text search by default. Patterns are not regular expressions and characters
such as $ should not be escaped. results.jq() also accepts the legacy
results.jq(resultId, expression, { limit }) form. It uses real jq syntax and
returns jq’s ordered output stream as { values, truncated }. AgentUse runs
bundled jq 1.8.2 WebAssembly in a
resource-limited worker, without a shell, host environment access, jq modules,
or a machine-level jq installation.
Code Mode has no direct filesystem, network, environment, process, package,
shell, or import access. It can invoke only capabilities exposed through the
host tool bridge. This includes Bash commands from the agent’s effective
auto-run allowlist; the same command validator, path policy, safe environment,
plugin policy, mocks, cancellation, and effect journal apply on both paths.
Commands matching tools.bash.gated are rejected in Code Mode and must use the
direct Bash tool after human approval. Approval gates, outcome tools, internal
creator tools, subagents, and any tool marked as approval-required remain
direct-only. Binary media results also stay on the direct path because the guest
bridge is JSON-only. Programs are bounded by source, memory, stack, computation,
output, nested-call, and concurrency limits. Time awaiting Bash is governed by
the Bash timeout and does not consume the guest computation timeout. The 128 nested
tool-call limit applies to each code_exec program, including failed attempts.
A later program receives a fresh allowance; continue unfinished work without
repeating completed effects. Active nested calls (up to 8) and guest memory
(32 MiB) remain shared across programs in the same run. The agent run retains
its own execution timeout and step limits.
Use Code Mode for eligible JSON tools, including a single call, and for an
auto-run Bash command whose output feeds parsing, filtering, branching,
batching, or another tool. Bash remains separately visible for gated commands
and simple standalone process execution. Call other tools directly when they
require suspension, approval, subagent delegation, binary or provider-native
result delivery, or outcome submission. A transport-sensitive tool can appear
on both paths: use its Code Mode form for JSON/text composition and its direct
form only for the special result delivery. Code Mode does not make several
store calls atomic; use store_claim or store_update_if for concurrent state
transitions.
Controlled comparison and rollback
Code Mode is part of the runtime, not.agentuse syntax. To compare the same
agent with and without deterministic tool composition, run it once normally
and once with the runtime-only kill switch:
--no-code-mode applies to the whole run tree, including delegated agents.
Use AGENTUSE_CODE_MODE=0 for test harnesses or to disable Code Mode for every
run launched by a serve daemon. agentuse serve --no-code-mode is the CLI
equivalent for one daemon process. These switches are for evaluation,
debugging, and emergency rollback, not permanent per-agent configuration.
Path Matching Behavior
Filesystem paths and BashallowedPaths use containment-based path matching by default:
Rule: If the path contains glob characters (
*, ?, [), it uses glob matching. Otherwise, it uses containment (path = path/**).
Filesystem Tool
Controls access for Read, Write, and Edit operations. Granting a permission exposes the matching tool to the agent:read → filesystem_read, write → filesystem_write (and filesystem_edit), edit → filesystem_edit.
Configuration
Fields
Permission Model
The capability hierarchy isread < edit < write:
read, read file contents.edit, replace strings inside an existing file (cannot create new files or overwrite a file wholesale).write, create or overwrite any file. Because this is strictly stronger thanedit, grantingwritealso grantsedit(the agent gets both the write and edit tools).
[read, write] is all most agents need, the agent can read, do targeted edits, and do full writes. List edit on its own only when you want the narrower grant: modify existing files but never create or clobber them.
Prefer
[read, write] over [read, write, edit], edit is redundant alongside write. Use [read, edit] deliberately when an agent should tweak existing files without the ability to create or overwrite them.Reading images and PDFs
filesystem_read returns text files with line numbers, but it also reads images (PNG, JPEG, GIF, WebP) and PDFs, handing the actual image or document to the model so an agent can reason over a chart, screenshot, or scanned page directly. File type is detected by content (magic bytes), not by extension.
This requires a model whose input modalities include image/pdf (Claude, GPT-4o, Gemini, and most modern models). On a text-only model, an image/PDF read returns a clear error instead of breaking the run. The same path allowlist applies as for text reads. Size caps: images up to ~5MB, PDFs up to ~32MB; offset/limit are ignored for media files.
Path Variables
Examples
Read Limits
read_file returns up to 2000 lines per call by default (AGENTUSE_TOOL_MAX_LINES); pass an explicit limit to read more, or an offset to page through a larger file. Individual lines longer than 2000 characters (AGENTUSE_TOOL_MAX_LINE_LENGTH) are truncated with a ... (truncated) suffix. See Tool Output environment variables to tune these.
Edit Operations
The edit tool replaces exact strings rather than rewriting whole files. It uses fuzzy matching to tolerate minor whitespace, indentation, and line-ending differences. Prefer editing over full writes on large files, rewriting a large file regenerates its entire contents as output tokens, which is slow and can exhaust a run’s time budget. A single edit replaces one string:
To make several changes in one call, pass an
edits array instead of the top-level old_string/new_string:
Batched edits apply sequentially (each to the result of the previous) and are all-or-nothing: if any edit fails to match, the file is left unchanged. Provide either the single form or the
edits array, not both.
Artifact Tools
Artifact tools let an agent save substantial deliverables for the user to view in the session UI without granting broad filesystem write access.Configuration
artifact_save writes under .agentuse/artifacts/ by default, records metadata in a manifest, and returns a viewable session URL when a session is active. Markdown artifacts can include title and tags, which are merged into frontmatter. To read an artifact’s content later, use filesystem_read on the returned path.
Metrics Tool
The metrics tool lets an agent record business-metric facts (counts and amounts) about work it just completed, e.g. “chased 4 invoices totaling $11,200”. Records land in the reserved sharedmetrics store and roll up on the serve Home page and over the JSON API, so the numbers a dashboard shows are deterministic tool writes, never model-computed sums.
Configuration
Fields
At least one of
value or count is required.
Idempotency
Records are upserted keyed on(sessionId, metric): a retried or resumed run overwrites its own earlier record instead of double-counting, and recording the same metric twice in one run is last-write-wins. This is what makes the numbers trustworthy enough to display. Runs without a session id (rare) cannot be deduplicated and always create a new record.
The runtime stamps sessionId and the agent id onto every record; the model cannot spoof provenance. Each record is a regular store item (type: "metric", tagged with the metric name) in .agentuse/store/metrics/, browsable at /stores/metrics and readable programmatically via GET /api/stores/metrics. In mock mode the write is not skipped, it is isolated: the metrics store is re-rooted along with every other store under .agentuse/store-mock/, so a test run’s records are inspectable there and never enter the real metrics store the dashboards read. See the store guide’s note on mock runs.
Bash Tool
Controls which shell commands can be executed and in which directories.Configuration
Fields
When no config timeout is set, the model may pass a per-call
timeout: a duration string, or a bare number meaning milliseconds (kept for model familiarity). A bare per-call number under 1000 is rejected with a corrective error (it is always a seconds-vs-milliseconds mixup; a real sub-second timeout must be written as "500ms").
Command Patterns
Bothcommands and gated support ${root}, ${agentDir}, and ${tmpDir}.
These resolve from the runtime path context before matching, including approval
enforcement and gated mocks. For example, python3 ${agentDir}/check.py *
matches the resolved script path. This does not expand shell input or environment
variables. ${agentDir} requires an agent loaded from a file. Replacement paths
containing whitespace or shell/wildcard metacharacters fail configuration rather
than changing pattern meaning; use a project-relative command for those paths.
Commands use simple wildcard matching:
allowedPaths Behavior
TheallowedPaths field uses containment - a path grants access to all files and subdirectories within it:
Project root is always accessible for bash commands. Use
allowedPaths for directories outside the project.Examples
Output Limits
The Bash tool retains at most 4MB (AGENTUSE_BASH_CAPTURE_BYTES) as canonical output for reusable results. When command output exceeds that cap, AgentUse keeps a head + tail slice (40% head / 60% tail by default), saves the complete stream as a session artifact, and marks the result as truncated. A truncated capture never receives a reusable resultId because that ID must always represent complete data.
Independently, any direct JSON or text result over 10KB (AGENTUSE_TOOL_INLINE_RESULT_BYTES) becomes a reusable result handle before it reaches the model. This smaller limit prevents a large result from being re-sent on every subsequent model step.
Queries through the results tool can return up to 20KB (AGENTUSE_RESULT_QUERY_BYTES) inline. They never create recursive result handles.
Bash stream fields use the same bounded captures as output. Check
metadata.truncated before parsing them. Keeping output for compatibility
means the serialized result also contains a combined copy of the streams, so it
can reach the inline-result limit sooner. Stored results preserve the separate
fields: use a query such as .stdout | fromjson | .items to inspect JSON stdout
without including stderr diagnostics.
Provider-native and binary results with native model-output conversion are excluded from this inline-result limit.
Because every tool result is re-sent to the model on each subsequent step, a single large output inflates input-token usage for the rest of the run. Prefer commands that emit only what you need, for example git diff --stat instead of a full git diff over high-churn files. See Tool Output environment variables to tune the caps.
Sandbox Tool
When asandbox is configured in the agent frontmatter, the sandbox__exec tool is injected for running commands inside the Docker container. File I/O is handled by the filesystem tool, no separate sandbox file tools are needed.
The sandbox tool is only available when
sandbox is configured. See the Sandbox guide for setup instructions.sandbox__exec
Execute a shell command inside the Docker container.
Returns
stdout, stderr, and exitCode.
Mount Mode
Each filesystem path is mounted at its real host path with per-path mode derived from permissions:- Read-only, No
writeoreditpermissions for that path - Read-write,
writeoreditpermissions granted for that path
/workspace/ alias). Changes made by the filesystem tool on the host are visible inside the container via the bind mount.
Run Outcome Tool
Thereport_outcome tool is always available, no configuration needed. Every agent ends its run by calling it once with one of three statuses.
report_outcome
What each status does:
complete: the objective was achieved. The call is the run’s answer and ends the run. The session iscompleted.idle: a successful check found no action due: nothing eligible, no alert condition met, or the intended work was already handled before this run. Routine logs, checkpoints, watermarks, audit notes, and monitoring-state updates do not turn an idle check into delivered work; exclude them fromartifacts, which must be[]. A requested report, analysis, or other substantive deliverable actually produced by this run iscomplete, even when it recommends no action. An idle run ends successfully and its durable session status remainscompleted; the run JSON carriesresult.idle: true, and the sessions list exposes an Idle badge and filter separately from Done. A stretch of idle runs can help reveal changes in an agent’s supply of eligible work.incomplete: a required outcome was not delivered because something failed, including due work the agent found but could not move. The run keeps going only for bookkeeping. When it ends, the session is persisted aserrorwith codeINCOMPLETE. The blocker kind is stored inerror.cause, witherror.subjectanderror.causeSource(runtime,approval,agent, orinferred) saying how it was established:- A failed tool call about the same thing (
command not found,No module named,Cannot find module) overrides the agent’s chosen kind - A sub-agent’s proven blocker passes to its parent
waiting_on_humanandrejected_by_humanshow asincompletein the session list, exit 0, and settle the Slack run card as “run incomplete” without a failure push- Any other blocker, or none, stays an
error: thefailureevent fires and the exit code is 1 agentuse sessions showdisplays the headline under the error section
- A failed tool call about the same thing (
incomplete when unsure between idle and incomplete. Agent files should not restate this mechanic; only add domain judgment about what counts as blocked, done, or genuinely empty for that agent.
Sessions recorded before
report_outcome used two tools, report_complete and report_incomplete. They still display and reconcile the same way, and a session suspended before the change resumes with those tools. An old report_complete never reads as idle, since it cannot say whether the run did anything.Security Considerations
Filesystem Tool
- Sensitive files blocked:
.env,.env.local, etc. are blocked by default - Symlink resolution: Symlinks are resolved to prevent escape attacks
- Path traversal prevention:
../sequences are normalized and validated
Bash Tool
- Command allowlist: Only explicitly allowed commands can run
- Directory restrictions: Commands can only access project root and
allowedPaths - Environment sanitization: Dangerous environment variables are cleared
- Timeout enforcement: Commands are killed after timeout