Choose what to test
Run
agentuse test --help to see both choices. Put options after workflow
or result. Both commands use real model calls and can incur provider costs.
Compare a result from a past job
<stateRoot>/.agentuse/test-evidence/. Later tests of that
source reuse the selection, even after you edit the instructions or change the
writer model. --selector-model <model> chooses the extractor on first use;
otherwise it uses the agent model. It does not replace already saved evidence.
The writer receives current agent instructions, explicit preloaded skills,
applicable learnings, the original additional prompt, and the selected evidence.
Explicit read-only file grants are loaded as current references and their hashes
appear in the report. Directory and wildcard grants are not expanded into fresh
references. A changed reference can change the result even with fixed evidence.
Relative dates use the original session date. The original output and human
feedback stay out of the writer and judge context and appear only in the report.
This tests the deliverable, not research, delegation, tool behavior, approval,
or publishing. No workflow tool or Code Mode executes. If required evidence or
capabilities are missing, the writer should report incomplete. Model-assisted
selection can make mistakes; inspect the saved audit before relying on a test.
Sources must be real sessions that have finished or reached approval and contain
recorded tool evidence. Unsupported/media-only evidence, invalid selections,
and inputs exceeding the context budget fail explicitly without truncation.
--judge <local-agent> supplies quality criteria from a separate agent. Its
current explicit skills, learnings and read-only file references are available,
but it has no tools. It evaluates the new result against the task and evidence;
a pass is not a claim of improvement over the original. Use --model to change
the writer model; the judge uses its own configured model.
--json emits the structured report. The printed report path contains the
comparison, evidence provenance, current reference hashes, and any judgment.
Each invocation creates a test session excluded from operational views; it
cannot be resumed as a live run. --timeout bounds selection, generation and
judging together. -C and --env-file work as on run.
Test workflow steps
An agent that sends email, pushes commits, posts to Slack, or deletes rows is hard to debug by running it. The first real run is also the first irreversible one, and reading the agent file only tells you what you wrote, not what the model does with it.agentuse test workflow closes that gap. It runs the agent all the way through, on its
own model, with its own instructions and its own tools, but fabricates the side
effects instead of performing them. You get a full session log of a complete run
with simulated external tool responses and isolated store writes by default.
Run a workflow test
--quiet, --debug, --no-tty,
--compact, --timeout, -C, --env-file, -m/--model, --json) plus the workflow options below.
A Mock Model Is Required
Mock mode fires an LLM call for every tool result, so it needs a model it can reach:AGENTUSE_MOCK_MODEL in your shell,
~/.agentuse/.env, or the env block of
~/.agentuse/config.json.
The flag wins over all of them.
Choose which tools to simulate
Newtest workflow invocations default to all: external tool results are
simulated. --scope gated explicitly simulates only matching gated bash
commands; other tools execute for real. Approval decisions are simulated in both
modes, and stores are isolated.
tools.bash.gated patterns run under gated scope gets a
warning: nothing is mocked, everything runs real, with gates still auto-resolved.
Approval Gates, Unattended
A test run must finish without a human, so the gate resolves deterministically instead of suspending. No model plays the reviewer.
The comment applies to the first gate only; the re-gate that follows is
approved, so the run exercises the revision path and still finishes. Repeating
it on every gate would loop an obedient agent until it gave up.
What Is Isolated
- Stores. Every store, including the reserved
metricsstore behind the dashboards, is re-rooted at<project-root>/.agentuse/store-mock/<timestamp>-<pid>/. Reads are seeded by copying the real store, so the agent still sees realistic data; writes land in the scratch copy, which is kept for inspection and swept after 7 days. See the store guide. - Operational views. Mock sessions are excluded by default from the serve home dashboard, agent health, and the sessions list, and they never fire push notifications, so a test loop does not buzz anyone’s phone or distort the picture of production.
The Closed Loop
Testing is meant to be iterated, not performed once:- Edit the agent file.
agentuse test workflow my-agent.agentuse.- Read the session log: what the agent actually did, in order, with what arguments.
- Repeat.
only mock (API: ?mock=only). A mock session’s detail page is always
reachable by id whatever the filter says, and the CLI marks these runs with a
trailing · mock (hide them with agentuse sessions list --no-mock).
For a tools.bash.gated flow, the log shows the whole gate sequence: the
await_human call with its changes[], the deterministic decision, and the
fabricated command result. In the effect WAL, a fabricated gated command has a
lease-approved entry but no bash-spawn entry, which is the audit signal
that it never executed.
What a Passing Test Does and Doesn’t Prove
- Mock outputs are non-deterministic. They are LLM-generated, so two runs can hand the agent different plausible results. That is useful for probing how the agent copes with variation, and useless as a regression assertion.
- Mocked approval decisions are deterministic. The branch you asked for is the branch you get, every time.
- MCP servers still connect at startup for tool discovery. Under full mock, their tool calls are simulated; gated scope leaves MCP calls live.
- Gate enforcement is still real. Mock mode does not relax the permission model: a gated command the agent tries to run without a lease is still denied, exactly as in production.
Lower-Level Plumbing
agentuse test workflow is the front door. The same machinery is exposed on run for
scripts and unusual cases:
run --mock keeps the approval gate real: await_human still
suspends, which is useful when you specifically want to verify the agent pauses.
Add --mock-approval for the unattended behaviour agentuse test workflow defaults to.
See CLI Commands for the full flag
reference and Environment Variables
for the AGENTUSE_MOCK_* equivalents.
Related
Approval Gates
Gate configuration, reviewer flows, and lease enforcement
Session Logs
Reading what a run actually did
Store
How mock runs isolate persistent state
Creating Agents
Writing the agent you are about to test
Compatibility and advanced debugging
agentuse test <agent> retains the older adaptive scope: gated when the agent
has gated bash patterns, otherwise all. Existing scripts keep that behavior.
The new test workflow command defaults to all even when gated patterns exist.
For an agent file literally named workflow or result, use an explicit path
such as ./workflow with the compatibility form.
agentuse test <agent> --replay <session-id> remains an advanced strict
recorded-call debugger. A changed tool argument without a matching recording
stops it. It is not the same as test result, which supplies evidence directly.
See Replay recorded inputs.