10 · Local SDK · CLI Test Framework
Write repeatable tests for agent outputs, tool call sequences, cost budgets, latency bounds, and consistency across runs. A local TypeScript SDK and CLI with deterministic mock runner and CI-friendly report artifacts — no hosted runner required.
Agent outputs change across prompts, models, tool configurations, and runs. Teams catch regressions through manual testing, production incidents, or not at all. There is no standard way to assert that a specific tool was called with a specific input, that the response stays within a cost budget, that two consecutive runs return the same output, or that a refusal pattern holds after a model upgrade.
Agent Test Harness brings unit and integration testing discipline to agent pipelines. Write JSON fixtures that describe prompts, mock tool responses, scripted tool call sequences, and assertions. Run the fixture suite with a single CLI command. Get a pass/fail exit code and JSON, Markdown, or JUnit XML reports that CI pipelines can consume without any hosted runner or account setup.
The harness ships with a deterministic mock agent runner so test suites execute
predictably without live LLM calls. When you are ready to test against a real agent,
supply your own AgentRunner callback via the SDK.
Four steps: write a fixture, run the CLI, read the report, and maintain history. All execution is local — no provider calls, no accounts.
Describe each test case with a prompt, optional
mockTools, scriptedToolCalls, and an
assertions array. Use repeat to run the
same case multiple times and assert consistency. Tag cases as
"smoke", "regression", or "safety"
to filter runs by purpose.
agent-test suite.json --reporter json --reporter junit loads
your fixture, executes it against the deterministic mock runner, evaluates
all assertions, and prints a formatted pass/fail summary. Exit code 1 if any
test fails — wires directly into shell scripts and CI steps without extra tooling.
The --out directory receives latest.json,
latest.md, and/or latest.xml depending on the reporters
selected. history.json is always appended with a summary row per run,
giving you a local trend record over time.
Import runSuite from the SDK and pass your own
AgentRunner callback to test against a real agent, a staging
environment, or any async function that returns { output, toolCalls, meta }.
custom assertions with JS functions are available in SDK mode only.
A fixture file is a JSON object with a tests array.
Each test case specifies a prompt, optional mock tools and scripted tool calls,
and an assertions array. The fixture loader validates the schema before execution
and rejects custom assertions with a clear error.
id (required), description, prompt (required),
systemPrompt?, mockTools?, scriptedToolCalls?,
assertions (required), repeat?, tags?, skip?.
·
custom assertion type is rejected in JSON fixtures
— use the SDK with a JS callback for custom assertion logic.
Thirteen assertion types across four groups. Twelve are available in JSON fixtures;
custom requires a JavaScript function and is SDK-only.
| Type | Group | What it asserts | Value format |
|---|---|---|---|
contains |
Output | Agent output contains a substring (case-insensitive) | String to search for |
not_contains |
Output | Agent output does not contain a substring (case-insensitive) | String that must be absent |
matches |
Output | Agent output matches a regular expression (case-insensitive) | Regex pattern string |
json_field |
Output | Parsed JSON output has a field at a dot-path equal to a given value | value + field dot-path |
tool_called |
Tool | A named tool appears at least once in the tool call trace | Tool name |
tool_not_called |
Tool | A named tool does not appear in the trace | Tool name |
tool_input |
Tool | A tool was called with a specific input field value | toolName.field.path=expectedValue |
tool_output |
Tool | A tool returned a specific output field value | toolName.field.path=expectedValue |
tool_sequence |
Tool | Tools were called in a specific exact order (complete sequence match) | Comma-separated tool names, e.g. search,format,respond |
tool_call_count |
Tool | A tool was called an exact number of times | toolName=N |
cost_under |
Budget | Actual costUsd in TestMeta is below a threshold |
USD threshold as a number, e.g. 0.01 |
latency_under |
Budget | Measured durationMs is below a threshold in milliseconds |
Milliseconds threshold as a number, e.g. 3000 |
custom |
SDK only | User-provided check(output, toolCalls, meta) => boolean function |
Requires a JS function via assertion.check — rejected in JSON fixtures |
The built-in mock agent runner executes scripted tool calls deterministically, records full tool traces, and constructs output from tool responses — no live LLM required. This makes fixture suites fast, stable, and safe to run in CI.
repeat
Three reporters write to an --out directory (default: .agent-test-results/).
history.json is always appended regardless of which reporters are selected.
Use --reporter junit for JUnit XML that CI servers such as GitHub Actions,
GitLab CI, and Jenkins can parse to show per-test results inline.
Full RunSuiteResult — all test results, assertion outcomes, tool call traces, output strings, and per-test metadata. Machine-readable for downstream tooling.
Summary table with test ID, pass/fail/skipped status, and per-assertion outcomes. Ready to commit to a repo or paste into a PR description.
JUnit-compatible XML output. CI-friendly — GitHub Actions, GitLab CI, and Jenkins parse this format to display per-test results in pull request checks. Exit code 1 on failure.
Append-only local history. Each run appends a summary row: runId, startedAt, total, passed, failed, skipped, durationMs. Local trend record with no cloud dependency.
Package source is published as @certaworks/agent-test-harness (v0.1.0) on npm as shown below.
The CLI binary is agent-test.
--tag flag for targeted runshistory.jsonAgentRunner interface for testing against real agents or staging environments--bail flag to stop after the first failing case@certaworks/agent-test-harnessAgent Test Harness ships as a local SDK + CLI. No hosted endpoint required to get started.
Once built locally, run `agent-test` to execute repeatable unit and integration tests against agent outputs, tool traces, and cost budgets.
View on npm →npm install -g @certaworks/agent-test-harness
npx -y @certaworks/agent-test-harness --help
Early Access
Get early access to Agent Test Harness — repeatable local tests for outputs, traces, cost budgets, and consistency. Write fixtures once, run them in CI, and catch regressions before production does.