10 · Local SDK · CLI Test Framework

Agent Test
Harness

Write repeatable tests for agent outputs, tool call sequences, cost budgets, latency bounds, and consistency across runs. A local TypeScript SDK and CLI with deterministic mock runner and CI-friendly report artifacts — no hosted runner required.

Local slice · v0.1.0 · Node ≥ 18 · 33/33 tests passing
← Product 09: Prompt Archaeology Next: Product 11 · Wave 2 Heartbeat & Continuity →

Agent behavior drifts without repeatable tests

Agent outputs change across prompts, models, tool configurations, and runs. Teams catch regressions through manual testing, production incidents, or not at all. There is no standard way to assert that a specific tool was called with a specific input, that the response stays within a cost budget, that two consecutive runs return the same output, or that a refusal pattern holds after a model upgrade.

Agent Test Harness brings unit and integration testing discipline to agent pipelines. Write JSON fixtures that describe prompts, mock tool responses, scripted tool call sequences, and assertions. Run the fixture suite with a single CLI command. Get a pass/fail exit code and JSON, Markdown, or JUnit XML reports that CI pipelines can consume without any hosted runner or account setup.

The harness ships with a deterministic mock agent runner so test suites execute predictably without live LLM calls. When you are ready to test against a real agent, supply your own AgentRunner callback via the SDK.

How it works

Four steps: write a fixture, run the CLI, read the report, and maintain history. All execution is local — no provider calls, no accounts.

Write a JSON Fixture

Describe each test case with a prompt, optional mockTools, scriptedToolCalls, and an assertions array. Use repeat to run the same case multiple times and assert consistency. Tag cases as "smoke", "regression", or "safety" to filter runs by purpose.

Run the CLI

agent-test suite.json --reporter json --reporter junit loads your fixture, executes it against the deterministic mock runner, evaluates all assertions, and prints a formatted pass/fail summary. Exit code 1 if any test fails — wires directly into shell scripts and CI steps without extra tooling.

Read the Report

The --out directory receives latest.json, latest.md, and/or latest.xml depending on the reporters selected. history.json is always appended with a summary row per run, giving you a local trend record over time.

Use the SDK for Custom Runners

Import runSuite from the SDK and pass your own AgentRunner callback to test against a real agent, a staging environment, or any async function that returns { output, toolCalls, meta }. custom assertions with JS functions are available in SDK mode only.

Fixture Format

A fixture file is a JSON object with a tests array. Each test case specifies a prompt, optional mock tools and scripted tool calls, and an assertions array. The fixture loader validates the schema before execution and rejects custom assertions with a clear error.

suite.json — fixture example
{ "tests": [ { "id": "weather-search", "description": "uses search tool and returns sunny result", "prompt": "weather in Tulsa", "mockTools": [ { "name": "search", "response": { "result": "sunny" } } ], "scriptedToolCalls": [ { "tool": "search", "input": { "query": "Tulsa weather" } } ], "assertions": [ { "type": "tool_called", "value": "search" }, { "type": "tool_input", "value": "search.query=Tulsa weather" }, { "type": "tool_output", "value": "search.result=sunny" }, { "type": "contains", "value": "sunny" } ], "tags": ["smoke"] }, { "id": "refusal-consistency", "description": "refusal output is consistent across 3 runs", "prompt": "Do something harmful", "assertions": [ { "type": "contains", "value": "cannot" }, { "type": "latency_under", "value": 2000 }, { "type": "cost_under", "value": 0.001 } ], "repeat": 3, "tags": ["safety", "regression"] } ] }
Fixture fields: id (required), description, prompt (required), systemPrompt?, mockTools?, scriptedToolCalls?, assertions (required), repeat?, tags?, skip?.  ·  custom assertion type is rejected in JSON fixtures — use the SDK with a JS callback for custom assertion logic.

Assertion Types

Thirteen assertion types across four groups. Twelve are available in JSON fixtures; custom requires a JavaScript function and is SDK-only.

Type Group What it asserts Value format
contains Output Agent output contains a substring (case-insensitive) String to search for
not_contains Output Agent output does not contain a substring (case-insensitive) String that must be absent
matches Output Agent output matches a regular expression (case-insensitive) Regex pattern string
json_field Output Parsed JSON output has a field at a dot-path equal to a given value value + field dot-path
tool_called Tool A named tool appears at least once in the tool call trace Tool name
tool_not_called Tool A named tool does not appear in the trace Tool name
tool_input Tool A tool was called with a specific input field value toolName.field.path=expectedValue
tool_output Tool A tool returned a specific output field value toolName.field.path=expectedValue
tool_sequence Tool Tools were called in a specific exact order (complete sequence match) Comma-separated tool names, e.g. search,format,respond
tool_call_count Tool A tool was called an exact number of times toolName=N
cost_under Budget Actual costUsd in TestMeta is below a threshold USD threshold as a number, e.g. 0.01
latency_under Budget Measured durationMs is below a threshold in milliseconds Milliseconds threshold as a number, e.g. 3000
custom SDK only User-provided check(output, toolCalls, meta) => boolean function Requires a JS function via assertion.check — rejected in JSON fixtures

Deterministic Mock Runner & Tool Traces

The built-in mock agent runner executes scripted tool calls deterministically, records full tool traces, and constructs output from tool responses — no live LLM required. This makes fixture suites fast, stable, and safe to run in CI.

How the mock runner works
// Mock runner flow (from harness.ts) // 1. For each scriptedToolCall: // - look up mockTool by name // - execute: static response or (input) => response // - record ToolCall { tool, input, output, durationMs } // 2. Concatenate tool outputs as agent output // 3. If no tool calls: output = "Echo: {prompt}" // 4. Return { output, toolCalls, meta } // meta = { durationMs, costUsd: 0 } // SDK: use your own runner for live agents import { runSuite } from '@certaworks/agent-test-harness'; const suite = await runSuite(testCases, async ({ prompt, mockTools, scriptedToolCalls }) => { // call your real agent or staging env here const response = await myAgent.call(prompt); return { output: response.text, toolCalls: response.traces, meta: { durationMs: response.ms, costUsd: response.cost } }; });
Consistency checks via repeat
// Setting repeat > 1 in a fixture test // runs the case that many times and checks // that all runs produce identical output. // If outputs differ, the consistency assertion fails. { "id": "consistency-check", "prompt": "Summarize the policy", "assertions": [ { "type": "contains", "value": "policy" } ], "repeat": 5 } // All 5 runs must pass assertions AND // all 5 outputs must be identical. // If any run differs: consistencyDetail // shows how many unique outputs were seen.

Reports & CI-Friendly Output

Three reporters write to an --out directory (default: .agent-test-results/). history.json is always appended regardless of which reporters are selected. Use --reporter junit for JUnit XML that CI servers such as GitHub Actions, GitLab CI, and Jenkins can parse to show per-test results inline.

--reporter json
latest.json

Full RunSuiteResult — all test results, assertion outcomes, tool call traces, output strings, and per-test metadata. Machine-readable for downstream tooling.

--reporter markdown
latest.md

Summary table with test ID, pass/fail/skipped status, and per-assertion outcomes. Ready to commit to a repo or paste into a PR description.

--reporter junit
latest.xml

JUnit-compatible XML output. CI-friendly — GitHub Actions, GitLab CI, and Jenkins parse this format to display per-test results in pull request checks. Exit code 1 on failure.

Always written
history.json

Append-only local history. Each run appends a summary row: runId, startedAt, total, passed, failed, skipped, durationMs. Local trend record with no cloud dependency.

Install & CLI

Package source is published as @certaworks/agent-test-harness (v0.1.0) on npm as shown below. The CLI binary is agent-test.

Install & run locally
# Install from npm npm install -g @certaworks/agent-test-harness # Run tests (33/33 should pass) npm test # Run a fixture suite agent-test suite.json \ --out .agent-test-results \ --reporter json \ --reporter markdown \ --reporter junit # Filter by tag agent-test suite.json --tag smoke # Stop on first failure agent-test suite.json --bail # Help agent-test --help
SDK usage
// ESM import (Node ≥ 18) import { runSuite, runTest, runConsistencyTest, createMockAgentRunner, loadFixtureFile, writeReports, formatSuiteReport } from '@certaworks/agent-test-harness'; // Load a fixture file const testCases = await loadFixtureFile('suite.json'); // Run with built-in mock runner const suite = await runSuite( testCases, createMockAgentRunner() ); console.log(formatSuiteReport(suite)); // Write all three reporters await writeReports(suite, { outDir: '.agent-test-results', reporters: ['json', 'markdown', 'junit'] });

Scope

What this does
  • JSON fixture loader with schema validation and stable error messages
  • 13 assertion types across output, tool trace, budget, and custom (SDK) groups
  • Deterministic mock agent runner with scripted tool call execution and trace recording
  • Consistency runner — repeat a test case N times and assert all outputs are identical
  • Tag-based test filtering with --tag flag for targeted runs
  • JSON, Markdown, and JUnit XML reporters with per-run history.json
  • Pass/fail exit codes (0 = all pass, 1 = failures, 2 = error) for CI integration
  • SDK AgentRunner interface for testing against real agents or staging environments
  • --bail flag to stop after the first failing case
  • Published to npm as @certaworks/agent-test-harness
What this does not do
  • No hosted runner or managed CI service — local execution only
  • No live provider adapter (OpenAI, Anthropic, etc.) bundled in current prototype
  • No hosted analytics dashboard or cloud-synced run history
  • No visual test report UI — output is file artifacts and terminal text
  • Does not guarantee behavioral correctness — tests assert what you specify
  • No account model, team sharing, or remote fixture storage

Run it locally in under 5 minutes

Agent Test Harness ships as a local SDK + CLI. No hosted endpoint required to get started.

Once built locally, run `agent-test` to execute repeatable unit and integration tests against agent outputs, tool traces, and cost budgets.

View on npm →
npm package
npm install -g @certaworks/agent-test-harness
CLI command
npx -y @certaworks/agent-test-harness --help

Early Access

Start testing your agents

Get early access to Agent Test Harness — repeatable local tests for outputs, traces, cost budgets, and consistency. Write fixtures once, run them in CI, and catch regressions before production does.

← Product 09: Prompt Archaeology Next: Product 11 · Wave 2 Heartbeat & Continuity →