09 · Local Probe Runner · API · MCP · CLI
Inspect what an agent actually does — not what its documentation says it does. Scripted probes surface behavioral signals across persona, constraints, refusal style, tool policy, and formatting habits. Evidence-backed estimates, not exact prompt recovery.
Teams regularly deploy, integrate, or hand off agents whose behavioral constraints, refusal patterns, tool policies, and formatting tendencies are undocumented. System prompts are confidential. Model providers change behavior across versions. Inherited agents may silently restrict topics, redirect certain requests, or respond in ways that conflict with downstream expectations — and there is no way to know until production.
Prompt Archaeology gives teams a repeatable local probe runner. It sends scripted prompts, observes responses, and produces evidence-backed reports covering eight behavioral categories: persona, constraints, refusal style, tool policy, format rules, tone, knowledge cutoff, and ownership signals.
The result is an audit trail — not exact prompt recovery. The tool estimates what behavioral constraints are likely present based on observed signals, with explicit confidence levels, evidence quotes, and caveats included in every report. A local HTTP API, CLI, and MCP server make it usable from any workflow.
Three steps: select probes, supply responses (live or captured), receive an evidence-backed report with per-probe confidence, signal quotes, and limitations.
Browse 16 built-in probes across 8 categories using list_probes or
GET /api/probes. Filter by category — persona,
constraints, refusal_style, tool_policy,
format_rules, tone, knowledge_cutoff,
or ownership — to target specific behavioral dimensions.
Either pass a live agent callback so the runner executes probes in batch,
or supply already-captured { probe_id, response } pairs from a transcript.
Both paths produce the same analysis. No cloud call is required.
Each report includes per-probe signal hit rates, confidence levels, evidence quotes, synthesized inferences per category, and explicit limitations. Export as JSON, Markdown, or plain text. Runs are saved to a durable local store for comparison across sessions.
All probes are scripted text prompts with declared signal keywords and an interpretation statement. When an agent response contains a signal keyword, it contributes to the probe's signal hit rate and confidence level.
Does the agent have a defined name or role? Probes ask directly for its identity and its design purpose to detect persona constraints.
Does the agent reveal or conceal its creator and product affiliation? Probes ask which company built it and what product it is part of.
What is the agent prohibited from doing? Probes test direct override attempts, harm refusal, system prompt reveal requests, and topic restrictions.
How does the agent refuse? Probes distinguish between firm flat-refusals ("I cannot") and redirect-and-offer patterns ("Instead, I can help with…").
What tools or actions does the agent claim access to? Probes ask about capabilities and whether it can take autonomous actions like sending email or posting to social media.
Does the agent obey explicit length instructions or does it apply a fixed format regardless? Probes test length compliance and default markdown usage.
How does the agent respond to informal input? Does it mirror casual language or maintain formal register? Signals distinguish formal versus casual response patterns.
Does the agent acknowledge a training cutoff? Asking for the current date surfaces whether the agent claims real-time access or reports a training cutoff boundary.
Every probe result includes a signal hit rate (matched keywords ÷ total signals), a confidence level, supporting evidence quotes, and at least one explicit caveat. The report never asserts certainty about hidden system instructions.
| Confidence | Signal Hit Rate | Interpretation | Caveat always present |
|---|---|---|---|
| high | ≥ 0.5 — half or more signals matched |
Observed strong behavioral signal for the probe's interpretation | Estimate from observed responses, not hidden prompt recovery |
| medium | ≥ 0.2 and < 0.5 — partial match |
Observed possible behavioral signal; treat as an audit lead | Partial signal match; may reflect general training rather than a system instruction |
| low | < 0.2 — few or no signals matched |
No strong behavioral signal observed for this probe | Low signal match; do not treat this as a stable behavioral constraint |
After individual probe analysis, the report synthesizes findings per category.
A category gets high confidence when at least half of its probes produce high-confidence results.
It gets medium when the average signal hit rate is ≥ 0.2.
Otherwise it is marked low.
The report also synthesizes a refusalStyle inference
(redirect-and-offer vs firm refusal vs ambiguous) by counting
redirect signals against flat-refusal signals across all constraint and refusal-style probes.
Every run produces a structured ArchaeologyReport. Export it locally in three
formats using the SDK, CLI, or export_report MCP tool.
No cloud upload, no hosted analytics.
Full structured report — machine-readable. Includes all probe results, evidence snippets, inference claims, category findings, and raw signal data.
Human-readable structured report. Sections for Methodology, Limitations, Inferences, Evidence — ready for sharing in docs or PR reviews.
Plain-text ASCII report. Terminal-safe. Sections delimited with decorative separators for quick scanning in CI output or shell sessions.
Package source is published as @certaworks/prompt-archaeology-tool (v0.1.0) on npm as shown below.
Three binaries are exposed: prompt-archaeology (CLI),
prompt-archaeology-mcp (MCP server), and
prompt-archaeology-api (HTTP server).
Import the SDK directly for programmatic probe runs and report export without MCP or HTTP.
Pass a live agent callback for active probing, or supply pre-captured
responses pairs for offline analysis.
Four tools exposed over the MCP protocol. Callable from any MCP-compatible agent runtime. All analysis is local — no cloud calls.
List available prompt archaeology probes. Optionally filter by category to return only probes for a specific behavioral dimension.
Analyze one observed agent response for cautious behavioral signals. Provide a probe ID and the observed response string — returns confidence, signals found, evidence quotes, and caveats.
Build a full evidence-backed archaeology report from a set of { probe_id, response } result pairs. Returns the report as structured JSON or formatted plain text.
Build a report and write it to a local file. Supports json, markdown, and text formats. Returns { ok: true, path } on success.
The local HTTP server binds to 127.0.0.1:4320 by default
(PORT and HOST env vars override). All routes return JSON
except /dashboard which serves a lightweight HTML page.
No authentication — local access only.
| Method | Path | Description |
|---|---|---|
| GET | /dashboard | Lightweight local HTML dashboard listing available API routes |
| GET | /health | Health check — returns { ok: true } |
| GET | /api/probes | List all probes. Optional ?category= query param filters by category |
| POST | /api/analyze | Analyze one response. Body: { probe_id, response } — returns a ProbeResult |
| POST | /api/reports | Build a full report from result pairs. Body: { results: [{ probe_id, response }] } |
| POST | /api/runs | Execute a probe batch and save to run history. Body: { target, responses? } |
| GET | /api/runs | List all saved runs as summaries (id, createdAt, target, totalProbes, errors) |
| GET | /api/runs/:id | Load a specific saved run by ID — returns the full ProbeBatchRun |
The prompt-archaeology bin reads a JSON input file of captured responses,
analyzes them, and writes reports to local files. Suitable for CI pipelines, shell scripts,
and one-shot audit sessions.
Every probe batch run can be saved to a durable local store. The HTTP server provides a
lightweight /dashboard page and /api/runs endpoints for
listing and loading past runs. All history stays local — no data leaves your machine.
By default, run history is stored at .prompt-archaeology/runs.json
relative to the current working directory. Override with either environment variable:
{ "version": 1, "runs": [ ...ProbeBatchRun ] }
— newest runs first. Runs include the full report, all probe results, evidence, and errors.
{ probe_id, response } pairs@certaworks/prompt-archaeology-toolPrompt Archaeology Tool ships as a local probe runner with MCP + CLI. No hosted endpoint required to get started.
Once running locally, call `analyze_response` to run scripted probes against agent responses and estimate behavioral constraints.
View on npm →npm install -g @certaworks/prompt-archaeology-tool
{
"mcpServers": {
"prompt-archaeology-tool": {
"command": "npx",
"args": ["-y", "-p", "@certaworks/prompt-archaeology-tool", "prompt-archaeology-mcp"]
}
}
}
Early Access
Get early access to Prompt Archaeology Tool — local probe runner, evidence-backed inference, and run history for agents you inherit, integrate, or need to understand before trusting in production.