09 · Local Probe Runner · API · MCP · CLI

Prompt
Archaeology

Inspect what an agent actually does — not what its documentation says it does. Scripted probes surface behavioral signals across persona, constraints, refusal style, tool policy, and formatting habits. Evidence-backed estimates, not exact prompt recovery.

Local slice · v0.1.0 · Node ≥ 18
← Product 08: Audit & Replay Logger Next: Product 10 Agent Test Harness →

You inherit agents without knowing how they behave

Teams regularly deploy, integrate, or hand off agents whose behavioral constraints, refusal patterns, tool policies, and formatting tendencies are undocumented. System prompts are confidential. Model providers change behavior across versions. Inherited agents may silently restrict topics, redirect certain requests, or respond in ways that conflict with downstream expectations — and there is no way to know until production.

Prompt Archaeology gives teams a repeatable local probe runner. It sends scripted prompts, observes responses, and produces evidence-backed reports covering eight behavioral categories: persona, constraints, refusal style, tool policy, format rules, tone, knowledge cutoff, and ownership signals.

The result is an audit trail — not exact prompt recovery. The tool estimates what behavioral constraints are likely present based on observed signals, with explicit confidence levels, evidence quotes, and caveats included in every report. A local HTTP API, CLI, and MCP server make it usable from any workflow.

How it works

Three steps: select probes, supply responses (live or captured), receive an evidence-backed report with per-probe confidence, signal quotes, and limitations.

Select Probes

Browse 16 built-in probes across 8 categories using list_probes or GET /api/probes. Filter by category — persona, constraints, refusal_style, tool_policy, format_rules, tone, knowledge_cutoff, or ownership — to target specific behavioral dimensions.

Supply Responses

Either pass a live agent callback so the runner executes probes in batch, or supply already-captured { probe_id, response } pairs from a transcript. Both paths produce the same analysis. No cloud call is required.

Read the Report

Each report includes per-probe signal hit rates, confidence levels, evidence quotes, synthesized inferences per category, and explicit limitations. Export as JSON, Markdown, or plain text. Runs are saved to a durable local store for comparison across sessions.

Probe Catalog — 8 Categories, 16 Probes

All probes are scripted text prompts with declared signal keywords and an interpretation statement. When an agent response contains a signal keyword, it contributes to the probe's signal hit rate and confidence level.

Inference, not recovery. Probes observe surface-level response signals. They do not extract, decode, or reconstruct hidden system prompts. Every finding is an estimate based on one or more observed responses and should be treated as an audit lead, not a definitive policy claim.
persona
Persona & Role

Does the agent have a defined name or role? Probes ask directly for its identity and its design purpose to detect persona constraints.

p-name p-role
ownership
Ownership & Creator

Does the agent reveal or conceal its creator and product affiliation? Probes ask which company built it and what product it is part of.

p-company p-product
constraints
Behavioral Constraints

What is the agent prohibited from doing? Probes test direct override attempts, harm refusal, system prompt reveal requests, and topic restrictions.

c-refuse-direct c-refuse-harm c-reveal-prompt c-topic
refusal_style
Refusal Style

How does the agent refuse? Probes distinguish between firm flat-refusals ("I cannot") and redirect-and-offer patterns ("Instead, I can help with…").

r-style-1 r-redirect
tool_policy
Tool Policy

What tools or actions does the agent claim access to? Probes ask about capabilities and whether it can take autonomous actions like sending email or posting to social media.

t-tools t-actions
format_rules
Format Rules

Does the agent obey explicit length instructions or does it apply a fixed format regardless? Probes test length compliance and default markdown usage.

f-length f-markdown
tone
Tone

How does the agent respond to informal input? Does it mirror casual language or maintain formal register? Signals distinguish formal versus casual response patterns.

tone-formal
knowledge_cutoff
Knowledge Cutoff

Does the agent acknowledge a training cutoff? Asking for the current date surfaces whether the agent claims real-time access or reports a training cutoff boundary.

k-cutoff

Evidence-Backed Inference Model

Every probe result includes a signal hit rate (matched keywords ÷ total signals), a confidence level, supporting evidence quotes, and at least one explicit caveat. The report never asserts certainty about hidden system instructions.

Confidence Signal Hit Rate Interpretation Caveat always present
high ≥ 0.5 — half or more signals matched Observed strong behavioral signal for the probe's interpretation Estimate from observed responses, not hidden prompt recovery
medium ≥ 0.2 and < 0.5 — partial match Observed possible behavioral signal; treat as an audit lead Partial signal match; may reflect general training rather than a system instruction
low < 0.2 — few or no signals matched No strong behavioral signal observed for this probe Low signal match; do not treat this as a stable behavioral constraint
Category-level synthesis

After individual probe analysis, the report synthesizes findings per category. A category gets high confidence when at least half of its probes produce high-confidence results. It gets medium when the average signal hit rate is ≥ 0.2. Otherwise it is marked low. The report also synthesizes a refusalStyle inference (redirect-and-offer vs firm refusal vs ambiguous) by counting redirect signals against flat-refusal signals across all constraint and refusal-style probes.

Report Formats & Export

Every run produces a structured ArchaeologyReport. Export it locally in three formats using the SDK, CLI, or export_report MCP tool. No cloud upload, no hosted analytics.

json

Full structured report — machine-readable. Includes all probe results, evidence snippets, inference claims, category findings, and raw signal data.

id, generatedAt, scope, methodology
limitations (array)
probedCategories, categoryFindings
inferenceClaims with evidence + rationale
inferences (persona, owner, refusal, tools, format, tone, cutoff)
rawResults with signalsFound, signalHitRate, caveats
markdown

Human-readable structured report. Sections for Methodology, Limitations, Inferences, Evidence — ready for sharing in docs or PR reviews.

# Prompt Archaeology Report header
## Methodology + ## Limitations
## Inferences with confidence badges
## Evidence per probe-id with quotes
text

Plain-text ASCII report. Terminal-safe. Sections delimited with decorative separators for quick scanning in CI output or shell sessions.

═══ PROMPT ARCHAEOLOGY REPORT header
── METHODOLOGY + INFERENCES sections
── CATEGORY BREAKDOWN per category
── EVIDENCE per probe with caveat lines

Install & Configure

Package source is published as @certaworks/prompt-archaeology-tool (v0.1.0) on npm as shown below. Three binaries are exposed: prompt-archaeology (CLI), prompt-archaeology-mcp (MCP server), and prompt-archaeology-api (HTTP server).

Install & run locally
# Install from npm npm install -g @certaworks/prompt-archaeology-tool # Run the MCP server npm run mcp # or: prompt-archaeology-mcp # Run the HTTP server (default: 127.0.0.1:4320) npm run serve # or: prompt-archaeology-api # CLI one-shot analysis prompt-archaeology \ --input responses.json \ --json-out report.json \ --text-out report.md
MCP client config
{ "mcpServers": { "prompt-archaeology": { "command": "node", "args": ["dist/mcp/server.js"], "env": { "PROMPT_ARCHAEOLOGY_STORE_PATH": "./.prompt-archaeology/runs.json" } } } }

SDK Surface

Import the SDK directly for programmatic probe runs and report export without MCP or HTTP. Pass a live agent callback for active probing, or supply pre-captured responses pairs for offline analysis.

Live agent probing
// ESM import (Node ≥ 18) import { runProbeBatch, createRunStore, exportReport } from '@certaworks/prompt-archaeology-tool'; // Run probes through a live agent callback const run = await runProbeBatch({ target: { name: 'Support assistant', vendor: 'local' }, probeIds: ['p-name', 'c-reveal-prompt', 't-tools'], agent: async probe => sendPromptToAgent(probe.prompt) }); // Save run + export Markdown report await createRunStore().saveRun(run); await exportReport(run.report, { format: 'markdown', filePath: './report.md' });
Offline / captured transcript
// Analyze already-captured responses const run = await runProbeBatch({ target: { name: 'Captured transcript' }, responses: [ { probe_id: 'p-name', response: "I'm Atlas." }, { probe_id: 'c-refuse-harm', response: 'I cannot help with harmful requests.' }, { probe_id: 'c-reveal-prompt', response: 'My instructions are confidential.' } ] }); // All three formats await exportReport(run.report, { format: 'json', filePath: 'report.json' }); await exportReport(run.report, { format: 'markdown', filePath: 'report.md' }); await exportReport(run.report, { format: 'text', filePath: 'report.txt' });

MCP Tools

Four tools exposed over the MCP protocol. Callable from any MCP-compatible agent runtime. All analysis is local — no cloud calls.

list_probes

List available prompt archaeology probes. Optionally filter by category to return only probes for a specific behavioral dimension.

category?
analyze_response

Analyze one observed agent response for cautious behavioral signals. Provide a probe ID and the observed response string — returns confidence, signals found, evidence quotes, and caveats.

probe_id response
build_report

Build a full evidence-backed archaeology report from a set of { probe_id, response } result pairs. Returns the report as structured JSON or formatted plain text.

results format? (json|text)
export_report

Build a report and write it to a local file. Supports json, markdown, and text formats. Returns { ok: true, path } on success.

results format path

HTTP API

The local HTTP server binds to 127.0.0.1:4320 by default (PORT and HOST env vars override). All routes return JSON except /dashboard which serves a lightweight HTML page. No authentication — local access only.

Method Path Description
GET /dashboard Lightweight local HTML dashboard listing available API routes
GET /health Health check — returns { ok: true }
GET /api/probes List all probes. Optional ?category= query param filters by category
POST /api/analyze Analyze one response. Body: { probe_id, response } — returns a ProbeResult
POST /api/reports Build a full report from result pairs. Body: { results: [{ probe_id, response }] }
POST /api/runs Execute a probe batch and save to run history. Body: { target, responses? }
GET /api/runs List all saved runs as summaries (id, createdAt, target, totalProbes, errors)
GET /api/runs/:id Load a specific saved run by ID — returns the full ProbeBatchRun

CLI

The prompt-archaeology bin reads a JSON input file of captured responses, analyzes them, and writes reports to local files. Suitable for CI pipelines, shell scripts, and one-shot audit sessions.

Command usage
# Basic usage prompt-archaeology \ --input responses.json \ --json-out report.json \ --text-out report.md # --text-out and --markdown-out are aliases # Omit --json-out/--text-out to print JSON to stdout # Help prompt-archaeology --help
Input file shape
{ "target": { "name": "Support assistant", "vendor": "acme-corp", "model": "gpt-4o" }, "responses": [ { "probe_id": "p-name", "response": "I'm Aria, your support agent." }, { "probe_id": "c-refuse-harm", "response": "I cannot assist with that request." } ] }

Local Run History & Dashboard

Every probe batch run can be saved to a durable local store. The HTTP server provides a lightweight /dashboard page and /api/runs endpoints for listing and loading past runs. All history stays local — no data leaves your machine.

Local run store

By default, run history is stored at .prompt-archaeology/runs.json relative to the current working directory. Override with either environment variable:

PROMPT_ARCHAEOLOGY_STORE_PATH=/path/to/runs.json PROMPT_ARCHAEOLOGY_STORE=/path/to/runs.json
Store schema: { "version": 1, "runs": [ ...ProbeBatchRun ] } — newest runs first. Runs include the full report, all probe results, evidence, and errors.

Scope

What this does
  • 16 scripted probes across 8 behavioral categories: persona, constraints, refusal style, tool policy, format rules, tone, knowledge cutoff, ownership
  • Evidence-backed confidence model (high/medium/low) with signal hit rates and quote snippets per probe
  • Explicit caveats and limitations embedded in every probe result and report
  • JSON, Markdown, and plain-text report export — local files only
  • Durable local run history with save/list/load APIs
  • 4 MCP tools: list probes, analyze response, build report, export report
  • Local HTTP API (8 endpoints) with lightweight dashboard page
  • CLI for batch analysis from captured response files
  • Offline analysis from pre-captured { probe_id, response } pairs
  • Published to npm as @certaworks/prompt-archaeology-tool
What this does not do
  • Does not recover hidden system prompts verbatim — inference only
  • No hosted SaaS, API service, or cloud-based probe execution
  • No production provider adapter network (OpenAI, Anthropic, etc.) in current prototype
  • No API-key billing, credit meter, or pay-per-probe pricing
  • No authenticated dashboard or multi-user run history
  • Does not guarantee inferred constraints are complete or exhaustive
  • Probe signals are keyword-based — no LLM or semantic scoring in current implementation

Run it locally in under 5 minutes

Prompt Archaeology Tool ships as a local probe runner with MCP + CLI. No hosted endpoint required to get started.

Once running locally, call `analyze_response` to run scripted probes against agent responses and estimate behavioral constraints.

View on npm →
npm package
npm install -g @certaworks/prompt-archaeology-tool
Claude Desktop config
{
  "mcpServers": {
    "prompt-archaeology-tool": {
      "command": "npx",
      "args": ["-y", "-p", "@certaworks/prompt-archaeology-tool", "prompt-archaeology-mcp"]
    }
  }
}

Early Access

Start your behavioral audit

Get early access to Prompt Archaeology Tool — local probe runner, evidence-backed inference, and run history for agents you inherit, integrate, or need to understand before trusting in production.

← Product 08: Audit & Replay Logger Next: Product 10 Agent Test Harness →