← Product 01: Identity Engine SDK Next: Product 03 Multi-Voice Deliberation Framework →

Agents act. Then you find out.

AI agents call tools, write to files, send customer emails, and push production changes without any pre-action control point. By the time a team reviews the log, the action has already happened — a destructive database query has run, an irreversible API call has been made, or a commit has altered code that touches critical paths.

Most agent frameworks offer no systematic way to evaluate a proposed action before it reaches a tool. Operators see the result of what the agent did, not a checkpoint where the action could have been evaluated, held, or blocked before execution.

The window for review closes before it opens. Confidence Gate adds that window back.

Score → Route → Trace

Agent proposes an action

Your agent calls the check_gate tool with the proposed action text and any available context — touched files, test evidence, source, project, or risk metadata.

Confidence Gate scores it

The gate runs a 0–100 confidence score, checks the action against configured policy rules, and produces a factor analysis covering linguistic certainty, completeness, contextual alignment, and factual grounding.

Returns a decision with trace

The gate returns allow, review, or block along with the reasoning trace and a durable local log entry. The agent proceeds, pauses for human review, or halts based on the result.

Confidence Gate is a decision-support layer. It adds a pre-action control point between agent intent and tool execution. It does not replace human judgment or production controls.

Default confidence thresholds

Thresholds are configurable. These are the defaults shipped with Confidence Gate.

Domain / Scope Default threshold Outcome below threshold
Default 0.75 Review or block based on policy
General 0.70 Review or block based on policy
Code execution 0.85 Review required
Financial 0.90 Review required
Legal 0.90 Review required
Medical 0.92 Review required
Risk level: low 0.60 Review or block based on policy
Risk level: medium 0.75 Review or block based on policy
Risk level: high 0.90 Review required
Risk level: critical 0.95 Block unless explicitly overridden

Thresholds are set per policy rule. The values above reflect factory defaults from config.ts. All thresholds are overridable via the set_threshold MCP tool or environment variables.

Three outcomes. One clear interface.

Allow

Confidence is above the policy threshold and supporting evidence is present. The action proceeds with a trace record logged to local durable storage. No human intervention required.

Example: low-risk file read, documentation update, non-destructive query

Review

Confidence is below the configured threshold or a high-risk policy rule was matched. The action is held in the review queue until a human approves, rejects, or resolves it before the agent proceeds.

Example: authentication change, customer-facing message, production deploy

Block

A hard policy rule was matched — the action is rejected immediately. The reason is logged with a full trace record. The agent does not proceed without a policy change or explicit override.

Example: destructive data deletion without migration evidence, policy-prohibited action type

Run it locally in under 5 minutes

Confidence Gate ships as a local MCP server. No hosted endpoint required to get started.

Once the server is running, any MCP-compatible agent runtime can call check_gate to score and route proposed actions.

View on npm →
npm package
npm install -g @certaworks/confidence-gate-mcp-server
Claude Desktop config
{
  "mcpServers": {
    "confidence-gate": {
      "command": "npx",
      "args": ["-y", "@certaworks/confidence-gate-mcp-server"]
    }
  }
}

Available MCP tools

Confidence Gate exposes the following tools to any MCP-compatible agent runtime.

Scoring
score_confidence
Score agent output text from 0 to 1 with a full factor breakdown.
check_gate
Gate an action and return allow, review, or block with confidence trace.
explain_confidence
Retrieve a saved confidence trace by gate_id.
Policy & Thresholds
set_threshold
Update the default, domain, or risk-level threshold at runtime.
get_config
Inspect the active gate configuration including all current thresholds.
list_policies
List named project policies stored in the local policy store.
upsert_policy
Create or version a named project policy with full audit metadata.
list_policy_audit
List policy version and threshold-change audit events.
Trace & Audit
list_traces
List stored gate traces from the local trace store.
query_traces
Filter traces by decision, approval status, project, source, policy, domain, risk, or time range.
resolve_approval
Approve, reject, or resolve a review-queue item with approver identity and note.
record_enforcement_receipt
Record what a client actually did after receiving a gate decision.
Dashboard
get_dashboard_summary
Return dashboard counts, recent checks, review queue, config, policies, and policy audit events.

Local dashboard — prototype status

Confidence Gate includes a local HTML dashboard for reviewing scored decisions, exploring policy settings, and inspecting trace history. The dashboard runs locally alongside the MCP server and does not require any cloud connection.

Roadmap — not currently built

Hosted dashboard, team seats, persistent cloud trace storage, and managed alerting are roadmap items. They are not currently built or available. What is available today is the local MCP server and the local prototype dashboard.

View local dashboard prototype →

What Confidence Gate does and does not do

Does
  • Score proposed agent actions with configurable policy rules
  • Return allow / review / block with confidence factors and trace context
  • Log decisions to local durable storage with hash-chain integrity
  • Provide a local MCP server interface for any MCP-compatible agent runtime
Does Not
  • Guarantee safety or prevent all harmful agent actions
  • Replace human review, legal review, or production controls
  • Provide hosted storage, team accounts, or managed alerting (roadmap)
  • Cover actions the agent takes without calling the gate
  • Recover hidden prompts or provide full agent observability

Get early access to hosted Confidence Gate

The local MCP server is available now for private testing. Hosted beta with cloud trace and team features is in development. Leave your details and we will reach out when hosted access opens.

← Product 01: Identity Engine SDK Next: Product 03 Multi-Voice Deliberation Framework →