Product 03 · Framework · SDK / MCP / CLI · Local Prototype

Multi-Voice Deliberation

Four voices — Builder, Skeptic, Evolver, Guardian — review every decision before your agent commits to an answer. Structured debate for high-stakes agent reasoning.

Local TypeScript prototype — provider support is narrow; hosted SaaS not yet live
← Product 02: Confidence Gate Next: Product 04 Agent Cost Router →

Single-pass answers miss risks, gaps, and better alternatives.

LLMs generate answers in a single forward pass. They have no built-in adversarial review, no mechanism for challenging their own assumptions, and no structured way to flag when a proposed action might be harmful, incomplete, or simply wrong. The first answer is often the only answer.

For low-stakes tasks this is fine. For high-stakes decisions — code changes that affect production, responses to upset customers, actions with irreversible consequences, or reasoning chains that inform further agent steps — a single-pass answer is a risk.

Multi-Voice Deliberation adds a structured panel review step before the agent commits. Four named voices each read the same question and respond from a defined perspective. Their positions, concerns, and suggestions are aggregated into a synthesis with a weighted confidence score. If panel confidence falls below the configured threshold, the result is flagged for human review before the agent proceeds.

Submit → Panel → Synthesize

Submit the question or decision

Call deliberate(adapter, options) with the prompt and optional context. Specify which voices to include, the model to use, the confidence threshold, and whether to run voices in parallel (default: true). Requires an API key for the selected provider.

Each voice reads and responds

Builder, Skeptic, Evolver, and Guardian each receive the same prompt with their own system prompt defining their role. Each voice returns a structured JSON object: position, concerns, suggestions, and a confidence score (0–1). Guardian carries a weight of 1.2 in the panel average.

Synthesize — flag for review if needed

The framework computes a weighted average panelConfidence, aggregates all concerns and suggestions, and produces a markdown synthesis. If panelConfidence falls below the approvalThreshold (default 0.6), requiresHumanReview is set true.

Each voice call is an LLM API call. Running all four in parallel is the default. Cost and latency scale with the model and provider selected. The framework does not make final decisions — it returns a synthesis and a review flag. Acting on the result is the responsibility of the calling code or the human reviewer.

Four voices. Four perspectives.

Each voice is defined by a name, a label, and a system prompt that shapes its response. All four are active by default. Voices can be filtered per call.

Builder

Constructs the best possible answer — concrete, actionable, and useful.

Focus
  • What should be done and how
  • Concrete steps and actionable detail
  • Making the response as useful as possible
Skeptic

Challenges assumptions and finds weaknesses in the proposed approach.

Focus
  • What could go wrong
  • Hidden assumptions and edge cases
  • What information is missing
Evolver

Improves and iterates — creative alternatives, hybrid approaches, second-order effects.

Focus
  • How to enhance the Builder's proposal
  • Approaches that address the Skeptic's concerns
  • Long-term implications and trade-offs
Guardian

Protects against harm, risk, and irreversibility. Flags when human review is needed.

Focus
  • Safety and ethical implications
  • Irreversible consequences
  • Safeguards and conditions for safe execution

Guardian carries panel weight 1.2 — heavier than the other three.

What deliberate() returns

The DeliberationResult interface is the exact return type of every deliberate() call.

FieldTypeDescription
promptstringThe original prompt passed to the panel
voiceResponsesVoiceResponse[]Per-voice position, concerns, suggestions, and confidence score
synthesisstringFormatted markdown with all voice positions, concerns, suggestions, and panel confidence
panelConfidencenumber (0–1)Weighted average confidence across all voices — Guardian weight is 1.2
requiresHumanReviewbooleanTrue when panelConfidence falls below approvalThreshold (default 0.6)
allConcernsstring[]Aggregate of all concerns raised across every voice
allSuggestionsstring[]Aggregate of all suggestions raised across every voice
durationMsnumberWall-clock time for the full deliberation in milliseconds

approvalThreshold defaults to 0.6. When requiresHumanReview is true, the calling code is responsible for pausing and routing the decision to a human reviewer. The framework flags — it does not block execution.

TypeScript SDK, MCP server, and CLI

Import deliberate, adapterFromEnv, and DEFAULT_VOICES into any Node ≥ 18 TypeScript project, or wire the MCP server into any MCP-compatible agent runtime. The CLI binary is deliberate.

Requires an API key for the selected provider. Set OPENROUTER_API_KEY, OPENAI_API_KEY, or GROQ_API_KEY in your environment. Provider support uses the OpenAI-compatible SDK.

View on npm →
npm package
# Install locally for MCP config and SDK imports
npm install @certaworks/multi-voice-deliberation-framework

# Optional: install globally for direct CLI commands
npm install -g @certaworks/multi-voice-deliberation-framework
SDK quickstart
import { deliberate, adapterFromEnv } from '@certaworks/multi-voice-deliberation-framework';

// Requires OPENROUTER_API_KEY, OPENAI_API_KEY, or GROQ_API_KEY
const adapter = adapterFromEnv();

const result = await deliberate(adapter, {
  prompt: 'Should we delete the legacy billing table after migration?',
  context: 'Migration completed 2 hours ago. No rollback confirmed.',
  model: 'openai/gpt-4o-mini', // via OpenRouter
  approvalThreshold: 0.65,     // flag for review below this
  parallel: true               // run voices concurrently
});

console.log(result.panelConfidence);    // e.g. 0.54
console.log(result.requiresHumanReview); // true — pause before acting
console.log(result.allConcerns);
console.log(result.synthesis);
CLI usage
# Run a deliberation from the command line
deliberate --prompt "Refactor auth middleware after failing tests?" \
           --model "openai/gpt-4o-mini" \
           --threshold 0.6

OpenAI-compatible providers via env vars

adapterFromEnv() checks for the following API keys in order and returns a ready-to-use adapter. Any OpenAI-compatible endpoint can also be wired manually via OpenAICompatAdapter.

OpenRouter ·  OPENROUTER_API_KEY · default model: openai/gpt-4o-mini
OpenAI ·  OPENAI_API_KEY · default model: gpt-4o-mini
Groq ·  GROQ_API_KEY · default model: llama-3.1-8b-instant
Provider support is implemented via the openai npm package with a configurable baseURL. Anthropic and other non-OpenAI-compatible providers are not supported by the current adapter. Model behavior and output quality will vary by provider and model selected. The default model is gpt-4o-mini — override with the DELIBERATION_MODEL env var or per-call option.

Available MCP tools

Multi-Voice Deliberation exposes the following tools to any MCP-compatible agent runtime. Both tools require a configured provider API key in the server environment.

deliberate
Run a multi-voice deliberation panel (Builder, Skeptic, Evolver, Guardian) on a question or proposed action before committing. Returns a synthesis, per-voice positions, concerns, suggestions, and a human-review flag when confidence is low.
list_voices
List the available deliberation voices and their roles.
Project-local Claude Desktop config
{
  "mcpServers": {
    "multi-voice-deliberation-framework": {
      "command": "node",
      "args": ["./node_modules/@certaworks/multi-voice-deliberation-framework/dist/mcp/server.js"]
    }
  }
}

What Multi-Voice Deliberation does and does not do

Does
  • Run Builder, Skeptic, Evolver, and Guardian in parallel or sequentially on any prompt
  • Return a weighted panel confidence score with Guardian at 1.2× weight
  • Aggregate all voice concerns and suggestions into the result
  • Flag requiresHumanReview when confidence falls below the configured threshold
  • Support OpenRouter, OpenAI, and Groq via the OpenAI-compatible adapter
Does Not
  • Guarantee correctness — it improves review coverage, not accuracy
  • Replace human review — when the flag is raised, a human must decide
  • Support Anthropic or other non-OpenAI-compatible providers in the current adapter
  • Block agent execution — the result must be acted on by the calling code
  • Provide hosted SaaS, team accounts, or persistent deliberation history (roadmap)
  • Provide legal, safety, or compliance certification of any kind

Get early access to Multi-Voice Deliberation

The local TypeScript SDK and MCP server are available for private testing. Hosted deliberation API, persistent panel history, and team features are in development. Leave your details and we will reach out when broader access opens.

← Product 02: Confidence Gate Next: Product 04 Agent Cost Router →