Call write_memory with a
type of
episodic or
semantic.
Attach tags, set an importance score (0–1), and optionally scope the record to a
session ID. The record is written to the local JSON store immediately and persists
across agent runs.
Product 06 · Local MCP Server · Local Slice
Cognitive Memory
Give agents durable local memory that persists across sessions. Store episodic and semantic records, retrieve by keyword or tag, and pack relevant context into the next prompt — all from a local JSON store with no cloud required.
Agents forget everything between sessions.
Every new agent session starts from zero. Decisions made in the previous run, context accumulated across tool calls, and knowledge the agent worked hard to assemble — all of it disappears the moment the context window closes. There is no structured way to carry useful signal forward.
Without a memory layer, agents repeat work they have already done because nothing was retained. They fill context windows with stale or reconstructed content instead of retrieving what actually matters. Long-running workflows lose continuity; short sessions cannot benefit from earlier sessions.
Cognitive Memory MCP Server gives agents a local memory store they can write to explicitly, search by relevance, and pack into the next prompt — without any hosted service required to get started.
Store → Recall → Pack
Call search_memory with a query
string. Search uses deterministic lexical relevance — token overlap between the query
and stored content and tags, weighted by recency and importance. No vector embeddings
or external model calls are required. Results can be filtered by type, session ID,
tags, minimum score, and result limit.
Call pack_context or
summarize_context with a query and
a character budget. The store runs a relevance search, then greedily packs the highest-scoring
records that fit within the budget. summarize_context
returns a numbered plaintext summary formatted for prompt injection. This is
deterministic text composition — not LLM-generated synthesis.
.cognitive-memory/memory.json
by default. Override the path with the
COGNITIVE_MEMORY_STORE_PATH
or COGNITIVE_MEMORY_STORE
environment variables. There is no sync, no cloud connection, and no hosted endpoint required.
Memory persists only on the local file system.
Two durable record types
The store supports two memory types pulled directly from the source. Both share the same record structure and storage format. The type field controls how the agent categorizes and filters memories.
Records of things that happened — agent actions, decisions made, tool calls executed, outcomes observed, or events that occurred during a session. Episodic memories are anchored in time and specific context. Use them to capture what the agent did and what it learned from doing it.
- type: "episodic"
- content: string — the memory text
- tags: string[] — filterable labels
- importance: number (0–1, default 0.5)
- sessionId: string (optional scope)
- createdAt, lastAccessedAt, accessCount
General knowledge, facts, rules, patterns, or durable observations that are not tied to a single session moment. Semantic memories describe what the agent knows rather than what it did. Use them to capture reusable facts, project conventions, learned constraints, or domain knowledge that should persist across many sessions.
- type: "semantic"
- content: string — the memory text
- tags: string[] — filterable labels
- importance: number (0–1, default 0.5)
- sessionId: string (optional scope)
- createdAt, lastAccessedAt, accessCount
Deterministic lexical search, no embeddings required
Search is fully deterministic and runs without any API key or external model call. The relevance function computes token overlap between the query and stored content and tags (after stopword removal), blended with tag match rate and a recency boost. Records are ranked by score and can be filtered by type, session ID, tags, minimum score threshold, and result count limit.
Relevance = (token overlap / query tokens) × 0.6 + (tag match rate) × 0.3 +
(recency boost over 30 days) × 0.1. Final record score multiplies relevance
by (0.5 + importance × 0.5)
so higher-importance records score proportionally better. Default minimum score
threshold: 0.05. Default result limit: 10 (maximum: 100).
Context packing behavior
pack_context runs a relevance search
(up to 50 candidates, minimum score 0.01), then greedily adds records to the output
in score order until the character budget is exhausted. Each record contributes its
content length plus 50 characters of overhead. Records that would exceed the budget
are counted as omitted. The result reports total characters used, budget fraction
consumed, and omitted count.
summarize_context runs the same
packing logic, then formats the packed records as a numbered list:
1. episodic [tag1, tag2]: content text.
This is deterministic text composition from the stored records — not an LLM call.
The default character budget is 8,000 characters.
Search does not use vector embeddings, cosine similarity, or semantic model calls. Embedding-based search is listed in the README as future work. Retrieval accuracy depends on lexical overlap between the query and the stored content — well-tagged records with clear content text will return the best results.
MCP server, SDK, and local JSON store
Cognitive Memory MCP Server ships as a local Node package with a stdio MCP server, SDK exports, and a file-backed JSON store. Once running, any MCP-compatible agent runtime can call memory tools directly.
The package bin is
cognitive-memory-mcp.
The MCP server reads from and writes to the local store on every tool call.
npm install -g @certaworks/cognitive-memory-mcp-server
{
"mcpServers": {
"cognitive-memory-mcp-server": {
"command": "npx",
"args": ["-y", "@certaworks/cognitive-memory-mcp-server"]
}
}
}
Available MCP tools
Cognitive Memory MCP Server exposes the following tools to any MCP-compatible agent
runtime. Tool names match the TOOLS array in
src/mcp/server.ts exactly.
File-backed local JSON store
All episodic and semantic memory records are written to a single versioned JSON file. The store is loaded from disk on every tool call and saved after every write or delete. There is no separate file per memory type — all records are stored together in one file.
Default storage path:
.cognitive-memory/memory.json
(relative to the working directory when the server starts). Override with either:
COGNITIVE_MEMORY_STORE_PATH
or
COGNITIVE_MEMORY_STORE
environment variables.
{
"version": 1,
"records": [
{
"id": "uuid",
"type": "episodic",
"content": "...",
"tags": ["tag1"],
"importance": 0.8,
"createdAt": "...",
"accessCount": 0,
"lastAccessedAt": "...",
"sessionId": "optional"
}
]
}
Hosted durable memory storage, team memory workspaces, encrypted remote retention, multi-user accounts, and vector/embedding search are roadmap items. They are not currently built or available. What is available today is the local MCP server, SDK exports, and the file-backed local JSON store.
What Cognitive Memory does and does not do
- Store durable episodic and semantic memory records in a local JSON file
- Retrieve memories by deterministic lexical relevance — keyword, tag, recency, and importance weighted
- Filter search by memory type, session ID, tags, minimum score, and result limit
- Pack and summarize relevant memories within a configurable character budget for prompt injection
- Expose MCP tools for write, search, list, pack, summarize, delete, clear, and stats operations
- Persist memories across agent sessions using the local file system
- Provide hosted memory SaaS or team memory accounts (roadmap)
- Use vector embeddings or semantic search — lexical retrieval only; embedding search is future work
- Provide encrypted remote retention or multi-user access control
- Guarantee perfect recall — lexical search returns best-match records, not exhaustive retrieval
- Provide privacy or compliance guarantees — memory content is stored in plaintext locally
-
Generate LLM-synthesized summaries —
summarize_contextis deterministic text formatting, not a model call
Get early access to hosted Cognitive Memory
The local MCP server and JSON store are available now for private testing. Hosted memory storage with cloud retention, team workspaces, and managed recall is in development. Leave your details and we will reach out when hosted access opens.