# Knowledge Bases

Keep owner-scoped knowledge in immutable revisions and retrieve either published current content or one exact addressed version.

- Canonical: https://docs.xmemo.dev/docs/concepts/knowledge-bases
- Locale: en-US
- Content-Locale: en-US
- Canonical-Content-Digest: 32c23d8ee7a9929c817fa0ccde3c21e727d025b885c4b0f451176422d70a7844
- Edition-Digest: 9a7ddb5eb39965c36d86d311c0fff7f850eb42e2a43ba11ccce8a5b2ec20aa98
- Source-Revision: sha256:dc475d4b6dabf512907c33eb893d58db1074ae2b84d6be8c1e2a4333e1cb7a94

## A Knowledge Base is a durable knowledge layer

A Knowledge Base is an owner-scoped collection of Knowledge items. Each item has a title, status, and current_revision_id pointer; each revision stores immutable canonical content, a revision number, content hash, source metadata, and provenance. Publishing a newer revision moves the item pointer while leaving the earlier revision available as an exact historical version. This separates stable reference material from event-shaped memory and makes citations reproducible.

- A base must be active for retrieval.
- A query search considers published items and only the revision named by each item's current_revision_id.
- The same server boundary can resolve a personal owner or an authenticated team-bound context; the caller cannot widen it by supplying a filter.

## Choose query search or addressed read

search_knowledge has two mutually exclusive read shapes. Query search supplies query and optionally narrows to knowledge_base_id; it returns a typed knowledge_search envelope with ranked result objects. Addressed read supplies knowledge_item_id and optionally knowledge_revision_id; when the revision is omitted, the item current revision is read, and when it is supplied, that exact immutable revision is read. The handler rejects mixing query with item or revision identifiers.

```json
{
  "type": "knowledge_search",
  "results": [
    {
      "knowledge_item_id": "<item-id>",
      "knowledge_revision_id": "<current-revision-id>",
      "title": "Runbook",
      "content": "<ranked excerpt>",
      "reference": "knowledge:<item-id>@<current-revision-id>",
      "score": 3.42
    }
  ]
}

{
  "type": "knowledge",
  "knowledge_item_id": "<item-id>",
  "knowledge_revision_id": "<exact-revision-id>",
  "content": "<bounded canonical content>",
  "content_offset": 0,
  "content_limit": 4000,
  "content_total_chars": 12345,
  "content_truncated": true
}
```

## Select the retrieval mode for the job

hybrid is the default and combines normalized lexical and semantic signals with alpha: alpha=0.5 gives equal weight, while alpha closer to 1 favors lexical evidence and alpha closer to 0 favors semantic evidence. Choose lexical when exact terms, identifiers, or phrase overlap are the strongest signal. Choose semantic when meaning matters more than wording and embeddings are available. Choose rrf when independent lexical and vector rankings should be fused by rank rather than compared on one score scale.

- hybrid — linear weighted combination of lexical and semantic scores; the semantic side falls back to lexical when no semantic score is available.
- lexical — normalized term/phrase matching only; useful for names, IDs, and exact wording.
- semantic — vector similarity when available, with the implementation's lexical fallback when it is not.
- rrf — reciprocal rank fusion of lexical and vector rankings; the service uses its configured rank constant.

## Strict currentness prevents stale search results

Query search first limits candidates to active Knowledge Bases, published Knowledge items, and the revision currently pointed to by each item. A superseded revision therefore cannot appear as a query result even if it is still present in the append-only revision chain. An addressed read is intentionally different: an explicit knowledge_revision_id asks for that exact revision, subject to owner, base, and published-item authorization, while omitting it follows the current pointer. This gives search freshness without losing historical reproducibility.

```text
query search       -> published item + current_revision_id only
addressed no pin   -> published item's current revision
addressed with pin -> exact immutable revision, bounded by offset/limit_chars
```

## The read surface is strict, bounded, and read-only

The MCP handler requires knowledge:read, accepts no caller-supplied owner or scope, emits no audit event, and bounds addressed content with offset and limit_chars. Query results are bounded by max_results. The committed OpenAPI contract also contains the authenticated REST companion POST /api/v1/knowledge/search; its full request/response reference and the separate Knowledge authoring routes remain outside this tool-focused slice.

```text
search_knowledge(
  query="deployment runbook",
  mode="hybrid",
  alpha=0.5,
  max_results=10
)

search_knowledge(
  knowledge_item_id="<item-id>",
  knowledge_revision_id="<revision-id>",
  offset=0,
  limit_chars=4000
)
```

## GA availability still follows the runtime switch

Knowledge Base is documented here as a normal GA capability under Han's 2026-09-02 production enablement confirmation, not as Preview. The repository's knowledge_runtime_enabled default remains false, so an environment that has not enabled the runtime returns knowledge_unavailable with code knowledge_runtime_disabled. That operational check is separate from the tool's read-only contract and does not make a disabled environment silently return stale or unauthorized content.

## ChatGPT user

Give ChatGPT durable access to your XMemo preferences, project facts, decisions, and TODOs without pasting bearer tokens into a chat.

Connect the hosted XMemo MCP server through the ChatGPT/OpenAI app OAuth flow, then approve the memory:read and memory:write grant for your XMemo account.

Save a synthetic preference or project note, start a new chat, then ask ChatGPT to recall it through XMemo before continuing work.

If OAuth fails or tools do not appear, sign out of the MCP server in the host app, reconnect the XMemo server URL, and retry before creating direct tokens.

## Copilot / Codex developer

Carry repo decisions, coding conventions, bug-fix notes, and task history between IDE and CLI agents.

Use OAuth for VS Code / GitHub Copilot and Gemini CLI when available. For Copilot CLI, Codex, Cursor, or other direct MCP clients, keep XMEMO_KEY in the local environment or secret store and set a stable XMEMO_AGENT_INSTANCE_ID.

Record a codebase decision or bug fix, then ask the next IDE or CLI agent to recall the relevant XMemo context before editing.

If recalls are empty, verify the selected MCP config path, the XMEMO_KEY environment variable for direct clients, and any stale OAuth credential in the host app.

## Team / enterprise pilot owner

Evaluate shared memory with account controls, source attribution, export/delete workflows, and reviewer-safe setup evidence.

Create or enter the protected XMemo workspace, invite approved users, then connect each client through OAuth or a scoped direct credential according to the readiness badges.

Have a pilot member save a synthetic team memory, confirm source attribution in XMemo, then review delete/export and support paths.

If a member cannot connect, check role permissions, OAuth approval, client readiness status, and support guidance before issuing a new token.

## Autonomous agent operator

Let headless or scheduled agents record progress, retrieve prior decisions, and keep a stable non-secret instance identity.

Fetch /api/v1/mcp/config/autonomous-agent. Prefer auth_modes.oauth when the runner supports OAuth + custom headers; use auth_modes.xmemo_key with XMEMO_KEY from a secret store only for fully headless runners.

Run one synthetic task that writes progress to XMemo, restart the runner, and confirm it recalls that progress using the same XMEMO_AGENT_INSTANCE_ID.

If attribution changes or recalls split across instances, persist XMEMO_AGENT_INSTANCE_ID outside git and verify the runner is not regenerating it on every start.
