# Evaluate & compare

Architectural evaluation criteria, vendor-neutral memory layer checklist, scenario-based selection, reproducible benchmarks, and migration guide.

- Canonical: https://docs.xmemo.dev/docs/guides/evaluate
- Locale: en-US
- Content-Locale: en-US
- Canonical-Content-Digest: 0dc4c982fd162e4faed2b187bd734cf1fbb96e4fa173b80642ac830808a98226
- Edition-Digest: 0eeeb3d71c697917b026f5ddb575e25f032c5c87a61563921de5f121e4763dbb
- Source-Revision: sha256:f430baffad02dd08150d44a87dc0225b25a48bb27d1b6363de2c481e4c753c9b

## Architectural evaluation dimensions

Evaluating agent memory systems requires looking beyond surface vector retrieval to core production architecture: data model, tenancy boundaries, consistency guarantees, protocol compliance, latency, and machine readability.

- Storage model: Hybrid relational metadata with vector embeddings versus vector-only flat stores. Relational grounding ensures exact filter pushdown, immutable audit logs, and deterministic scope partitioning.
- Multi-tenant isolation: Hardware- or server-enforced tenant partitions versus application-level metadata filtering. Shared agent spaces require cryptographically signed space bindings.
- Consistency guarantees: Read-your-writes consistency across distributed tool calls versus eventual consistency windows that cause hallucinations in continuous agent loops.
- Protocol compliance: Native Model Context Protocol (MCP) Streamable HTTP endpoints versus proprietary wrappers requiring per-framework SDK adapters.
- Machine readability: Structured, deterministic JSON responses adhering strictly to tool schemas without prose tail or conversational wrapper text.

## What to evaluate in any memory layer

Evaluate any agent memory system against seven foundational production dimensions: scope isolation, caller attribution, lifecycle deletion, protocol integration, concurrency rate limits, reproducible benchmarks, and plan-dependent features.

- Scope isolation: Ask any vendor how user, session, agent, and workspace memories are partitioned. XMemo enforces a 4-tier orthogonal scope hierarchy (user, session, agent, team/space) with first-class project isolation at the storage boundary (see /docs/concepts/scopes and /docs/concepts/projects).
- Attribution and provenance: Ask any vendor whether every memory write and mutation traces back to a verified caller identity and source event. XMemo records caller attribution, agent instance fingerprints, and provenance metadata in audit logs and preserves version history (superseded versions) with update DAGs (see /docs/concepts/agent-instance-identity and /docs/concepts/memory-model).
- Deletion and export: Ask any vendor how the platform supports tombstone soft-deletion, hard content redaction, and bulk data export. XMemo provides governed lifecycle states (active, archived, deleted), soft-delete with restoration, permanent content redaction, and structured JSON export (see /docs/concepts/memory-deletion and /docs/api/memory).
- MCP and REST integration: Ask any vendor whether the memory layer provides native Model Context Protocol (MCP) support alongside standard REST endpoints. XMemo delivers a native Streamable HTTP MCP server exposing a frozen standard tool catalog alongside authenticated REST endpoints under /v1/ (see /docs/mcp/overview and /docs/api/authentication).
- Rate limits: Ask any vendor how request throughput and concurrency limits are enforced and communicated to clients. XMemo enforces runtime rate limiting governed by the api-authentication layer rather than published static quotas, and clients must not hard-code quota assumptions (see /docs/api/authentication).
- Reproducible benchmarks: Ask any vendor whether retrieval accuracy and latency assertions are independently verifiable with open harnesses and deterministic datasets. XMemo publishes repository-owned benchmark contracts using deterministic golden set datasets (recall_golden_set.jsonl) with immutable run records, distinguishing in-memory evaluation from network I/O (see /docs/guides/evaluate#reproducible-benchmark-methodology).
- Plan-dependent features: Ask any vendor which background processes and advanced capabilities are included in base runtime versus gated by commercial tiers. XMemo includes core memory storage, search, recall, and the 5-action reflect tool in the base runtime, while scheduled background Dream reflection is feature-flagged and plan/entitlement-dependent (presets: manual, daily, or weekly; see /docs/concepts/how-xmemo-works and /docs/concepts/dream-reflection).

```json
{
  "evaluation_checklist": {
    "dimensions": [
      {
        "dimension": "Scope isolation",        "question": "How are user, session, agent, and workspace memories segregated?",        "xmemo_answer": "4-tier orthogonal scope hierarchy (user, session, agent, team/space) enforced at storage boundary (/docs/concepts/scopes)"      },
      {
        "dimension": "Attribution and provenance",        "question": "Can every memory mutation be traced to a specific caller and event?",        "xmemo_answer": "Caller attribution and agent instance fingerprints recorded in audit logs with version history (/docs/concepts/agent-instance-identity)"      },
      {
        "dimension": "Deletion and export",        "question": "Does the system support soft deletion, hard redaction, and data export?",        "xmemo_answer": "Tombstone soft-deletion with restoration, hard redaction, and structured export (/docs/concepts/memory-deletion)"      },
      {
        "dimension": "MCP and REST integration",        "question": "Is the memory layer accessible via native MCP as well as HTTP REST?",        "xmemo_answer": "Native Streamable HTTP MCP server with frozen tool catalog and authenticated REST endpoints (/docs/mcp/overview)"      },
      {
        "dimension": "Rate limits",        "question": "How are request quotas communicated and enforced?",        "xmemo_answer": "Runtime rate limiting governed by api-authentication without hard-coded client quotas (/docs/api/authentication)"      },
      {
        "dimension": "Reproducible benchmarks",        "question": "Are retrieval accuracy and performance metrics independently verifiable?",        "xmemo_answer": "Open repository-owned benchmark harness with deterministic golden set datasets and run traces (/docs/guides/evaluate#reproducible-benchmark-methodology)"      },
      {
        "dimension": "Plan-dependent features",        "question": "Which features are core runtime versus gated by tier entitlements?",        "xmemo_answer": "Core operations and 5-action reflect tool in base runtime; background Dream reflection is plan/entitlement-dependent (/docs/concepts/dream-reflection)"      }
    ]
  }
}
```

## Scenario-based selection guide

Evaluate your workload topology, team structure, and governance requirements to select the appropriate memory architecture pattern.

- Multi-agent workflows across multiple environments: Choose a memory layer with native MCP support, shared project scopes, and cross-client synchronization when orchestrating agents across Claude Code, Cursor, ChatGPT, Codex, and Gemini CLI.
- Regulated and enterprise workspaces: Prioritize architectures providing orthogonal scope hierarchies (user, session, agent, team/space), caller attribution audit logs, and hardware/server-enforced tenant boundaries.
- High-concurrency autonomous loops: Select architectures featuring versioned update trees, optimistic locking, and read-your-writes consistency to prevent conflicting updates and stale memory hallucinations.

## Reproducible benchmark methodology

XMemo provides an open, repository-owned benchmark harness to evaluate retrieval accuracy and governance constraints under controlled conditions. Benchmarks distinguish strictly between internal synthetic evaluation, production latency, and verified public claims.

- Evaluation Methodology: Benchmarks follow frozen method contract docs/reports/T06_REPRODUCIBLE_BENCHMARK_METHOD_CONTRACT_2026-09-16.json with deterministic golden set datasets and recorded verification traces.
- Internal Synthetic Benchmarks: Run hermetically via scripts/benchmark_runner.py using tests/data/recall_golden_set.jsonl (SHA-256 7533FBDC...). These measure in-memory retrieval algorithm quality and ranking logic without network or database I/O.
- Real Production Latency: End-to-end latency includes TLS handshake, network transit, database transactions, and tenant isolation overhead. Production latency varies by hosting topology (cloud SaaS vs self-hosted Docker) and geographical distance, and is not represented by in-memory synthetic figures.
- Public Performance Claims: All reported metrics must cite an immutable benchmark_run_id and raw verifier report under docs/reports/t06_reproducibility/. Unverified vendor claims and synthetic shortcuts are excluded from baseline assertions.
- Failure Mode Analysis: Documented edge failure observed when ambiguous entity names ('config.py' across two distinct project scopes) competed for ranking without explicit project scope qualification; resolved by adding project scope constraints.

```bash
# Execute reproducible retrieval benchmark against golden set
python scripts/benchmark_runner.py \
  --track standard \
  --dataset tests/data/recall_golden_set.jsonl \
  --cache-state COLD_CACHE \
  --output docs/reports/t06_reproducibility/custom_run_raw.json
```

## Migration to XMemo

Migrating from legacy key-value, vector-only, or external memory systems to XMemo follows a four-stage structured transition.

- Stage 1: Export legacy memory items as JSON Lines containing text, user_id, session_id, and created_at timestamps.
- Stage 2: Map identifiers to XMemo scope hierarchy — map user_id to owner, session_id to agent_instance_id or project_id, and categorize facts by durability.
- Stage 3: Ingest via the batch remember tool or REST API (/v1/memories) with appropriate scope and provenance attributes.
- Stage 4: Validate recall accuracy and isolation boundaries using recall_context before enabling agent write permissions.

```python
import requests

def migrate_legacy_record(legacy_item, api_key):
    payload = {
        "content": legacy_item["content"],
        "path": f"migration/{legacy_item.get('id', 'item')}",
        "scope": "project" if legacy_item.get("project_id") else "user",
        "metadata": {"source": "legacy_migration", "original_id": legacy_item.get("id")}
    }
    response = requests.post(
        "https://xmemo.dev/v1/memories",
        headers={"Authorization": f"Bearer {api_key}"},
        json=payload
    )
    response.raise_for_status()
    return response.json()
```

## Known limitations and operational boundaries

To prevent misapplication, understand the operational boundaries and explicit design trade-offs of XMemo.

- Not a general relational database: XMemo is optimized for agent facts, state, and skills. It does not replace application primary databases for arbitrary SQL transactions.
- Not an unstructured file dump: Documents exceeding 32 KB should be indexed in Knowledge Bases rather than stored as single inline memory entries.
- Dream reflection frequency: The reflect API and MCP tool surface does not expose a scheduler; background Dream scheduling is feature-flagged and plan/entitlement-dependent (presets: manual [default], daily at owner-local 03:00, or weekly), and consolidation is not synchronous with writes.
- Rate limits: Hosted endpoints enforce runtime rate limiting governed by api-authentication rather than published static quotas; clients must not hard-code quota assumptions.

## ChatGPT user

Give ChatGPT durable access to your XMemo preferences, project facts, decisions, and TODOs without pasting bearer tokens into a chat.

Connect the hosted XMemo MCP server through the ChatGPT/OpenAI app OAuth flow, then approve the memory:read and memory:write grant for your XMemo account.

Save a synthetic preference or project note, start a new chat, then ask ChatGPT to recall it through XMemo before continuing work.

If OAuth fails or tools do not appear, sign out of the MCP server in the host app, reconnect the XMemo server URL, and retry before creating direct tokens.

## Copilot / Codex developer

Carry repo decisions, coding conventions, bug-fix notes, and task history between IDE and CLI agents.

Use OAuth for VS Code / GitHub Copilot and Gemini CLI when available. For Copilot CLI, Codex, Cursor, or other direct MCP clients, keep XMEMO_KEY in the local environment or secret store and set a stable XMEMO_AGENT_INSTANCE_ID.

Record a codebase decision or bug fix, then ask the next IDE or CLI agent to recall the relevant XMemo context before editing.

If recalls are empty, verify the selected MCP config path, the XMEMO_KEY environment variable for direct clients, and any stale OAuth credential in the host app.

## Team / enterprise pilot owner

Evaluate shared memory with account controls, source attribution, export/delete workflows, and reviewer-safe setup evidence.

Create or enter the protected XMemo workspace, invite approved users, then connect each client through OAuth or a scoped direct credential according to the readiness badges.

Have a pilot member save a synthetic team memory, confirm source attribution in XMemo, then review delete/export and support paths.

If a member cannot connect, check role permissions, OAuth approval, client readiness status, and support guidance before issuing a new token.

## Autonomous agent operator

Let headless or scheduled agents record progress, retrieve prior decisions, and keep a stable non-secret instance identity.

Fetch /api/v1/mcp/config/autonomous-agent. Prefer auth_modes.oauth when the runner supports OAuth + custom headers; use auth_modes.xmemo_key with XMEMO_KEY from a secret store only for fully headless runners.

Run one synthetic task that writes progress to XMemo, restart the runner, and confirm it recalls that progress using the same XMEMO_AGENT_INSTANCE_ID.

If attribution changes or recalls split across instances, persist XMEMO_AGENT_INSTANCE_ID outside git and verify the runner is not regenerating it on every start.
