Recall docs

Recall is a memory layer for AI agents. You send it the messages your agents exchange with people; it works out what is true about each person, keeps track of what changed, and answers questions about them. It speaks Honcho's v3 REST API.

Building an app? Start with the quickstart. Giving a coding agent a memory? Go to integrations.

Quickstart

1. Get an API key. Sign in with your email, open API keys and create a key. It is shown once; keep it in your secret store.

2. Point your environment at Recall.

export RECALL_URL=https://your-recall-url
export RECALL_API_KEY=pk_live_...

3. Record a conversation. Workspaces, sessions and peers are created on first use.

curl $RECALL_URL/v3/workspaces/my-app/sessions/chat-1/messages \
  -H "Authorization: Bearer $RECALL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"messages":[
        {"peer_id":"alice","content":"I moved to Utrecht last month and I am vegetarian now."},
        {"peer_id":"assistant","content":"Noted! Utrecht has great veggie food."}
      ]}'

4. Ask about the person. Learning runs in the background and usually finishes within seconds; GET /v3/workspaces/my-app/queue/status shows progress.

curl $RECALL_URL/v3/workspaces/my-app/peers/alice/chat \
  -H "Authorization: Bearer $RECALL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"query":"What should I cook for Alice?","reasoning_level":"low"}'

{"content":"Something vegetarian — she switched to a vegetarian diet...",
 "citations":["..."],"confidence":0.9,"abstained":false}

5. Look at what it learned. Open the dashboard: Memory lists the facts Recall concluded about Alice, with when each held and the messages it came from.

Concepts

ThingWhat it is
WorkspaceAn isolated memory space, usually one per app or environment.
PeerAnyone who speaks: a user, an agent, a bot. Use one stable id per real person or agent.
SessionOne coherent conversation between peers. Messages belong to a session.
ConclusionA fact Recall derived about a peer, as seen by an observer, with when it became true, when it stopped, and the messages it came from.
Peer cardA short, always-up-to-date profile of a peer: the facts worth having in every prompt.
ScopeA group of sessions you want to recall together, e.g. one project or one customer account.
DreamA background pass that consolidates many conclusions into fewer, sharper ones. Included in the price.
Facts are bi-temporal. When Alice says she moved, "lives in Amsterdam" is closed and "lives in Utrecht" starts — at ingest, with relative dates like "last week" anchored to when the message was written.

Coming from Honcho

Recall implements Honcho's v3 REST API, so Honcho's SDKs work against it. Change the base URL and use a Recall key:

// TypeScript — @honcho-ai/sdk
const honcho = new Honcho({
  baseURL: process.env.RECALL_URL,
  apiKey: process.env.RECALL_API_KEY,
  workspaceId: "my-app",
});

# Python — honcho-ai
honcho = Honcho(base_url=os.environ["RECALL_URL"],
                api_key=os.environ["RECALL_API_KEY"],
                workspace_id="my-app")

What you get on top: answers with citations, confidence and abstained; facts that are superseded at ingest instead of later; an Idempotency-Key header on writes; POST …/sessions/{id}/turn to store messages and get context in one request; usage by task at GET /v3/usage; read-only GETs for one workspace, peer or session; and DELETE …/peers/{peer} to forget a person.

Ask questions

POST /v3/workspaces/{workspace}/peers/{peer}/chat answers a question from what the peer knows. Ask about another peer with target, restrict to a conversation with session_id, and stream with "stream": true.

FieldMeaning
queryThe question, in plain language.
reasoning_levelminimal, low (default), medium, high or max. Up to medium: one retrieval and one model call. High and max use a tool loop for hard, multi-step questions.
include_evidenceAlso return the facts and messages the answer used.
response_formatA JSON Schema: the answer comes back as matching JSON.

Answers are cached until the memory they depend on changes; a cached answer is still billed as one query.

Prompt context

GET /v3/workspaces/{workspace}/sessions/{session}/context?tokens=4000&peer_target=alice returns recent messages, a running summary and what is known about the peer, trimmed to your token budget. With search_query it adds the facts most relevant to the current turn. The SDKs turn it into OpenAI or Anthropic messages for you.

SDKs

Recall ships typed SDKs for TypeScript and Python with lazy handles (no request until you use them), idempotent writes, retries with backoff, and prompt builders.

import { Recall } from "@recall-memory/sdk";

const recall = new Recall();             // RECALL_URL, RECALL_API_KEY
const alice = recall.peer("alice");
const session = recall.session("chat-1");

const { context } = await session.turn([alice.message("Hi again!")], {
  peerTarget: alice, tokens: 4000,
});
const messages = context.toOpenAI(recall.peer("assistant"));

Install. The packages are on their way to npm and PyPI. Until then, install the Python SDK from GitHub, or use Honcho's published SDKs (see above), which work unchanged:

pip install "git+https://github.com/RedSix6/recallmemory#subdirectory=sdks/python"

Both SDKs also have peer.delete() to forget a person.

Command line

recall is the command-line client for a Recall server, and it starts one too. It ships with the server package: it is on the PATH in the Docker image, and from a checkout you build it with pnpm --filter @recall-memory/server build and run node packages/server/dist/cli.js. Save the connection once, then look at and edit memory:

recall init --url $RECALL_URL --api-key $RECALL_API_KEY -w my-app -p alice
recall doctor --url $RECALL_URL        # check the key, workspace, queue and peer
recall peer card                       # what Recall knows about you
recall peer chat "Where do I live?"
recall session messages <id> --last 10
recall conclusion create "Prefers tabs over spaces"
recall session upload notes.pdf <id> -p alice   # a file becomes messages
recall queue status                    # still learning from recent messages?
recall peer delete alice --yes         # forget a person
GroupCommands
workspacelist, create, inspect, delete, chat, search, queue-status
peerlist, create, inspect, card, chat, search, representation, get-metadata, set-metadata, dream, delete
sessionlist, create, inspect, messages, context, summaries, search, representation, peers, add-peers, remove-peers, get-metadata, set-metadata, upload, delete
scopelist, create, inspect, sessions, status, add-sessions, remove-session
messagelist, create, get
conclusionlist, search, get, derived, create, delete
queuestatus

Settings come from flags, then RECALL_URL, RECALL_API_KEY, RECALL_WORKSPACE_ID and RECALL_PEER_ID, then ~/.recall/config.json; Honcho's HONCHO_* variables and ~/.honcho/config.json are read as fallbacks. recall config shows each value and where it came from. Every command takes --json, and commands that delete need --yes.

recall doctor prints the configuration a server would start with (secrets are never shown): the database, which provider keys are set and the model each task routes to. With --url it also checks a running server the way a client would: /health, your API key, the workspace and its background queue, and the peer. It only reads, and exits 1 when a check fails. recall start, token, accounts and keys are for running a server.

Integrations

Recall plugs into coding agents and assistants so they remember you across sessions and projects. An integration recalls (adds your peer card and the facts Recall has learned about you to the agent's context), records (stores your messages and the agent's final replies, so Recall keeps learning) and, where the agent supports it, adds tools. The agent is a peer that Recall does not observe, so facts are learned about you, not from the agent's own statements.

AgentConnects throughMemory is added
Claude CodePlugin: hooks and MCPat session start
CodexPlugin or installer: hooks and MCPat session start
OpenClawPluginbefore each turn
Hermes AgentMemory providerbefore each turn
Agent SkillA SKILL.md the agent readswhen the agent decides to look
Any MCP clientMCP serverwhen the model calls a tool

Connect to your server. Every integration needs the URL of your Recall server and an API key, and optionally a workspace and a peer id for you. Set RECALL_URL, RECALL_API_KEY, RECALL_WORKSPACE_ID and RECALL_PEER_ID, or save them once with recall init, which writes ~/.recall/config.json. With neither, an integration uses http://localhost:8000, the workspace default and your OS user name as your peer id, which suits a local server (recall start) that needs no key. Keep the peer id stable: it is how Recall knows you.

Recording sends your prompts and the agent's final replies, each cut to 24,000 characters, to the Recall server you configured. Tool calls and their output are not sent. The plugins and the Hermes provider can stop recording (RECALL_RECORD=false).

If Recall is unreachable, the plugins and the Hermes provider fail quietly: they skip it for a minute, keep what could not be sent, and deliver it in order once Recall is back. The agent behaves as it would without them.

Claude Code

The plugin recalls when a session starts, resumes or is compacted, records each prompt and Claude's final reply in a Recall session named cc-<claude session id>, and adds Recall's MCP tools. It needs Claude Code with plugin support and Node 18+.

/plugin marketplace add RedSix6/recallmemory
/plugin install recall-memory@recall

From a shell, claude plugin marketplace add RedSix6/recallmemory and claude plugin install recall-memory@recall do the same. Then tell it where Recall is, by running the CLI once:

recall init --url https://your-recall-url --api-key pk_live_... -w my-workspace -p your-name

or with the environment variables above. Restart Claude Code (or run /reload-plugins) and run /recall-memory:status: it prints the settings in use and whether the server answers and accepts your key. /mcp shows whether the recall tools connected.

Setting (default)What it does
RECALL_LOOKUP (session)session adds memory at session start; prompt also looks up facts related to each prompt; off turns lookup off.
RECALL_RECORD (true)false stops recording; lookup still works.
RECALL_CONTEXT_TOKENS (1500)Size of the memory added to the session.
RECALL_MCP_PROFILE (default)full exposes all 40 MCP tools instead of the curated eight.
RECALL_SESSION_STRATEGY (per-session)per-directory records every session of a project into one Recall session.
RECALL_ENABLED (true)false turns the hooks off. Set it under env in a project's .claude/settings.json to switch Recall off for that project.

Check it. Send Claude a prompt, then look at what was recorded; in a new session ask "what do you know about me?".

recall session list
recall session messages cc-<session id> --last 2

Hooks never print errors into the session; they log to hooks.log in the plugin's data directory (RECALL_DEBUG=1 also prints them to stderr). Full reference: integrations/claude-code.

Codex

The Codex integration uses Codex hooks (SessionStart, UserPromptSubmit, Stop) and a stdio MCP server: it adds your peer card and facts at session start (also after /clear or a compaction), records each prompt and Codex's final reply in a session named codex-<codex session id>, and adds the same eight MCP tools. It needs Codex CLI with hooks (tested with 0.162.0) and Node 18+. Pick one of the two routes; installing both handles every turn twice.

As a Codex plugin:

codex plugin marketplace add RedSix6/recallmemory
codex plugin add recall-memory@recall

Start Codex and run /hooks: Codex skips new hooks until you trust them, so trust the three Recall hooks.

With the installer (no plugin system needed):

git clone https://github.com/RedSix6/recallmemory
node recallmemory/integrations/codex/install.mjs \
  --url https://your-recall-url --api-key pk_live_... --workspace my-workspace --peer your-name

It copies the scripts to ~/.codex/recall/scripts, adds the three hooks to ~/.codex/hooks.json and the recall MCP server to ~/.codex/config.toml (in a marked block; your other settings are untouched), and saves the connection to ~/.recall/config.json. Run it again to update, or with --uninstall to remove everything it added; then start Codex and trust the hooks with /hooks. Other options: --codex-home DIR, --no-hooks, --no-mcp, --skill (also installs the Agent Skill) and --agents-md (adds a short "Recall memory" section to ~/.codex/AGENTS.md).

The settings are the same as for Claude Code, except that RECALL_ASSISTANT_PEER defaults to codex. Codex's hook output limit is about 2,500 tokens, so keep RECALL_CONTEXT_TOKENS at or below that.

Check it.

codex mcp list                                # shows the recall server
codex exec "I prefer pnpm over npm"           # prints "session id: <id>"
recall session messages codex-<id> --last 2   # your prompt and Codex's reply

If nothing is recorded, the hooks are probably not trusted yet: run /hooks in Codex. Full reference: integrations/codex.

OpenClaw

A plugin that adds what Recall knows about you, focused on the message at hand, before each turn (in a <recall-memory> block marked as background information, not instructions), records your message and the assistant's reply after each turn in a session named oc-<session key>, and gives the assistant four tools: recall_search, recall_ask, recall_remember and recall_forget. It runs next to OpenClaw's own memory and does not take the memory slot. It needs OpenClaw 2026.9.9 or later and is installed from a checkout, because it is not on ClawHub or npm yet.

git clone https://github.com/RedSix6/recallmemory
openclaw plugins install ./recallmemory/integrations/openclaw --force --accept-capabilities

Both hooks read the conversation, which OpenClaw only allows a third-party plugin to do once you grant it. Then say where Recall is:

openclaw config set plugins.entries.recall-memory.hooks '{"allowConversationAccess":true}' --strict-json --merge
openclaw config set plugins.entries.recall-memory.config \
  '{"url":"https://your-recall-url","workspace":"my-workspace","peer":"your-name"}' --strict-json --merge

Put the API key in RECALL_API_KEY in the Gateway's environment, or in ~/.recall/config.json, rather than in openclaw.json. Restart the Gateway, or run openclaw plugins reload recall-memory. The plugin's lookup setting chooses when memory is added: turn (default) on every user turn, first-turn once per session, off for tools only. Heartbeat and cron turns are not recorded unless you set recordAutomated.

Check it.

openclaw plugins inspect recall-memory --runtime --json      # lists the four recall_* tools
openclaw agent --local -m "I moved to Utrecht last month" --session-key agent:main:recall-check
recall session messages oc-agent-main-recall-check --last 2   # your message and the reply

Full reference: integrations/openclaw.

Hermes Agent

A memory provider for Hermes Agent. Hermes' own MEMORY.md and USER.md memory keeps working; Recall is the one external provider next to it. Before each turn it returns your peer card and the facts most related to the message, after each turn it stores the exchange in a session named hermes-<session id>, it copies the facts Hermes' built-in memory saves about you, and it adds the tools recall_search, recall_ask, recall_remember and recall_forget. It uses only the Python standard library.

git clone https://github.com/RedSix6/recallmemory
mkdir -p ~/.hermes/plugins
cp -r recallmemory/integrations/hermes/recall ~/.hermes/plugins/recall
hermes memory setup        # pick "recall" and enter the URL, API key, workspace and your peer id

hermes memory setup writes the API key to ~/.hermes/.env (RECALL_API_KEY) and the rest to ~/.hermes/recall.json. Instead of the wizard you can set memory.provider: recall in ~/.hermes/config.yaml.

Check it.

hermes memory status                               # "recall ... ← active"
hermes chat -q "I moved to Utrecht last month" -Q  # prints session_id: <id>
recall session messages hermes-<id> --last 2       # your message and the reply

Then, in a new chat, ask "where do I live?". Full reference: integrations/hermes.

Agent Skill

A portable skill in the Agent Skills format: a folder with a SKILL.md that teaches any agent that supports skills when and how to use Recall. It looks up what is known about the user, saves facts they state, forgets on request, and records conversations when nothing else does. It works through Recall's MCP tools when they are connected, and otherwise through scripts/recall.mjs, a dependency-free Node 18+ helper. Nothing happens automatically: the agent reads the skill when a task calls for memory and decides what to call. Use it alone or next to an integration; the skill tells the agent not to record messages itself when an integration already does.

Copy the recall-memory folder to where your agent looks for skills:

git clone https://github.com/RedSix6/recallmemory
mkdir -p ~/.agents/skills && cp -r recallmemory/integrations/agent-skill/recall-memory ~/.agents/skills/
AgentPersonal skills folder
Claude Code~/.claude/skills/
Codex~/.agents/skills/, or run the Codex installer with --skill
OpenClaw~/.agents/skills/, or <workspace>/skills/ for one agent
Hermes Agent~/.hermes/skills/, or add ~/.agents/skills to its external skill directories

The skill needs no configuration when the agent has Recall's MCP tools. For the helper script, set the usual variables or run recall init once. Check it:

node ~/.agents/skills/recall-memory/scripts/recall.mjs status        # health ok, access accepted
node ~/.agents/skills/recall-memory/scripts/recall.mjs remember "Prefers tea over coffee"
node ~/.agents/skills/recall-memory/scripts/recall.mjs context --query drinks

Then ask your agent "what do you remember about my drinks?". Full reference: integrations/agent-skill.

MCP

Recall serves the Model Context Protocol at https://your-recall-url/mcp (streamable HTTP, stateless). Authenticate with Authorization: Bearer pk_live_…; OAuth is not supported yet. By default it exposes eight tools for recording and recalling: chat, search, get_session_context, get_peer_card, list_conclusions, create_conclusions, delete_conclusion and add_messages_to_session. Send x-recall-mcp-profile: full (or use /mcp?profile=full) for all 40 tools (workspaces, peers, sessions, conclusions, scopes), which have Honcho's tool and argument names. Defaults for the workspace, your peer id and the session come from the headers X-Recall-Workspace-ID, X-Recall-User-Name and X-Recall-Session-ID.

{
  "mcpServers": {
    "recall": {
      "type": "http",
      "url": "https://your-recall-url/mcp",
      "headers": { "Authorization": "Bearer pk_live_..." }
    }
  }
}

Webhooks

Register an endpoint with POST /v3/workspaces/{workspace}/webhooks and a body of {"url": "https://…"}. Recall sends a queue.empty event when background work for a session finishes; GET …/webhooks/test sends a test.event. A workspace can have 10 endpoints, and URLs that point at private addresses are refused.

POST https://your-endpoint.example/recall
X-Honcho-Signature: <hex HMAC-SHA256 of the body>
X-Recall-Event-Id: <event id>

{"data":{"observed":"alice","observer":"alice","queue_type":"representation",
 "session_id":"chat-1","workspace_id":"my-app"},
 "timestamp":"2026-10-09T10:00:00Z","type":"queue.empty"}

Dashboard

The dashboard at /app is where you see and manage what Recall holds for your account. Sign in with your email; there is no password.

PageWhat it does
OverviewA getting-started checklist (create a key, send messages, ask a question) that ticks itself off from your account's data, spend this month, and charts of messages and questions per day.
MemoryThe memory explorer. Browse workspaces, then peers and sessions. A peer shows its card and its facts: Current or All, so superseded facts stay visible as history, each with when it held, the messages it came from and what replaced it. A session shows the transcript, its summary and a free context preview (no model calls).
PlaygroundPick a workspace and a peer, ask a question at any reasoning level, and see the evidence behind the answer. It runs the same engine, billing and spend caps as POST …/peers/{peer}/chat.
API keysCreate keys, optionally scoped to a workspace, peer or session, and revoke them.
UsageThe last 30 days: model calls, tokens or charges per day, and a table by day.
Billing and SettingsCredit, top-ups and invoices; account name, monthly budget, webhook secret, export and deletion of your data.

REST API

Base URL https://your-recall-url. Every request carries Authorization: Bearer <key>; bodies are JSON. Writes accept an Idempotency-Key header. In the table, … stands for /v3/workspaces/{workspace}.

EndpointDoes
POST /v3/workspaces · /listGet or create a workspace · list them
GET /v3/workspaces/{w}RecallRead one workspace; never creates
PUT /v3/workspaces/{w}/peers/{p}Get or create a peer (metadata, configuration)
GET …/peers/{p}RecallRead one peer; never creates
DELETE …/peers/{p}RecallForget a person
GET|PUT …/peers/{p}/cardRead or set the peer card
PUT …/sessions/{s} · …/peersGet or create a session · manage who is in it
GET …/sessions/{s}RecallRead one session; never creates
POST …/sessions/{s}/messagesAdd messages
POST …/sessions/{s}/messages/listPage through messages, with filters
POST …/sessions/{s}/turnRecallAdd messages and get the next context in one call
POST …/peers/{p}/chat · POST …/chatAsk about a peer · ask across the workspace
GET …/sessions/{s}/context · GET …/peers/{p}/contextPrompt-ready context
POST …/search · …/peers/{p}/search · …/sessions/{s}/searchHybrid search over messages
POST …/conclusions/list · /query · POST …/conclusionsList, semantically query or add facts
GET …/queue/statusBackground work still pending
POST …/schedule_dreamConsolidate memory now
GET /v3/usage · GET …/usageRecallUsage and charges for the account · a workspace

The Recall tag marks what Honcho's API does not have.

Read without creating

POST /v3/workspaces, PUT …/peers/{peer} and PUT …/sessions/{session} get or create, so a typo in an id quietly makes something new. To only look, use GET:

GET /v3/workspaces/{workspace}
GET /v3/workspaces/{workspace}/peers/{peer}
GET /v3/workspaces/{workspace}/sessions/{session}

Each returns the same object as the get-or-create call and never creates anything. A workspace, peer or session that does not exist answers 404; the hidden peer id of a scope answers 422.

Search filters

The three search routes take query, limit (1 to 100, default 10) and filters. Besides field and metadata filters, filters accepts peer_perspective: a peer id. Results are then limited to what that peer saw, meaning messages in sessions it belongs to (or belonged to) that were written while it was a member, plus everything it wrote itself.

curl $RECALL_URL/v3/workspaces/my-app/search \
  -H "Authorization: Bearer $RECALL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"query":"tent","filters":{"peer_perspective":"bob"},"limit":10}'

peer_perspective must be a non-empty peer id, otherwise the request is a 422. A peer that does not exist saw nothing, so the result is empty.

Forget a person

DELETE /v3/workspaces/{workspace}/peers/{peer} erases a person, for example when they ask you to under GDPR. It removes:

It keeps facts about other peers that merely mention the person, and messages other people wrote. Delete those separately if you need them gone: DELETE …/conclusions/{id}, or the session with DELETE …/sessions/{session}.

curl -X DELETE $RECALL_URL/v3/workspaces/my-app/peers/alice \
  -H "Authorization: Bearer $RECALL_API_KEY"

{"peer_id":"alice","deleted":true,"messages_deleted":12,
 "conclusions_deleted":9,"sessions_affected":2}

The response says what was removed (the numbers here are examples). The call is permanent. It needs a key for the whole workspace or account: a key scoped to one peer or one session answers 401. A peer that does not exist answers 404, so if the response to a retry is a 404, the peer is already gone. The SDKs and the CLI do the same:

await recall.peer("alice").delete();   // TypeScript
recall.peer("alice").delete()          # Python
recall peer delete alice --yes         # CLI; --yes confirms

The TypeScript SDK returns { peerId, messagesDeleted, conclusionsDeleted, sessionsAffected }; the Python SDK returns the response as a dict.

Browser access

By default any origin may call the API, because it authenticates with bearer keys, not cookies. If you run Recall yourself, restrict it with RECALL_CORS_ORIGINS, a comma-separated list of origins; unset, or *, allows any:

RECALL_CORS_ORIGINS=https://app.example.com,https://admin.example.com

Code that runs in a browser should hold a key scoped to one workspace, peer or session (set the scope when you create it under API keys), never an account-wide key.

Errors and limits

Errors are JSON: {"detail": "…"}.

StatusMeaning
401Missing or invalid key, or a key whose scope does not cover this resource. Recall never answers 403.
402Out of credit or over your monthly cap. Messages are still stored; learning and chat resume after a top-up.
404 · 409 · 422Not found · conflict · invalid input (the message says which field)
429Rate limited; retry after the Retry-After header

Models and providers

Recall sends each model task to a provider/model reference you choose, so any model can serve any task. These are server settings, set as environment variables when you run Recall yourself. With no provider key at all, Recall runs on an offline mock model, so you can try the API.

ProviderKey and model reference
OpenAIOPENAI_API_KEY
openai/<model>
AnthropicANTHROPIC_API_KEY
anthropic/<model>
GeminiGEMINI_API_KEY
gemini/<model>
OpenRouterOPENROUTER_API_KEY
openrouter/<vendor>/<model>
GroqGROQ_API_KEY
groq/<model>
DeepSeekDEEPSEEK_API_KEY
deepseek/<model>
TogetherTOGETHER_API_KEY
together/<model>, or together-dedicated/<project>/<endpoint> for a fine-tune
FireworksFIREWORKS_API_KEY
fireworks/accounts/<account>/models/<model>
Any OpenAI-compatible endpoint (Ollama, vLLM, LM Studio)RECALL_COMPAT_BASE_URL and RECALL_COMPAT_API_KEY
compat/<model>

Set a model per task. Each can come from a different provider:

VariableTask
RECALL_MODEL_REASONERLearning: extracting facts from messages in the background
RECALL_MODEL_SUMMARYSession summaries
RECALL_MODEL_CARDPeer cards
RECALL_MODEL_DREAMDreams (consolidation)
RECALL_MODEL_CHAT_MINIMAL to RECALL_MODEL_CHAT_MAXChat at each reasoning level: MINIMAL, LOW, MEDIUM, HIGH, MAX
RECALL_EMBEDDING_MODEL, RECALL_EMBEDDING_DIMENSIONSEmbeddings (1536 dimensions by default)
RECALL_DEFAULT_PROVIDERanthropic, openai or gemini: the provider for every task you have not set

A task you leave unset gets the cheapest suitable model among the providers that have a key (Anthropic, OpenAI or Gemini), unless you pin one with RECALL_DEFAULT_PROVIDER. recall doctor prints the model each task will use.

OPENROUTER_API_KEY=sk-or-...
ANTHROPIC_API_KEY=sk-ant-...
RECALL_MODEL_REASONER=openrouter/openai/gpt-oss-120b
RECALL_MODEL_CHAT_HIGH=anthropic/claude-sonnet-5-5
RECALL_MODEL_CHAT_MAX=anthropic/claude-opus-5-5

OpenRouter: fallbacks and price-sorted routing

RECALL_OPENROUTER_ROUTING is a JSON object keyed by task, plus default, which applies to every task and is overridden field by field by the task's own entry:

RECALL_OPENROUTER_ROUTING='{
  "default":  {"sort": "price", "data_collection": "deny"},
  "reasoner": {"fallbacks": ["qwen/qwen3.5-9b"], "allow_fallbacks": true}
}'
FieldMeaning
fallbacksModels tried in order when the first is down or refuses.
sortprice (cheapest endpoint first), throughput or latency.
data_collection, zdrdeny allows only providers that do not store or train on prompts; zdr only zero-data-retention endpoints.
max_priceA ceiling in USD per 1M tokens.
allow_fallbacks, order, only, ignore, require_parametersProvider selection, as in OpenRouter's request body.

Model suffixes pass through: openrouter/google/gemma-4-26b-a4b-it:floor sorts that model's providers by price, :nitro by throughput.

Hosted mode (RECALL_HOSTED=true, for customer data). Every OpenRouter request is sent with data_collection: "deny", whatever the routing says. :free models are refused, at startup when you configure one and per call otherwise, because free endpoints are often served by providers that log or train on prompts. RECALL_OPENROUTER_ALLOW_FREE=true overrides it.

Cost metering

Every model and embedding call is recorded with its tokens and cost, by workspace, task and model, at GET /v3/workspaces/{id}/usage, GET /v3/usage and in the dashboard. OpenRouter reports what it charged for each call and Recall records exactly that. For other providers the cost is the call's tokens at list price; override prices with RECALL_PRICING_JSON. OpenRouter models missing from Recall's price table are priced from OpenRouter's public catalog (turn that off with RECALL_OPENROUTER_LIVE_PRICES=false).

Pricing

Prepaid credit, no subscription. New accounts get $5. Learning from messages costs $1.75 per 1M message tokens; chat costs $0.0005 (minimal), $0.003 (low), $0.015 (medium), $0.03 (high) or $0.15 (max) per query. Context, search, storage, summaries and dreams are included. Set a monthly cap and auto top-up under Billing.

Benchmarks

We ran Recall and a self-hosted Honcho side by side on 75 held-out questions from LoCoMo (15 per conversation, single-hop, multi-hop, temporal and open-domain), with every model call in both systems on gpt-6-luna and the same judge (gpt-4.1). Token use was measured at the provider by a metering proxy.

RecallHoncho
Accuracy, strict judge60.0%60.0%
Accuracy, lenient judge, LoCoMo mix79.2%79.8%
Median answer time1.6 s4.1 s
Model spend for the run$0.127$0.198
Bill for the run at list prices$0.38$0.92

Published LoCoMo scores are usually higher because of lenient judges and a question mix weighted to easy categories; compare like with like. The harness is in evals/.

LongMemEval-S and BEAM. The same open harness also runs LongMemEval-S and BEAM (the 100K, 500K, 1M and 10M splits, scored with BEAM's own metric) against any Honcho-compatible URL. We have not run them yet, so there are no results to show. From a checkout, you can see what a run would cost without sending anything:

pnpm --filter @recall-memory/evals download --dataset beam --beam-split 100k
pnpm --filter @recall-memory/evals estimate --dataset beam --beam-split 100k

The method, the judge and the offline cost estimates are in docs/BENCHMARKS.md.