Recall docs
Recall is a memory layer for AI agents. You send it the messages your agents exchange with people; it works out what is true about each person, keeps track of what changed, and answers questions about them. It speaks Honcho's v3 REST API.
Building an app? Start with the quickstart. Giving a coding agent a memory? Go to integrations.
Quickstart
1. Get an API key. Sign in with your email, open API keys and create a key. It is shown once; keep it in your secret store.
2. Point your environment at Recall.
export RECALL_URL=https://your-recall-url
export RECALL_API_KEY=pk_live_...
3. Record a conversation. Workspaces, sessions and peers are created on first use.
curl $RECALL_URL/v3/workspaces/my-app/sessions/chat-1/messages \
-H "Authorization: Bearer $RECALL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"messages":[
{"peer_id":"alice","content":"I moved to Utrecht last month and I am vegetarian now."},
{"peer_id":"assistant","content":"Noted! Utrecht has great veggie food."}
]}'
4. Ask about the person. Learning runs in the background and usually finishes within seconds; GET /v3/workspaces/my-app/queue/status shows progress.
curl $RECALL_URL/v3/workspaces/my-app/peers/alice/chat \
-H "Authorization: Bearer $RECALL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"query":"What should I cook for Alice?","reasoning_level":"low"}'
{"content":"Something vegetarian — she switched to a vegetarian diet...",
"citations":["..."],"confidence":0.9,"abstained":false}
5. Look at what it learned. Open the dashboard: Memory lists the facts Recall concluded about Alice, with when each held and the messages it came from.
Concepts
| Thing | What it is |
|---|---|
| Workspace | An isolated memory space, usually one per app or environment. |
| Peer | Anyone who speaks: a user, an agent, a bot. Use one stable id per real person or agent. |
| Session | One coherent conversation between peers. Messages belong to a session. |
| Conclusion | A fact Recall derived about a peer, as seen by an observer, with when it became true, when it stopped, and the messages it came from. |
| Peer card | A short, always-up-to-date profile of a peer: the facts worth having in every prompt. |
| Scope | A group of sessions you want to recall together, e.g. one project or one customer account. |
| Dream | A background pass that consolidates many conclusions into fewer, sharper ones. Included in the price. |
Coming from Honcho
Recall implements Honcho's v3 REST API, so Honcho's SDKs work against it. Change the base URL and use a Recall key:
// TypeScript — @honcho-ai/sdk
const honcho = new Honcho({
baseURL: process.env.RECALL_URL,
apiKey: process.env.RECALL_API_KEY,
workspaceId: "my-app",
});
# Python — honcho-ai
honcho = Honcho(base_url=os.environ["RECALL_URL"],
api_key=os.environ["RECALL_API_KEY"],
workspace_id="my-app")
What you get on top: answers with citations, confidence and abstained; facts that are superseded at ingest instead of later; an Idempotency-Key header on writes; POST …/sessions/{id}/turn to store messages and get context in one request; usage by task at GET /v3/usage; read-only GETs for one workspace, peer or session; and DELETE …/peers/{peer} to forget a person.
Ask questions
POST /v3/workspaces/{workspace}/peers/{peer}/chat answers a question from what the peer knows. Ask about another peer with target, restrict to a conversation with session_id, and stream with "stream": true.
| Field | Meaning |
|---|---|
query | The question, in plain language. |
reasoning_level | minimal, low (default), medium, high or max. Up to medium: one retrieval and one model call. High and max use a tool loop for hard, multi-step questions. |
include_evidence | Also return the facts and messages the answer used. |
response_format | A JSON Schema: the answer comes back as matching JSON. |
Answers are cached until the memory they depend on changes; a cached answer is still billed as one query.
Prompt context
GET /v3/workspaces/{workspace}/sessions/{session}/context?tokens=4000&peer_target=alice returns recent messages, a running summary and what is known about the peer, trimmed to your token budget. With search_query it adds the facts most relevant to the current turn. The SDKs turn it into OpenAI or Anthropic messages for you.
SDKs
Recall ships typed SDKs for TypeScript and Python with lazy handles (no request until you use them), idempotent writes, retries with backoff, and prompt builders.
import { Recall } from "@recall-memory/sdk";
const recall = new Recall(); // RECALL_URL, RECALL_API_KEY
const alice = recall.peer("alice");
const session = recall.session("chat-1");
const { context } = await session.turn([alice.message("Hi again!")], {
peerTarget: alice, tokens: 4000,
});
const messages = context.toOpenAI(recall.peer("assistant"));
Install. The packages are on their way to npm and PyPI. Until then, install the Python SDK from GitHub, or use Honcho's published SDKs (see above), which work unchanged:
pip install "git+https://github.com/RedSix6/recallmemory#subdirectory=sdks/python"
Both SDKs also have peer.delete() to forget a person.
Command line
recall is the command-line client for a Recall server, and it starts one too. It ships with the server package: it is on the PATH in the Docker image, and from a checkout you build it with pnpm --filter @recall-memory/server build and run node packages/server/dist/cli.js. Save the connection once, then look at and edit memory:
recall init --url $RECALL_URL --api-key $RECALL_API_KEY -w my-app -p alice
recall doctor --url $RECALL_URL # check the key, workspace, queue and peer
recall peer card # what Recall knows about you
recall peer chat "Where do I live?"
recall session messages <id> --last 10
recall conclusion create "Prefers tabs over spaces"
recall session upload notes.pdf <id> -p alice # a file becomes messages
recall queue status # still learning from recent messages?
recall peer delete alice --yes # forget a person
| Group | Commands |
|---|---|
workspace | list, create, inspect, delete, chat, search, queue-status |
peer | list, create, inspect, card, chat, search, representation, get-metadata, set-metadata, dream, delete |
session | list, create, inspect, messages, context, summaries, search, representation, peers, add-peers, remove-peers, get-metadata, set-metadata, upload, delete |
scope | list, create, inspect, sessions, status, add-sessions, remove-session |
message | list, create, get |
conclusion | list, search, get, derived, create, delete |
queue | status |
Settings come from flags, then RECALL_URL, RECALL_API_KEY, RECALL_WORKSPACE_ID and RECALL_PEER_ID, then ~/.recall/config.json; Honcho's HONCHO_* variables and ~/.honcho/config.json are read as fallbacks. recall config shows each value and where it came from. Every command takes --json, and commands that delete need --yes.
recall doctor prints the configuration a server would start with (secrets are never shown): the database, which provider keys are set and the model each task routes to. With --url it also checks a running server the way a client would: /health, your API key, the workspace and its background queue, and the peer. It only reads, and exits 1 when a check fails. recall start, token, accounts and keys are for running a server.
Integrations
Recall plugs into coding agents and assistants so they remember you across sessions and projects. An integration recalls (adds your peer card and the facts Recall has learned about you to the agent's context), records (stores your messages and the agent's final replies, so Recall keeps learning) and, where the agent supports it, adds tools. The agent is a peer that Recall does not observe, so facts are learned about you, not from the agent's own statements.
| Agent | Connects through | Memory is added |
|---|---|---|
| Claude Code | Plugin: hooks and MCP | at session start |
| Codex | Plugin or installer: hooks and MCP | at session start |
| OpenClaw | Plugin | before each turn |
| Hermes Agent | Memory provider | before each turn |
| Agent Skill | A SKILL.md the agent reads | when the agent decides to look |
| Any MCP client | MCP server | when the model calls a tool |
Connect to your server. Every integration needs the URL of your Recall server and an API key, and optionally a workspace and a peer id for you. Set RECALL_URL, RECALL_API_KEY, RECALL_WORKSPACE_ID and RECALL_PEER_ID, or save them once with recall init, which writes ~/.recall/config.json. With neither, an integration uses http://localhost:8000, the workspace default and your OS user name as your peer id, which suits a local server (recall start) that needs no key. Keep the peer id stable: it is how Recall knows you.
RECALL_RECORD=false).If Recall is unreachable, the plugins and the Hermes provider fail quietly: they skip it for a minute, keep what could not be sent, and deliver it in order once Recall is back. The agent behaves as it would without them.
Claude Code
The plugin recalls when a session starts, resumes or is compacted, records each prompt and Claude's final reply in a Recall session named cc-<claude session id>, and adds Recall's MCP tools. It needs Claude Code with plugin support and Node 18+.
/plugin marketplace add RedSix6/recallmemory
/plugin install recall-memory@recall
From a shell, claude plugin marketplace add RedSix6/recallmemory and claude plugin install recall-memory@recall do the same. Then tell it where Recall is, by running the CLI once:
recall init --url https://your-recall-url --api-key pk_live_... -w my-workspace -p your-name
or with the environment variables above. Restart Claude Code (or run /reload-plugins) and run /recall-memory:status: it prints the settings in use and whether the server answers and accepts your key. /mcp shows whether the recall tools connected.
| Setting (default) | What it does |
|---|---|
RECALL_LOOKUP (session) | session adds memory at session start; prompt also looks up facts related to each prompt; off turns lookup off. |
RECALL_RECORD (true) | false stops recording; lookup still works. |
RECALL_CONTEXT_TOKENS (1500) | Size of the memory added to the session. |
RECALL_MCP_PROFILE (default) | full exposes all 40 MCP tools instead of the curated eight. |
RECALL_SESSION_STRATEGY (per-session) | per-directory records every session of a project into one Recall session. |
RECALL_ENABLED (true) | false turns the hooks off. Set it under env in a project's .claude/settings.json to switch Recall off for that project. |
Check it. Send Claude a prompt, then look at what was recorded; in a new session ask "what do you know about me?".
recall session list
recall session messages cc-<session id> --last 2
Hooks never print errors into the session; they log to hooks.log in the plugin's data directory (RECALL_DEBUG=1 also prints them to stderr). Full reference: integrations/claude-code.
Codex
The Codex integration uses Codex hooks (SessionStart, UserPromptSubmit, Stop) and a stdio MCP server: it adds your peer card and facts at session start (also after /clear or a compaction), records each prompt and Codex's final reply in a session named codex-<codex session id>, and adds the same eight MCP tools. It needs Codex CLI with hooks (tested with 0.162.0) and Node 18+. Pick one of the two routes; installing both handles every turn twice.
As a Codex plugin:
codex plugin marketplace add RedSix6/recallmemory
codex plugin add recall-memory@recall
Start Codex and run /hooks: Codex skips new hooks until you trust them, so trust the three Recall hooks.
With the installer (no plugin system needed):
git clone https://github.com/RedSix6/recallmemory
node recallmemory/integrations/codex/install.mjs \
--url https://your-recall-url --api-key pk_live_... --workspace my-workspace --peer your-name
It copies the scripts to ~/.codex/recall/scripts, adds the three hooks to ~/.codex/hooks.json and the recall MCP server to ~/.codex/config.toml (in a marked block; your other settings are untouched), and saves the connection to ~/.recall/config.json. Run it again to update, or with --uninstall to remove everything it added; then start Codex and trust the hooks with /hooks. Other options: --codex-home DIR, --no-hooks, --no-mcp, --skill (also installs the Agent Skill) and --agents-md (adds a short "Recall memory" section to ~/.codex/AGENTS.md).
The settings are the same as for Claude Code, except that RECALL_ASSISTANT_PEER defaults to codex. Codex's hook output limit is about 2,500 tokens, so keep RECALL_CONTEXT_TOKENS at or below that.
Check it.
codex mcp list # shows the recall server
codex exec "I prefer pnpm over npm" # prints "session id: <id>"
recall session messages codex-<id> --last 2 # your prompt and Codex's reply
If nothing is recorded, the hooks are probably not trusted yet: run /hooks in Codex. Full reference: integrations/codex.
OpenClaw
A plugin that adds what Recall knows about you, focused on the message at hand, before each turn (in a <recall-memory> block marked as background information, not instructions), records your message and the assistant's reply after each turn in a session named oc-<session key>, and gives the assistant four tools: recall_search, recall_ask, recall_remember and recall_forget. It runs next to OpenClaw's own memory and does not take the memory slot. It needs OpenClaw 2026.9.9 or later and is installed from a checkout, because it is not on ClawHub or npm yet.
git clone https://github.com/RedSix6/recallmemory
openclaw plugins install ./recallmemory/integrations/openclaw --force --accept-capabilities
Both hooks read the conversation, which OpenClaw only allows a third-party plugin to do once you grant it. Then say where Recall is:
openclaw config set plugins.entries.recall-memory.hooks '{"allowConversationAccess":true}' --strict-json --merge
openclaw config set plugins.entries.recall-memory.config \
'{"url":"https://your-recall-url","workspace":"my-workspace","peer":"your-name"}' --strict-json --merge
Put the API key in RECALL_API_KEY in the Gateway's environment, or in ~/.recall/config.json, rather than in openclaw.json. Restart the Gateway, or run openclaw plugins reload recall-memory. The plugin's lookup setting chooses when memory is added: turn (default) on every user turn, first-turn once per session, off for tools only. Heartbeat and cron turns are not recorded unless you set recordAutomated.
Check it.
openclaw plugins inspect recall-memory --runtime --json # lists the four recall_* tools
openclaw agent --local -m "I moved to Utrecht last month" --session-key agent:main:recall-check
recall session messages oc-agent-main-recall-check --last 2 # your message and the reply
Full reference: integrations/openclaw.
Hermes Agent
A memory provider for Hermes Agent. Hermes' own MEMORY.md and USER.md memory keeps working; Recall is the one external provider next to it. Before each turn it returns your peer card and the facts most related to the message, after each turn it stores the exchange in a session named hermes-<session id>, it copies the facts Hermes' built-in memory saves about you, and it adds the tools recall_search, recall_ask, recall_remember and recall_forget. It uses only the Python standard library.
git clone https://github.com/RedSix6/recallmemory
mkdir -p ~/.hermes/plugins
cp -r recallmemory/integrations/hermes/recall ~/.hermes/plugins/recall
hermes memory setup # pick "recall" and enter the URL, API key, workspace and your peer id
hermes memory setup writes the API key to ~/.hermes/.env (RECALL_API_KEY) and the rest to ~/.hermes/recall.json. Instead of the wizard you can set memory.provider: recall in ~/.hermes/config.yaml.
Check it.
hermes memory status # "recall ... ← active"
hermes chat -q "I moved to Utrecht last month" -Q # prints session_id: <id>
recall session messages hermes-<id> --last 2 # your message and the reply
Then, in a new chat, ask "where do I live?". Full reference: integrations/hermes.
Agent Skill
A portable skill in the Agent Skills format: a folder with a SKILL.md that teaches any agent that supports skills when and how to use Recall. It looks up what is known about the user, saves facts they state, forgets on request, and records conversations when nothing else does. It works through Recall's MCP tools when they are connected, and otherwise through scripts/recall.mjs, a dependency-free Node 18+ helper. Nothing happens automatically: the agent reads the skill when a task calls for memory and decides what to call. Use it alone or next to an integration; the skill tells the agent not to record messages itself when an integration already does.
Copy the recall-memory folder to where your agent looks for skills:
git clone https://github.com/RedSix6/recallmemory
mkdir -p ~/.agents/skills && cp -r recallmemory/integrations/agent-skill/recall-memory ~/.agents/skills/
| Agent | Personal skills folder |
|---|---|
| Claude Code | ~/.claude/skills/ |
| Codex | ~/.agents/skills/, or run the Codex installer with --skill |
| OpenClaw | ~/.agents/skills/, or <workspace>/skills/ for one agent |
| Hermes Agent | ~/.hermes/skills/, or add ~/.agents/skills to its external skill directories |
The skill needs no configuration when the agent has Recall's MCP tools. For the helper script, set the usual variables or run recall init once. Check it:
node ~/.agents/skills/recall-memory/scripts/recall.mjs status # health ok, access accepted
node ~/.agents/skills/recall-memory/scripts/recall.mjs remember "Prefers tea over coffee"
node ~/.agents/skills/recall-memory/scripts/recall.mjs context --query drinks
Then ask your agent "what do you remember about my drinks?". Full reference: integrations/agent-skill.
MCP
Recall serves the Model Context Protocol at https://your-recall-url/mcp (streamable HTTP, stateless). Authenticate with Authorization: Bearer pk_live_…; OAuth is not supported yet. By default it exposes eight tools for recording and recalling: chat, search, get_session_context, get_peer_card, list_conclusions, create_conclusions, delete_conclusion and add_messages_to_session. Send x-recall-mcp-profile: full (or use /mcp?profile=full) for all 40 tools (workspaces, peers, sessions, conclusions, scopes), which have Honcho's tool and argument names. Defaults for the workspace, your peer id and the session come from the headers X-Recall-Workspace-ID, X-Recall-User-Name and X-Recall-Session-ID.
{
"mcpServers": {
"recall": {
"type": "http",
"url": "https://your-recall-url/mcp",
"headers": { "Authorization": "Bearer pk_live_..." }
}
}
}
Webhooks
Register an endpoint with POST /v3/workspaces/{workspace}/webhooks and a body of {"url": "https://…"}. Recall sends a queue.empty event when background work for a session finishes; GET …/webhooks/test sends a test.event. A workspace can have 10 endpoints, and URLs that point at private addresses are refused.
POST https://your-endpoint.example/recall
X-Honcho-Signature: <hex HMAC-SHA256 of the body>
X-Recall-Event-Id: <event id>
{"data":{"observed":"alice","observer":"alice","queue_type":"representation",
"session_id":"chat-1","workspace_id":"my-app"},
"timestamp":"2026-10-09T10:00:00Z","type":"queue.empty"}
- Verify the signature.
X-Honcho-Signatureis the hex HMAC-SHA256 of the raw request body, keyed with your webhook secret: your account's own secret (rotate it under Settings). The body is compact JSON with sorted keys, so verify the bytes you received. - Drop duplicates.
X-Recall-Event-Idis the same on every retry of an event, so keep the ids you have handled and ignore repeats. - Answer quickly with a 2xx. Recall retries with backoff when your endpoint cannot be reached, times out, or answers 5xx, 408 or 429. Other 3xx and 4xx answers are final.
Dashboard
The dashboard at /app is where you see and manage what Recall holds for your account. Sign in with your email; there is no password.
| Page | What it does |
|---|---|
| Overview | A getting-started checklist (create a key, send messages, ask a question) that ticks itself off from your account's data, spend this month, and charts of messages and questions per day. |
| Memory | The memory explorer. Browse workspaces, then peers and sessions. A peer shows its card and its facts: Current or All, so superseded facts stay visible as history, each with when it held, the messages it came from and what replaced it. A session shows the transcript, its summary and a free context preview (no model calls). |
| Playground | Pick a workspace and a peer, ask a question at any reasoning level, and see the evidence behind the answer. It runs the same engine, billing and spend caps as POST …/peers/{peer}/chat. |
| API keys | Create keys, optionally scoped to a workspace, peer or session, and revoke them. |
| Usage | The last 30 days: model calls, tokens or charges per day, and a table by day. |
| Billing and Settings | Credit, top-ups and invoices; account name, monthly budget, webhook secret, export and deletion of your data. |
REST API
Base URL https://your-recall-url. Every request carries Authorization: Bearer <key>; bodies are JSON. Writes accept an Idempotency-Key header. In the table, … stands for /v3/workspaces/{workspace}.
| Endpoint | Does |
|---|---|
POST /v3/workspaces · /list | Get or create a workspace · list them |
GET /v3/workspaces/{w}Recall | Read one workspace; never creates |
PUT /v3/workspaces/{w}/peers/{p} | Get or create a peer (metadata, configuration) |
GET …/peers/{p}Recall | Read one peer; never creates |
DELETE …/peers/{p}Recall | Forget a person |
GET|PUT …/peers/{p}/card | Read or set the peer card |
PUT …/sessions/{s} · …/peers | Get or create a session · manage who is in it |
GET …/sessions/{s}Recall | Read one session; never creates |
POST …/sessions/{s}/messages | Add messages |
POST …/sessions/{s}/messages/list | Page through messages, with filters |
POST …/sessions/{s}/turnRecall | Add messages and get the next context in one call |
POST …/peers/{p}/chat · POST …/chat | Ask about a peer · ask across the workspace |
GET …/sessions/{s}/context · GET …/peers/{p}/context | Prompt-ready context |
POST …/search · …/peers/{p}/search · …/sessions/{s}/search | Hybrid search over messages |
POST …/conclusions/list · /query · POST …/conclusions | List, semantically query or add facts |
GET …/queue/status | Background work still pending |
POST …/schedule_dream | Consolidate memory now |
GET /v3/usage · GET …/usageRecall | Usage and charges for the account · a workspace |
The Recall tag marks what Honcho's API does not have.
Read without creating
POST /v3/workspaces, PUT …/peers/{peer} and PUT …/sessions/{session} get or create, so a typo in an id quietly makes something new. To only look, use GET:
GET /v3/workspaces/{workspace}
GET /v3/workspaces/{workspace}/peers/{peer}
GET /v3/workspaces/{workspace}/sessions/{session}
Each returns the same object as the get-or-create call and never creates anything. A workspace, peer or session that does not exist answers 404; the hidden peer id of a scope answers 422.
Search filters
The three search routes take query, limit (1 to 100, default 10) and filters. Besides field and metadata filters, filters accepts peer_perspective: a peer id. Results are then limited to what that peer saw, meaning messages in sessions it belongs to (or belonged to) that were written while it was a member, plus everything it wrote itself.
curl $RECALL_URL/v3/workspaces/my-app/search \
-H "Authorization: Bearer $RECALL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"query":"tent","filters":{"peer_perspective":"bob"},"limit":10}'
peer_perspective must be a non-empty peer id, otherwise the request is a 422. A peer that does not exist saw nothing, so the result is empty.
Forget a person
DELETE /v3/workspaces/{workspace}/peers/{peer} erases a person, for example when they ask you to under GDPR. It removes:
- every message the peer wrote, and its memberships in sessions;
- everything memory holds about the peer and from its point of view: facts and peer cards in every collection it observes or is observed in, including what scopes remember of it;
- the summaries of sessions it wrote in. They quote it, so they are dropped and rebuilt from the remaining messages;
- cached answers, which may quote it.
It keeps facts about other peers that merely mention the person, and messages other people wrote. Delete those separately if you need them gone: DELETE …/conclusions/{id}, or the session with DELETE …/sessions/{session}.
curl -X DELETE $RECALL_URL/v3/workspaces/my-app/peers/alice \
-H "Authorization: Bearer $RECALL_API_KEY"
{"peer_id":"alice","deleted":true,"messages_deleted":12,
"conclusions_deleted":9,"sessions_affected":2}
The response says what was removed (the numbers here are examples). The call is permanent. It needs a key for the whole workspace or account: a key scoped to one peer or one session answers 401. A peer that does not exist answers 404, so if the response to a retry is a 404, the peer is already gone. The SDKs and the CLI do the same:
await recall.peer("alice").delete(); // TypeScript
recall.peer("alice").delete() # Python
recall peer delete alice --yes # CLI; --yes confirms
The TypeScript SDK returns { peerId, messagesDeleted, conclusionsDeleted, sessionsAffected }; the Python SDK returns the response as a dict.
Browser access
By default any origin may call the API, because it authenticates with bearer keys, not cookies. If you run Recall yourself, restrict it with RECALL_CORS_ORIGINS, a comma-separated list of origins; unset, or *, allows any:
RECALL_CORS_ORIGINS=https://app.example.com,https://admin.example.com
Code that runs in a browser should hold a key scoped to one workspace, peer or session (set the scope when you create it under API keys), never an account-wide key.
Errors and limits
Errors are JSON: {"detail": "…"}.
| Status | Meaning |
|---|---|
401 | Missing or invalid key, or a key whose scope does not cover this resource. Recall never answers 403. |
402 | Out of credit or over your monthly cap. Messages are still stored; learning and chat resume after a top-up. |
404 · 409 · 422 | Not found · conflict · invalid input (the message says which field) |
429 | Rate limited; retry after the Retry-After header |
Models and providers
Recall sends each model task to a provider/model reference you choose, so any model can serve any task. These are server settings, set as environment variables when you run Recall yourself. With no provider key at all, Recall runs on an offline mock model, so you can try the API.
| Provider | Key and model reference |
|---|---|
| OpenAI | OPENAI_API_KEYopenai/<model> |
| Anthropic | ANTHROPIC_API_KEYanthropic/<model> |
| Gemini | GEMINI_API_KEYgemini/<model> |
| OpenRouter | OPENROUTER_API_KEYopenrouter/<vendor>/<model> |
| Groq | GROQ_API_KEYgroq/<model> |
| DeepSeek | DEEPSEEK_API_KEYdeepseek/<model> |
| Together | TOGETHER_API_KEYtogether/<model>, or together-dedicated/<project>/<endpoint> for a fine-tune |
| Fireworks | FIREWORKS_API_KEYfireworks/accounts/<account>/models/<model> |
| Any OpenAI-compatible endpoint (Ollama, vLLM, LM Studio) | RECALL_COMPAT_BASE_URL and RECALL_COMPAT_API_KEYcompat/<model> |
Set a model per task. Each can come from a different provider:
| Variable | Task |
|---|---|
RECALL_MODEL_REASONER | Learning: extracting facts from messages in the background |
RECALL_MODEL_SUMMARY | Session summaries |
RECALL_MODEL_CARD | Peer cards |
RECALL_MODEL_DREAM | Dreams (consolidation) |
RECALL_MODEL_CHAT_MINIMAL to RECALL_MODEL_CHAT_MAX | Chat at each reasoning level: MINIMAL, LOW, MEDIUM, HIGH, MAX |
RECALL_EMBEDDING_MODEL, RECALL_EMBEDDING_DIMENSIONS | Embeddings (1536 dimensions by default) |
RECALL_DEFAULT_PROVIDER | anthropic, openai or gemini: the provider for every task you have not set |
A task you leave unset gets the cheapest suitable model among the providers that have a key (Anthropic, OpenAI or Gemini), unless you pin one with RECALL_DEFAULT_PROVIDER. recall doctor prints the model each task will use.
OPENROUTER_API_KEY=sk-or-...
ANTHROPIC_API_KEY=sk-ant-...
RECALL_MODEL_REASONER=openrouter/openai/gpt-oss-120b
RECALL_MODEL_CHAT_HIGH=anthropic/claude-sonnet-5-5
RECALL_MODEL_CHAT_MAX=anthropic/claude-opus-5-5
OpenRouter: fallbacks and price-sorted routing
RECALL_OPENROUTER_ROUTING is a JSON object keyed by task, plus default, which applies to every task and is overridden field by field by the task's own entry:
RECALL_OPENROUTER_ROUTING='{
"default": {"sort": "price", "data_collection": "deny"},
"reasoner": {"fallbacks": ["qwen/qwen3.5-9b"], "allow_fallbacks": true}
}'
| Field | Meaning |
|---|---|
fallbacks | Models tried in order when the first is down or refuses. |
sort | price (cheapest endpoint first), throughput or latency. |
data_collection, zdr | deny allows only providers that do not store or train on prompts; zdr only zero-data-retention endpoints. |
max_price | A ceiling in USD per 1M tokens. |
allow_fallbacks, order, only, ignore, require_parameters | Provider selection, as in OpenRouter's request body. |
Model suffixes pass through: openrouter/google/gemma-4-26b-a4b-it:floor sorts that model's providers by price, :nitro by throughput.
RECALL_HOSTED=true, for customer data). Every OpenRouter request is sent with data_collection: "deny", whatever the routing says. :free models are refused, at startup when you configure one and per call otherwise, because free endpoints are often served by providers that log or train on prompts. RECALL_OPENROUTER_ALLOW_FREE=true overrides it.Cost metering
Every model and embedding call is recorded with its tokens and cost, by workspace, task and model, at GET /v3/workspaces/{id}/usage, GET /v3/usage and in the dashboard. OpenRouter reports what it charged for each call and Recall records exactly that. For other providers the cost is the call's tokens at list price; override prices with RECALL_PRICING_JSON. OpenRouter models missing from Recall's price table are priced from OpenRouter's public catalog (turn that off with RECALL_OPENROUTER_LIVE_PRICES=false).
Pricing
Prepaid credit, no subscription. New accounts get $5. Learning from messages costs $1.75 per 1M message tokens; chat costs $0.0005 (minimal), $0.003 (low), $0.015 (medium), $0.03 (high) or $0.15 (max) per query. Context, search, storage, summaries and dreams are included. Set a monthly cap and auto top-up under Billing.
Benchmarks
We ran Recall and a self-hosted Honcho side by side on 75 held-out questions from LoCoMo (15 per conversation, single-hop, multi-hop, temporal and open-domain), with every model call in both systems on gpt-6-luna and the same judge (gpt-4.1). Token use was measured at the provider by a metering proxy.
| Recall | Honcho | |
|---|---|---|
| Accuracy, strict judge | 60.0% | 60.0% |
| Accuracy, lenient judge, LoCoMo mix | 79.2% | 79.8% |
| Median answer time | 1.6 s | 4.1 s |
| Model spend for the run | $0.127 | $0.198 |
| Bill for the run at list prices | $0.38 | $0.92 |
Published LoCoMo scores are usually higher because of lenient judges and a question mix weighted to easy categories; compare like with like. The harness is in evals/.
LongMemEval-S and BEAM. The same open harness also runs LongMemEval-S and BEAM (the 100K, 500K, 1M and 10M splits, scored with BEAM's own metric) against any Honcho-compatible URL. We have not run them yet, so there are no results to show. From a checkout, you can see what a run would cost without sending anything:
pnpm --filter @recall-memory/evals download --dataset beam --beam-split 100k
pnpm --filter @recall-memory/evals estimate --dataset beam --beam-split 100k
The method, the judge and the offline cost estimates are in docs/BENCHMARKS.md.