Skip to main content
The self-hosted server aims for zero configuration — the only required input is one LLM provider key, which the first-boot wizard collects interactively (or set via env var for non-interactive deployments). Embeddings default to local English; you can pick another provider in the optional wizard step or via env. Everything else below is opt-in. The installer writes API keys to ~/.supermemory/env, which is loaded on every launch. You can also set variables in your shell or a process manager.

Core

The currently published v0.0.8 binds all interfaces and its implicit local authentication is unsafe when exposed to untrusted networks. Do not expose it; restrict access with a firewall or run it on an isolated machine. The next release will bind loopback and require the generated key for every API request, including requests from localhost. Its browser welcome page will no longer reveal the key; enter the key printed at first boot in the Memory tab. If you later expose that release through a reverse proxy, protect the endpoint with TLS and access controls.

LLM providers

In production, Supermemory uses its own proprietary models tuned for long-horizon data understanding. Self-hosted, you bring your own LLM for the intelligent steps — summaries, contextual chunking, and memory extraction. Embeddings default to a local model (no API key) and can optionally use OpenAI, Gemini, or Ollama — see Embeddings. Configure at least one LLM provider:
No key set? The server walks you through it. On first boot, an interactive setup wizard asks which provider you want, securely prompts for the key, and saves it encrypted — including a custom base URL and model name if you pick an OpenAI-compatible endpoint.
With multiple providers configured, the first one in the order above is used.
Image, video, and high-fidelity PDF understanding require a Gemini or Vertex AI key. Text ingestion, memory extraction, and search work with any provider.

Fully offline with local models

OPENAI_API_KEY + OPENAI_BASE_URL covers any OpenAI-compatible endpoint: Ollama, LM Studio, vLLM, llama.cpp server, Together, Fireworks, and more.

File storage

Nothing to configure. Uploaded files (PDFs, images) are stored on local disk inside $SUPERMEMORY_DATA_DIR and served by the server at /files/:key.

Embeddings

By default, vectors are computed locally with Xenova/bge-base-en-v1.5 (768d) — no embedding API key. On interactive first boot you can pick a different provider after the LLM key step; for Docker/CI set env vars instead. The full provider table, multilingual guidance, remote examples (OpenAI, Gemini, Ollama) and the dimension-lock warning are in embeddings for self-hosting.

Embedding performance

Local embeddings are prewarmed at startup with conservative defaults — one worker, minimal CPU footprint. Turn these up if you’re ingesting heavily and prefer throughput over headroom (remote embedding providers ignore these — there’s no local worker pool to tune):

Memory limits & ingestion queue

The server manages memory for you and separates the two kinds of work you send it:
  • Searches are always served immediately. They never wait behind ingestion, regardless of how much is queued.
  • Adds are accepted instantly but processed through a queue. A POST /v3/documents call returns in milliseconds with status queued; extraction, embedding, and indexing happen in the background at a controlled pace.
Ingestion may grow the server’s memory usage by at most SUPERMEMORY_EMBEDDING_RAM_LIMIT (default 1 GB) above its post-boot baseline. Past that, new documents simply wait in the queue until memory drops back under the limit — nothing is dropped, ingestion just slows down. The limit is measured above the boot baseline because the built-in local embeddings and storage engine have a fixed footprint that exists before any document is processed. The limit is printed at boot, and whenever adds are waiting the binary shows a live status line in the terminal:
Raise the limit and concurrency on machines with spare RAM for faster bulk imports; lower them on small VPSes where you want the server to stay lean and don’t mind adds draining slowly.

Telemetry

Supermemory collects telemetry. You can disable it with SUPERMEMORY_DISABLE_TELEMETRY=1.

Platform-only features

These exist in the codebase but are exclusive to the hosted platform — the self-hosted binary doesn’t include them:
  • Connectors — Google Drive, Notion, Gmail, OneDrive background sync
  • Supermemory MCP — managed MCP server endpoints
  • Optimized memory extraction — the platform’s extraction pipeline is tuned for higher quality at lower cost than bring-your-own-key
  • Managed scale — globally distributed infrastructure, no capacity planning
Any other environment variables you may find referenced in the codebase are platform-only: the self-hosted binary ignores them even when set.

Example: production-ish .env

That’s enough for full ingestion, memory extraction, and hybrid search with the default local embeddings.