This documentation was generated by AI Agents, for AI Agents. If anything is incorrect or missing, contact contact@omnesis.dev.

Documentation

Overview

Omnesis indexes and searches your email, messages, notes, files, health data and calendars in one place, on machines you own. This page explains the parts it is made of and where your data lives. Read it before you install; every other page assumes these terms.

How it fits together

One gateway holds everything. Every other component is a device paired to it with its own token: the collector that syncs your accounts, the CLI, the portal in your browser, the phone apps and the browser extension. External AI agents reach the gateway too, through approved connections rather than as devices.

  1. SourcesGmail, WhatsApp, Apple Notes, Strava, Obsidian…
  2. Collectorsyncs each source on a schedule, on the machine that has it
  3. Gatewaystores, indexes and searches documents, runs the agent, serves the portal on port 7600

Push data in

  1. Phone apps and the browser extensionhealth, activity, location visits and photo text from phones; pages you read from the browser

Each holds write scopes for its own source types only; the extension can never read.

Read and manage

  1. CLI, portal and phone appssearch, SQL, the agent, sources, devices, settings, notifications

Device tokens, minted at pairing.

Ask over MCP

  1. External agentsClaude, ChatGPT, Codex and others

OAuth connections, limited by an access level.

Sources flow through a collector into the gateway. Phone apps and the browser extension also push their own data straight to the gateway; the CLI, portal and phone apps read and manage it; external agents ask questions over MCP.

The gateway owns all state: documents, the search index, analytics, people, devices, external-agent connections and configuration. The collector pulls data out of your accounts and local apps and sends it to the gateway. The gateway and a collector can share a machine or run on different ones.

The gateway

The gateway is one process serving HTTPS on port 7600. It hosts the document store (SQLite), the indexer, the hybrid search pipeline, the agent, the analytics database, the device and token registry, and the portal: its built-in web interface at https://<gateway-host>:7600/portal/.

It serves HTTPS only; there is no plain-HTTP mode. Which certificate it serves, and what that means for browsers and phones, is covered in Setup → Certificates. Every route declares the scope it requires, and a route that declares none can never be served. Only a small surface is public: the health check, pairing-code redemption, the portal's login page and static files, the OAuth sign-in endpoints that external agents and provider sign-ins redirect through, and a few non-secret lookups such as model logos. Everything else needs a token with the right scope.

What leaves your machines

Your documents are never sent to an Omnesis service. The gateway and collectors make these outbound requests, each of which you enabled or can switch off:

The collector

The collector is the process that syncs sources on a machine. It listens on no port of its own, apart from short-lived localhost listeners that catch a provider's OAuth redirect during sign-in. It sends documents to the gateway over HTTPS and holds one WebSocket open for live control. Which sources it runs, and their configuration, come from the gateway, so you manage sources in one place and the collector follows.

It can run on a different machine from the gateway. A common split is the gateway on an always-on server and the collector on the desktop where your accounts and local app data live. On the same LAN, a collector finds the gateway through mDNS. See Setup → Collector on another machine.

Devices and connections

A device is anything paired to the gateway with its own token: a collector, the CLI, a portal browser, an iOS or Android app, the browser extension, a machine running an OpenClaw or Hermes harness, or an integration (someone else's code that asks Omnesis questions for you, such as a voice-assistant handler). A device pairs by redeeming a single-use pairing code: 10 characters, valid for 10 minutes by default. The gateway then issues a token limited to the scopes that kind of device needs. How to pair each kind is on Setup.

Scope What it allows Held by default by
read Search and read documents, analytics and people. Collector, CLI, portal, phone apps.
read:bulk Enumerate the whole corpus at once: full document lists, dumps of analytics tables. Includes read. No device; add it to a token you create.
admin Manage sources, devices, tokens, models and configuration. Also covers answer, but not read. CLI, portal, phone apps.
answer Ask questions through the privacy-reviewed Answer endpoint, and nothing else. Integrations.
write:<source-type> Push documents and analytics rows for one source type. Phone apps, for the sources they host on the phone; the browser extension, as write:web and nothing else.
write:* Push documents and rows for any source type. Collector.
push:claim Collect the notifications queued for this device. Phone apps.

A harness machine's device token carries one narrow scope, subscriptions:receive, for the harness plugin's own traffic and none of the above. An integration can hold only answer, push:claim and write scopes, whatever token you create for it. A phone's write scopes follow the sources its app version hosts, so a phone picks up a new source's scope when it next connects, without pairing again. Revoking a device invalidates its tokens but keeps its identity, so it can be brought back; each device also reports its version, as described in version compatibility.

On first start the gateway writes a bootstrap admin token to ~/.config/omnesis/token, readable only by your user. The collector on the same machine uses it to pair itself and saves its own token to collector-token. The CLI on that machine reads it too, which is why commands such as omnesis devices pair work on the gateway host without a login.

An external AI agent that reads your data over MCP is not a device. Each agent you approve becomes a connection, which signs in with OAuth and is governed by an access level: a named set of permissions that says which of Answer, Direct and Notes it may use, which sources those reach, and which privacy policy reviews its answers. Removing a connection affects no other connection. See Agents & MCP.

Sources and providers

A source is one account's data from one platform, identified as type:account, for example gmail:maya@example.com. Its source type is the kind of data, such as gmail. A provider is one sign-in that serves several source types: authenticate Google once and get Gmail, Calendar, Drive and Contacts.

Sources run in one of three places:

The full catalogue is on Available sources, and adding, pausing and removing sources is under Managing sources. Writing your own is covered in Building sources.

Documents and the index

Everything a source syncs becomes a document: Markdown content plus metadata (type, timestamps, source URL, the people involved), stored in SQLite. Ingest never waits for indexing. Each new document wakes the indexer, a background worker in the gateway. It splits the document into chunks suited to its type, adds them to a keyword index (SQLite FTS5), and embeds them with the configured embedding model into a vector index (HNSW).

A search queries both indexes and merges the results, so exact terms and paraphrases both match. Without an embedding model, search falls back to keywords alone. How ranking works, and the query syntax, are on Search.

People and reference graphs

Documents mention people: senders, recipients, chat participants, contacts. The gateway resolves those mentions across sources, so an email address, a phone number and a contact-card name become one person. Ambiguous cases appear as merge suggestions you confirm in the portal.

It also records references between documents (links, shared attachments, threads) in a reference graph. Search results show how often each document is referenced, and a document's trail follows those references to reconstruct the story around it.

A link can be recorded before its target exists: a chat can mention a pull request before you connect GitHub. Once the target document is indexed, the link resolves to it. When the same page exists both as a browser capture and as a document from a dedicated source, the graph points at the dedicated source's document and keeps the capture as a second copy.

Analytics

Structured sources emit typed rows alongside their documents: health samples, screen-time sessions, calendar events, bank transactions, activities. These land in a DuckDB database next to the document store, one table per record type, with rows linked back to their documents.

You query them with plain SQL: omnesis sql "..." from the CLI, or Debug → SQL in the portal. Queries are read-only; each runs in a sandboxed connection with writes and external access disabled.

The agent

The agent is a chat interface over your data, running inside the gateway. It answers questions by searching documents, following trails, looking up people and querying analytics, and cites the documents it used. It never changes your documents; the only thing it writes is its own memory, which cites its evidence. It starts working once you assign a model to it, local or remote. See The agent.

Where your data lives

Omnesis keeps its state in one directory, ~/.config/omnesis (set OMNESIS_CONFIG_DIR to move it; a hardened gateway uses /var/lib/omnesis-gateway). omnesis backup is the supported way to capture it; see Backup.

With encryption at rest, the stores and stored credentials in that directory are encrypted with keys that a root key unlocks. The root key is kept in your operating system's keyring, or sealed by a passphrase file, and never in the directory itself. A copy of the directory alone therefore cannot be opened on another machine: also keep the recovery code from omnesis keyring export-recovery.

If the root key is missing, the gateway refuses to open the stores rather than fall back to plaintext. A collector on another machine keeps its own keys and encrypts the provider data it stores the same way. Your source apps' own databases and Omnesis's logs are outside this encryption. See Encryption at rest.

The last column of the table applies to an install with encryption at rest; without it, every file is plaintext, readable only by your user.

Path What it is Encrypted at rest
omnesis.json Configuration, owned by the gateway; collectors fetch it. No; secrets it names live in config-secrets/.
.env Environment for the services: gateway URL, certificate paths, bind address and, under Docker, the image tag. No.
docker-compose.yml The container stack, on a Docker install. No.
OMNESIS.md Optional standing instructions you write for the agent; see The agent. No.
token The bootstrap admin token, written on the gateway's first start. Yes.
collector-token The collector's own token, saved when it pairs. Yes.
cache/omnesis.cache.json The collector's copy of the gateway's configuration, rewritten after every successful fetch. A collector that starts while the gateway is unreachable runs on it if it is less than 24 hours old, and fetches the current configuration once it connects. Do not edit it. Yes.
config-secrets/ Secrets the configuration refers to, such as model API keys. Yes.
keyring/storage-keys/ The per-store keys, each usable only with the root key. Yes, by the root key.
<provider>-credentials.json An OAuth client you registered yourself, such as your Google client; set it with omnesis creds set <provider>. Yes.
<provider>/<account>/credentials.json An API key you pasted for one account, kept per account so several accounts of a source can be connected side by side. Yes.
<provider>/<account>/tokens.json OAuth tokens for a Google, Microsoft, Notion or Strava account. Yes.
whatsapp/<account>/auth/ WhatsApp session files. Yes.
whatsapp/<account>/store.db The WhatsApp message archive. Yes.
apple-imessage/transcripts.db Voice-note transcripts made on the collector. Yes.
enable-banking/<account>/bootstrap/ Transaction pages fetched at consent time, removed once imported. Yes.
omnesis.db SQLite: documents, sources, devices, tokens, people, links, sync state. Yes.
index.db SQLite: document chunks and the keyword index. Yes.
index.usearch, index.usearch.enc The vector index, a cache that can be rebuilt. With encryption at rest it is kept on disk as index.usearch.enc; a plaintext working copy exists only while the gateway runs and is removed when it stops. Yes, while the gateway is stopped.
analytics.db DuckDB: typed rows from structured sources. The analytics.db.tmp/ directory beside it holds temporary files while the gateway runs. Yes.
conversations/ Agent conversations, one JSON file each. No.
gateway.lock Marks the directory as owned by a running gateway; a second gateway refuses to start while it is held. No.
models/ Downloaded local model files. No.
tls/ The gateway's TLS certificate and key. No.
backups/, exports/ Backups and exports you create. Yes.

Help and support

If this documentation is wrong or incomplete, email contact@omnesis.dev. Ask usage questions in GitHub Discussions, report bugs through the issue tracker, and read the contribution guide before proposing a code change. Everyone taking part follows the code of conduct. For a phone-app problem, include what Mobile apps asks for.

Omnesis is open source under the GNU Affero General Public License, version 3 or later. The source code for everything described in these docs, including the apps and the browser extension, is in the omnesis-dev/Omnesis repository. The Omnesis name and logo are trademarks.

Never include your data in a report. Emails, messages, contacts, health records, names and other indexed details are private, and so are tokens and pairing codes. Reproduce a problem with invented examples. Report a vulnerability through the private security disclosure process, not a public issue.

Where next