Use

The agent

Omnesis has a built-in chat agent that runs inside the gateway. Ask it a question about your indexed data and it answers with citations to the source documents. This page is for anyone who chats with it: what it can do, what it remembers, how it is contained, and how to choose the model it runs on. To let an outside AI tool such as Claude, ChatGPT or Codex read your data, see Agents & MCP.

What it does

The agent is part of the gateway process, so there is no extra service to run. Given a question, it works through your data with a small set of tools:

Every answer cites the documents it drew from, and each citation opens the document itself. When a hard question splits into substantial independent parts, the agent can fan out workers in parallel. Each worker gets a fresh context and the same read-only retrieval tools, and reports its evidence; the agent then compares and merges what they found.

For a deliberate, fully orchestrated research pass, type / in the composer and choose Deep Research. Omnesis plans the question, runs focused readers in parallel, verifies their cited evidence, and writes one report. Neither needs any setup beyond an assigned model.

Open the agent from Ask in the portal (/portal/agent) or from Ask in the menu of the iOS and Android apps.

Memory

Ask the agent to remember a preference, a fact about you or another person, or useful context for a document. It saves the memory as an annotation that cites its evidence, and the annotation stays available across conversations and gateway restarts. Memory needs no model beyond the one the agent already uses. You can ask the agent to correct or forget a remembered fact.

A memory must cite a stored document or a message you wrote. The gateway checks the quoted evidence, and an agent reply never counts as something you said. Person pages show what Omnesis has learned about that person, your own person page shows your profile, and a document's details show its annotations under Enriched by Omnesis. Deleting a conversation also removes the annotations that cite it.

Sandboxed by design

The agent reads your documents and writes only its own memory. Its retrieval tools (search, documents, trails, people, SQL) read and never change anything:

The tools run against your own gateway only. Even when the model runs on a remote backend, it never touches your data stores directly; it sees only what the gateway's tools return.

Conversations

Conversations persist across gateway restarts. Each one is a plain JSON file under ~/.config/omnesis/conversations/, so your history is ordinary files you can inspect and copy. They are kept until you delete them, unless activityRetention.maxAge is set (see Configuration); pinned conversations are always kept. If the assigned model's backend is unreachable, you can still browse saved conversations; sending a message then shows the backend's error.

Returning to the portal or either mobile app within an hour reopens the conversation you were using. After an hour away, Omnesis opens a fresh conversation with the composer ready, and the previous one stays in the conversation list.

Unread conversations and notifications

A conversation the agent wrote into while you were not looking, such as a slow answer that finished later, is marked unread in the conversation list. Opening it clears the mark on every client, because the gateway holds the read state rather than each client. Only showing the conversation on screen counts; a client that merely synced it in the background has not read it.

When a conversation goes unread while you are away, your phone gets one notification, and tapping it opens that conversation. A turn that writes several messages, or further messages before you look, sends no second notification; the conversation can notify you again only after you have read it. A conversation already open on your screen never notifies. The Notifications page covers delivery.

Past conversations and length limits

Conversations are also indexed back into your data as the omnesis-chat source. Past conversations show up in search like any other document, and the agent can recall what you discussed with it before.

Omnesis never compacts, summarizes, prunes, or silently truncates a conversation to fit a model's context window. When a conversation reaches the model's limit, Omnesis keeps the failed turn and any partial answer, makes the conversation read-only, and asks you to start a new one. It stays read-only on every device, and changing the assigned model does not reopen it. When only a single answer reaches the model's output limit, Omnesis marks that answer as truncated and you can ask a follow-up.

Times and time zones

The portal and the iOS and Android apps send their own time zone when they open a conversation, and the agent answers in that zone. Ask when an evening appointment starts from a laptop at home, then ask again from a phone eight time zones away: each answer is in the hour that device's clock shows.

This matters because the gateway does not move. It stays on the machine that indexes your data, in whatever zone that machine is set to, while you travel. Relative windows follow the same rule: today and this week mean the day and week where you are, not where the gateway is. Two kinds of entry are left alone: an all-day event stays a date, and an event with its own zone (a flight landing abroad) is given in the local time of that place.

Search filters work differently: after: and before: resolve relative words to a UTC calendar date on the gateway's clock, so pass an absolute YYYY-MM-DD when a day boundary matters. The answer API receives no device time zone either, and uses the gateway's.

Pick the model

The agent is off until you assign a model to the agent role. You choose where that model runs:

# Anthropic: store the key, allow cloud inference, pick a model

❯ omnesis backend key set anthropic

❯ omnesis config set inference.allowRemoteInference true

❯ omnesis model list --role agent

❯ omnesis model assign agent anthropic/MODEL_ID

# fully local: declare an OpenAI-compatible server, then assign

❯ omnesis backend add ollama http://localhost:11434

❯ omnesis model assign agent ollama/qwen3:8b

# ChatGPT subscription: sign in to Codex once, then assign

❯ omnesis codex login --wait

❯ omnesis codex setup-agent MODEL_ID_FROM_STATUS

omnesis backend key set anthropic prompts for the key. omnesis model list --role agent prints the model IDs to assign, including anthropic/… entries once the key is set. Without remote inference enabled, the Anthropic assignment is saved but the agent stays off.

codex setup-agent takes a model from the status printed after login. It enables remote inference and assigns that model in one step, but only while the agent role is unassigned; it never replaces an existing assignment. Codex receives the conversation, the relevant Omnesis data, and tool results for these requests. When a codex command is installed and the agent role is unassigned, an interactive install offers to do this for you (see First run).

After assigning an HTTP or Codex model, open its role card in Settings → Models to choose the reasoning settings that model offers, such as a reasoning toggle, an effort level or a token budget; see Reasoning settings.

Nothing leaves your machine by default: cloud backends (Anthropic, Codex) and HTTP backends that do not resolve to loopback stay unused until you enable remote inference. Choosing a cloud model in the portal asks you to confirm before enabling this permission and saving the assignment. The permission covers all configured remote inference backends on the gateway; turn off the Allow cloud inference switch under Settings → Models. Backends, presets, the remote-inference setting, and the other model roles are covered in Operating → Models.

If the agent already has a cloud model assigned but cloud inference is disabled, the portal offers Enable cloud inference with confirmation, or Choose a local model. Enabling it sends your messages, conversation context, relevant Omnesis data, and tool results to the selected provider when you use the agent. After enabling it, use Retry message to send a failed message. No gateway restart or JSON editing is needed.

Standing instructions

OMNESIS.md is an optional Markdown file in the gateway's config directory. If it exists, the agent reads it on every run. Use it for what you would otherwise repeat in every conversation: who you are, how you want to be answered, conventions you expect it to keep.

Edit it in the portal under Settings → OMNESIS.md, or open the file in any editor. The file on disk is the only copy, so whichever wrote last is what the agent reads. There is no version history; backups include it.

~/.config/omnesis/OMNESIS.md
# House rules

- I work in engineering; assume technical vocabulary.
- Dates as YYYY-MM-DD, distances in kilometres.
- When you cite a message thread, name who sent it.

The file reaches the agent you chat with, the workers it delegates research to, and answers served to a connected external agent. It does not reach the privacy reviewer or the check on memory quotes, so nothing you write in OMNESIS.md can loosen your privacy policy.

The file is capped at 16 KB. Past that, only the first 16 KB reaches the agent, marked so the agent knows it is reading a fragment, and the portal says the file is over the cap. A file over 256 KB is not read at all, and the portal shows its size instead of opening an editor. The file is written readable by its owner only.

Changes take effect on the next run. A conversation already open keeps the instructions it started with, so start a new conversation to see an edit take hold.