Memorybeta

How an agent keeps a short prompt and still remembers a user: a rolling conversation summary, a per-user profile of durable facts, recall over earlier conversations, and the judgments Jev makes around them.

Long conversations used to be sent to the model almost whole: the last 20 messages, including every card the agent rendered. On a quiz-style agent that is 20–25k input tokens per turn for a one-line answer. Memory replaces that with three smaller things.

1. Rolling summary (per conversation)

Only the tail of the conversation is sent word for word (default: the last 6 messages). Everything older is folded into a running summary of at most 150 words that the runtime rewrites in the background after each reply, with a small model on your key (Haiku for Anthropic, gpt-6-mini for OpenAI, Flash for Google). The summary goes into the prompt as its own block above the recent messages.

Jev makes two judgments around it: which of the older messages are still needed verbatim to answer the next message (scored 0–3; “useful” and “essential” stay in the window), and whether a rewritten summary lost something the user said (the previous one is kept when it did).

The summary is visible in the conversation inspector. Turn it off per agent under Behaviour → Conversation summary; the window then falls back to the history size.

2. User memory (per user, across conversations)

Durable facts and preferences the user states (“I’m vegetarian”, “I live in Lisbon”, “I bought the 750 ml bottle”) are extracted after the reply, judged, and stored on the end user. Every later conversation starts with them in the prompt as what you know about this user, so the agent is personal from the first message without any transcript.

Jev decides, per turn, whether the message is worth remembering at all and whether it asks to forget something, so the extractor only runs when needed. Per candidate fact it judges durability (durable, situational, noise), how it relates to what is stored (none, same, update, contradiction) and how sensitive it is. Situational and noise are dropped, “same” refreshes the stored fact, “update” replaces it, a contradiction keeps both with the newer one first, and sensitive facts (health, finances, identity and similar) are dropped unless the agent allows them. At most 50 facts per user.

Only what the user explicitly said is stored, never inferences and never the assistant’s own words. The customer sees, edits, adds and deletes facts on the Users page, and “forget everything” clears them. A user asking the agent to forget something has the same effect.

3. Recall (per user, across conversations)

Every turn (the user’s message and the reply) and every summary rewrite is embedded and stored per user, so the agent can find the long tail that does not fit the profile. On each turn Jev first judges whether the message needs earlier conversations at all (“last time”, “my usual”, a past order); most turns do not, and nothing is retrieved. When it does, the runtime embeds the message, pulls the ten closest snippets from that user’s earlier conversations (never the current one), lets Jev score each 0–3 for usefulness, and injects at most three that scored “useful background” or better as from earlier conversations with this user. The inspector shows a memory.recall entry on turns that recalled something.

Indexing is asynchronous: a turn becomes searchable seconds to a few minutes after it happened, so a question asked within the same minute may not recall it yet (the user profile covers durable facts in that gap). Embeddings run on Kletso’s own account (Cloudflare Workers AI), not on your key, and vectors live in a Kletso-managed index filtered by end user. “Forget everything” on the Users page, or the user asking to forget everything, deletes that user’s vectors as well as their facts. Turn it off per agent under Behaviour → Recall earlier conversations.

Retention

Recall snippets and remembered facts are kept for the project’s memory retention window (Settings → Project; default 180 days, 0 keeps them until the user is forgotten). A nightly job removes older ones; facts the user restates are refreshed and stay. “Forget everything” removes a user’s facts and snippets immediately.

Cost and latency

The gate judgments ride on the one Jev call the turn already makes, so they add no round trip; recall adds one embedding and one Jev re-rank only on the turns that need it. Summaries and extraction run after the reply is sent and are billed on your key under memory:<model> in the cost view. On SeeCircles’ quiz agent the window change alone cuts input tokens per turn by roughly 80 percent.

Last updated 2026-09-28 · Report an issue with this page