Guides

Agentic Systems

Multi-user agent memory and credential isolation

How to key memory, credentials and channel defaults to the acting user and tenant, so one person's context cannot show up in a Slack group or another tenant.

By Tirth Gajjar · Founder & CTO

16 min

At a glance

A scoped tool token does not stop one user's memory or OAuth from reaching a Slack group or another tenant. Key memory and credentials to the acting user inside the tenant, and decide in code what a direct message may carry that a channel may not.

Use this guide to Agentic Systems to review the design choices and checks for your system.

Who this is for
Engineers building AI systems and technical leads reviewing the implementation.
Topics
  • Agent Memory
  • Model Context Protocol (MCP)
  • AI Agent
  • Tool / Function Calling
  • Guardrails
  • LLM Observability & Tracing

Published

How do you stop one user's memory or OAuth from showing up in a Slack group or another tenant? At Bigcircle we treat a scoped tool token as necessary and still incomplete. Once an AI agent serves more than one person, memory, credentials and the rule for a channel have to be keyed to the acting user and the tenant.

A direct message and a group channel are different rooms, and the code has to say which facts may cross. Leave a key unset and the next person can act on context they were never given, including a note or a calendar they cannot open. The token's scope does not repair a missing key.

The hard part is that this leak does not look like a permission error. The tool call succeeds, because the token is valid. The reply is fluent, because the memory it used is a fact someone really stored. The person reading the channel was not the person the fact was stored for. A success status has nothing to say about that mismatch.

The second difficulty is the shape most agent harnesses ship with. A store keyed by thread, a bot token for every invited channel, and a memory file owned by the process all behave on a laptop with one owner. In a shared workspace those three are now common to everyone who can mention the bot.

This guide starts with the three isolation units and what each one is allowed to hold, then names the three ways a fact or a token leaves its owner. The pattern after that keeps credentials, user memory and channel defaults on separate keys, then the guide decides when shared memory is allowed, who may see it, and how the keys sit in front of MCP.

What does each isolation unit own?

A multi-user agent separates three units, and each unit has its own store. The user holds private facts and credentials, the tenant holds the boundary between customers, and the conversation holds only the current thread. A fact written into the wrong store is read by the wrong unit later, while the read itself still looks like ordinary retrieval.

UnitIt ownsIt must not hold
UserPreferences, private facts they asked the agent to keep, and the OAuth tokens they grantedAnother person's facts, or a token minted for the app as a whole
TenantWhich users exist, procedures already visible across the workspace, and the prefix that stops one customer reading anotherA person's private memory, or any token that can cross to a second customer
ConversationMessages in this thread, scratch notes for this task, and which channel the thread sits inLong-term memory, credentials, or facts pulled from a different thread

Keying memory only by the conversation lets the next person in that thread, or a colliding identifier, read the previous person's notes. Keying it only by the tenant puts one preference into every user's prompt. Keying credentials only by the app makes every tool call run as the bot, with channel membership as the only check.

The channel is a property of the conversation, and it decides which of the user's facts may enter the prompt on this turn. A direct message may carry that user's private notes, while a group channel may carry only what was stored for the channel. Ignore that property and the correct user key still places a private note in a public room.

How does a fact or a token leave its owner?

Memory bleed, OAuth reuse and over-sharing into a group channel are three different misses, and they do not fall together. A scoped token limits reuse only when the call also refuses to attach anyone else's credential. Closing one miss leaves the other two standing, which is why a token rotation on its own does not finish the work.

Three misses on one shared botIllustrative model. Two people, one process. Not an incident and not a rate.Person A, in a direct messageRemember Thursday. Connects a calendar.Person B, in a group channelAsks what the calendar holds this week.One bot processMemory key is the bot. Token is the installer's.Memory bleedThursday's note is inthe prompt built for B.OAuth reuseA's calendar is fetchedfor B's question.Group over-shareThe reply is postedin the group channel.A narrower scope on the installer's token does not move any of the three.

This figure is an illustrative model. It walks one imagined bot through two people and three misses, and it is not a record of an incident or a rate we measured.

Take a Slack bot as the shape of the failure, rather than as a measured result from a line we publish. One person messages the bot in a direct message, asks it to remember a review on Thursday, and connects a calendar with OAuth.

A teammate later mentions the bot in a public channel and asks what the calendar holds this week. If the memory key is the bot, Thursday's note is already in the prompt. If the tool uses the first person's token, or the bot token granted calendar scope at install, the events come back into the channel.

Slack's agent design guidance says an agent should not use information the invoking user could not open on their own. Private information, in that same guidance, belongs in a direct message, a private channel or an ephemeral message. A reply that carries a private note into the group has already broken the room, whatever the token scope says.

We treat that guidance as their rule for the room, and not as a count of how often a bot breaks it, because we do not have a count to publish. The rule cannot be applied when the store cannot tell the invoking user from the previous one. The store has to know who is asking.

A token can be scoped, short-lived and stored encrypted, and still be the wrong token for the person who asked. Scope limits which operations the token may call, and it does not name the person this request is acting for. Encrypting that token does not change whose data it returns.

The usual form is an installer who connected once, after which every later caller uses that same grant. Presenting one user's token on another user's request is reuse, including when the scopes look narrow and the expiry looks careful. Until the grant is stored on the acting user, every caller shares the installer's reach.

A group channel adds a default a single-user agent never had to write down. The bot can read the channel because it was invited, so a bot-token call can return a record the asking user cannot open, and the reply goes to the whole room. Rotate only the token and the next channel question still loads the note.

What pattern keeps credentials, memory and the channel apart?

Keep the credential and the memory on the acting user, inside the tenant, and let a channel rule decide what may leave that store. The tool call uses a token that user granted, chosen by their id, and the tenant prefix stops two customers colliding on one key. Build the prompt only after that rule has run.

StoreKeyWritten whenRead when
User credentialTenant and userThe user finishes OAuth, and again on refreshA tool call for a request that user started
User memoryTenant, user and factThe user asks to remember, or confirms a writeA later request by that user, in a room the channel rule allows
Channel memoryTenant and channelThe fact was posted in that channel, or a member explicitly files it thereA request in that same channel
Conversation scratchTenant and conversationDuring this threadUntil the thread ends, then discard it

On the tool layers we build, the model never receives the token itself. The layer holds it, attaches it on the upstream call, and returns either a result or a typed refusal the model can repeat without guessing. Guessing past a refusal is how an empty result gets narrated as a fact.

Our guide on how to scope AI agent permissions binds that call to a principal, a task and a target, and this guide supplies the key that principal is looked up by. A boundary that says "the user" while the call runs on a bot token has named a person it did not use.

Drop a missing user id into a refusal, because a quieter fallback still answers the question. The fallback to a bot token turns the installer's grant into the grant for everyone who mentions the bot. Say the agent cannot act until that person connects, and wait. A wider credential during the wait is the leak.

Refresh tokens stay on the same key as the access tokens they replace. Writing the new access token onto the app, instead of onto the user, undoes the split at the moment the old token expires. Name the user on the refresh record, then decide which facts may leave that user for a workspace store.

When may a workspace share memory?

Shared workspace memory is allowed when the people in the room can already open the fact, and it is forbidden when sharing would widen that audience. A company procedure that is already visible inside the workspace can live in a tenant store and be read in any channel of that tenant.

When shared memory is allowedFramework. Who may receive the fact. No frequencies, because the mix is yours.DM with ownerSame channelOther userOther tenantCompany procedurealready visibleShareShareIn tenantRefuseNote postedin this channelPointerShareMembersRefusePersonalpreferenceThat userRefuseRefuseRefuseFact storedfrom a DMThat userIf filedRefuseRefuseOAuth orrefresh tokenNeverNeverNeverNeverA token never enters the prompt. Filing a DM fact into a channel is a new publication.

This figure is a framework. It sorts facts by who may receive them, and it reports no frequencies, because a leakage rate would be a number we do not have.

FactDirect message with the ownerChannel where it was postedAnother user, same tenantAnother tenant
Company procedure already visible in the workspaceYesYesYes, inside the tenantNo
Note posted in this channelOnly as a pointer back to the channelYesOnly if they are in the channelNo
Personal preferenceYes, for that userNoNoNo
Fact the user stored from a direct messageYes, for that userNo, unless they explicitly file it into the channelNoNo
OAuth token or refresh tokenNever placed in the promptNeverNeverNever

The row teams argue over is the channel note, which stays in that channel even when the tenant id matches a person outside it. A summary posted into a larger room is a new publication, and the agent may publish it only when a member of the source channel files it under their name.

A preference looks small until it pulls a document the room should not see, such as an unreleased pricing sheet on a later question. Store it on the user, and still ask the source whether this room may see the sheet before anything from it enters the prompt. The preference does not outrank the source.

Every row stops at the tenant boundary, including the rows that look harmless inside one company. An identifier that matches across two customers refers to two people, and the prefix on the key keeps their stores apart. The prefix is checked before the read, not after a result comes back.

Refuse a request with no tenant, and treat a demo workspace as a real tenant rather than a shortcut. A demo token copied into production is one way a test user's notes show up as a customer's context. A refusal from the matrix has to be explicit, and someone has to be allowed to see that it happened.

Who may see stored memory, and when does the agent refuse?

A fact the matrix refuses is not dropped quietly, because a silent skip leaves the model to invent a softer answer. Who may see the body, who may see only that an entry exists, and who may act on it are different permissions, and they belong in the record before anyone argues about a tool scope.

WhoWhat they may seeWhat they may do
User the fact is keyed toThe body, in a room the channel rule allowsRead it back, or delete it
Tenant administratorThat an entry exists, who wrote it, and whenDelete it. The body stays closed unless a separate grant says otherwise
Incident investigatorA redacted record by defaultThe body only under a named grant that expires and is logged
The agentThe body only inside that user's own requestNo search across other users to answer one person's question

The agent may act on a fact only for the user it is keyed to, by loading it into that user's request and nowhere else. A search across users, opened because the question was whether anyone else had Thursday held, builds a directory of private notes and then asks the model to read it.

Refuse the search and say so, because an empty list will be told back as if nothing were scheduled. Return a typed refusal the model can repeat, and keep the two outcomes distinct in the log. Our guide on dual-layer production agent ops shows why an empty success gets investigated by the wrong owner when the trace looks finished.

The audit line names the tenant, the acting user, the conversation, the channel type, the memory key and the credential id. The token stays in a vault, and the log holds a reference, because a token in the log is a second copy for everyone who can read logs. Keep the line in the same observability stack as the other tool logs.

Deletion has to reach the copies, or the fact is still available to the next summary. A note removed from the user store can still sit in a conversation summary, in a channel memory someone filed, and in a trace kept for a later eval. Compaction will copy whatever it can still see.

Delete or redact the copies before the next summary is written, including traces you plan to replay later. The record you keep for a challenge is the read or the refusal, not the body, unless a grant says otherwise. Our guide on how to build an audit trail for AI decisions is the longer account of what a challenged decision must still show. That rule does not loosen when the tool call goes out through an MCP server.

How does this sit in front of MCP tool access?

A tool scope on an MCP server is the credential half of the pattern, and it does not choose the memory key. The authorization specification of 25 November 2025 requires a protected MCP server to behave as an OAuth 2.1 resource server, and to reject a token that was not issued for that server.

The requirement stops a token minted for one server from being spent at another, which is a real boundary and a narrow one. It does not stop your process from attaching one user's token to a different user's request, and it does not choose which memory was loaded before the call.

Pick the user before the server checks the tokenFramework. Audience validation does not choose whose memory was loaded.RequestActing user, tenant, channel, conversationYour layerResolve the user and the tenant, or refuse.Load memory on that key. Apply the channel rule.Select that user's token. Do not fall back to the bot.The model does not receive the token.MCP serverReject a token issued for a different server.Run the tool. Return a result or a typed refusal.Reply and logA group channel receives no private note.Log the key and the refusal, not the secret.Shared installer sessionOne cached token forevery person in the workspace.The audience check can pass.The acting user is missing.Open a session per acting user, or pass the user on every call.

This figure is a framework. It places identity, memory and the channel rule in your layer, leaves audience validation on the MCP server, and describes that structure without standing in for any vendor's product.

Resolve the acting user and the tenant in your own layer, before a tool is selected, and load memory only on that key. Apply the channel rule, then call the tool with the token stored for that user, so the server validates a token you already chose. The server cannot repair a memory load that happened against the wrong person.

The consolidation guide checks the caller on every request and keeps the upstream credential inside the layer, out of the model's reach. That check still needs the acting user on the request, or it authenticates the installer and reads the installer's data. The governed context guide puts the same policy on every client, which holds only when each call arrives as the person who asked.

A playbook run as a tool, which the playbooks guide covers, has to record the user it ran as and refuse a memory that user does not own. A scope on the playbook tool does not supply that user, any more than a scope on a calendar tool does. Write the user onto the run record before the playbook reads anything.

Prompt injection raises the cost of a wrong key, and it does not change which key is the right one. A channel message can tell the agent to repeat a direct-message note, or to call a tool with a broader token than the asker holds. Both are the same miss, aimed at two stores.

The channel rule lives in code, outside the prompt, because injected text will argue with any instruction sitting beside it. Read the principal from the request, not from a name the model supplied, or the injection chooses the person and the token follows. Guardrails that trust the model for the principal will apply a correct scope to the wrong person.

The failure this arrangement still allows is a server-side session that outlives the person who opened it. The client connects as the installer, caches the tool list and the token, and then serves every person in the workspace through that one session. The cache is convenient, and it is also a shared credential with a long life.

The audience check passes, because the token really was issued for that server, and the acting user never entered the key. Open a session for each acting user, or pass the user on every call and ignore a cached token that belongs to someone else. When the grant is revoked, the next call has to refuse, and that refusal belongs in the infrastructure log rather than a run that continues on the cache. The same split, token in the layer and memory keyed to the acting user, is how we build MCP servers that more than one person can use.

Common questions

Why is a scoped tool token not enough?#

A scope limits which operations a token may call, and it says nothing about which person this request is for or which room will see the result. A narrowly scoped token that belongs to the installer still returns that person's data when someone else asks. Key the token and the memory to the acting user inside the tenant, and decide in code what a group channel is allowed to receive.

What is memory bleed between users?#

Memory bleed is a later request reading a fact that was stored for someone else, or for a different room than the one now asking. It shows up when the store is keyed by the bot, by the workspace, or by a thread more than one person can continue. The tool call can succeed and the reply can be fluent, because the fact is real. The failure is that the owner of the fact is not the person in the room.

Can one bot token serve every user?#

A bot token is the app, not a person. It can see the channels the app was invited into, and a call made with it is checked against the app. Use it for posts that should appear as the bot. When the call reads or changes a person's data, use a token that person granted. If the request has no such token, refuse, and do not fall back to the bot token in order to answer anyway.

When is shared workspace memory acceptable?#

Shared workspace memory is acceptable for a procedure or a fact the audience can already open, and only inside one tenant. A company playbook that is already visible across the workspace can be read in a channel. A note from a direct message, a personal preference, and any credential cannot. A fact posted in one channel stays there unless a person who can see it explicitly files a copy somewhere wider.

Who is allowed to read stored user memory?#

The user the fact is keyed to may read it, in a room the channel rule allows, and may delete it. A tenant administrator may see that an entry exists and may delete it, and does not get the private body by default. An investigator gets a redacted record, or the body under a named grant that expires and is logged. The agent does not search across users to answer a question one person asked.

How should a memory refusal be returned?#

Return a typed status that says this caller may not use that memory or that credential, and say the same thing to the user in plain language. Do not return an empty list. An empty list is easy to retell as "nothing is scheduled" or "no notes exist", which is a different fact. Log the refusal with the user, the tenant, the channel type and the key that was refused, and leave both the secret and the note body out of the log.

Do MCP audience checks replace per-user memory keys?#

No. The server should still reject a token that was issued for a different server, and that check stays. In front of it, your layer has to select the token that belongs to the acting user and load only that user's memory, after the channel rule has run. A session opened by the installer and then shared across a workspace can pass the audience check and still leak. Use one session per acting user, or send the user on every call.

What happens to memory when a user is deactivated?#

Lock or delete their credential immediately, so a refresh cannot mint a new access token for a person who has left. Keep the audit references that say reads and refusals happened. Remove the private body from the user store, from conversation summaries, and from any channel copy they did not explicitly file as workspace knowledge. A deactivated user's token that stays valid until it expires is still a credential someone can call with.

Further reading