How do you stop one user's memory or OAuth from showing up in a Slack group or another tenant? At Bigcircle we treat a scoped tool token as necessary and still incomplete. Once an AI agent serves more than one person, memory, credentials and the rule for a channel have to be keyed to the acting user and the tenant.
A direct message and a group channel are different rooms, and the code has to say which facts may cross. Leave a key unset and the next person can act on context they were never given, including a note or a calendar they cannot open. The token's scope does not repair a missing key.
The hard part is that this leak does not look like a permission error. The tool call succeeds, because the token is valid. The reply is fluent, because the memory it used is a fact someone really stored. The person reading the channel was not the person the fact was stored for. A success status has nothing to say about that mismatch.
The second difficulty is the shape most agent harnesses ship with. A store keyed by thread, a bot token for every invited channel, and a memory file owned by the process all behave on a laptop with one owner. In a shared workspace those three are now common to everyone who can mention the bot.
This guide starts with the three isolation units and what each one is allowed to hold, then names the three ways a fact or a token leaves its owner. The pattern after that keeps credentials, user memory and channel defaults on separate keys, then the guide decides when shared memory is allowed, who may see it, and how the keys sit in front of MCP.
What does each isolation unit own?
A multi-user agent separates three units, and each unit has its own store. The user holds private facts and credentials, the tenant holds the boundary between customers, and the conversation holds only the current thread. A fact written into the wrong store is read by the wrong unit later, while the read itself still looks like ordinary retrieval.
| Unit | It owns | It must not hold |
|---|---|---|
| User | Preferences, private facts they asked the agent to keep, and the OAuth tokens they granted | Another person's facts, or a token minted for the app as a whole |
| Tenant | Which users exist, procedures already visible across the workspace, and the prefix that stops one customer reading another | A person's private memory, or any token that can cross to a second customer |
| Conversation | Messages in this thread, scratch notes for this task, and which channel the thread sits in | Long-term memory, credentials, or facts pulled from a different thread |
Keying memory only by the conversation lets the next person in that thread, or a colliding identifier, read the previous person's notes. Keying it only by the tenant puts one preference into every user's prompt. Keying credentials only by the app makes every tool call run as the bot, with channel membership as the only check.
The channel is a property of the conversation, and it decides which of the user's facts may enter the prompt on this turn. A direct message may carry that user's private notes, while a group channel may carry only what was stored for the channel. Ignore that property and the correct user key still places a private note in a public room.
How does a fact or a token leave its owner?
Memory bleed, OAuth reuse and over-sharing into a group channel are three different misses, and they do not fall together. A scoped token limits reuse only when the call also refuses to attach anyone else's credential. Closing one miss leaves the other two standing, which is why a token rotation on its own does not finish the work.
This figure is an illustrative model. It walks one imagined bot through two people and three misses, and it is not a record of an incident or a rate we measured.
Take a Slack bot as the shape of the failure, rather than as a measured result from a line we publish. One person messages the bot in a direct message, asks it to remember a review on Thursday, and connects a calendar with OAuth.
A teammate later mentions the bot in a public channel and asks what the calendar holds this week. If the memory key is the bot, Thursday's note is already in the prompt. If the tool uses the first person's token, or the bot token granted calendar scope at install, the events come back into the channel.
Slack's agent design guidance says an agent should not use information the invoking user could not open on their own. Private information, in that same guidance, belongs in a direct message, a private channel or an ephemeral message. A reply that carries a private note into the group has already broken the room, whatever the token scope says.
We treat that guidance as their rule for the room, and not as a count of how often a bot breaks it, because we do not have a count to publish. The rule cannot be applied when the store cannot tell the invoking user from the previous one. The store has to know who is asking.
A token can be scoped, short-lived and stored encrypted, and still be the wrong token for the person who asked. Scope limits which operations the token may call, and it does not name the person this request is acting for. Encrypting that token does not change whose data it returns.
The usual form is an installer who connected once, after which every later caller uses that same grant. Presenting one user's token on another user's request is reuse, including when the scopes look narrow and the expiry looks careful. Until the grant is stored on the acting user, every caller shares the installer's reach.
A group channel adds a default a single-user agent never had to write down. The bot can read the channel because it was invited, so a bot-token call can return a record the asking user cannot open, and the reply goes to the whole room. Rotate only the token and the next channel question still loads the note.
What pattern keeps credentials, memory and the channel apart?
Keep the credential and the memory on the acting user, inside the tenant, and let a channel rule decide what may leave that store. The tool call uses a token that user granted, chosen by their id, and the tenant prefix stops two customers colliding on one key. Build the prompt only after that rule has run.
| Store | Key | Written when | Read when |
|---|---|---|---|
| User credential | Tenant and user | The user finishes OAuth, and again on refresh | A tool call for a request that user started |
| User memory | Tenant, user and fact | The user asks to remember, or confirms a write | A later request by that user, in a room the channel rule allows |
| Channel memory | Tenant and channel | The fact was posted in that channel, or a member explicitly files it there | A request in that same channel |
| Conversation scratch | Tenant and conversation | During this thread | Until the thread ends, then discard it |
On the tool layers we build, the model never receives the token itself. The layer holds it, attaches it on the upstream call, and returns either a result or a typed refusal the model can repeat without guessing. Guessing past a refusal is how an empty result gets narrated as a fact.
Our guide on how to scope AI agent permissions binds that call to a principal, a task and a target, and this guide supplies the key that principal is looked up by. A boundary that says "the user" while the call runs on a bot token has named a person it did not use.
Drop a missing user id into a refusal, because a quieter fallback still answers the question. The fallback to a bot token turns the installer's grant into the grant for everyone who mentions the bot. Say the agent cannot act until that person connects, and wait. A wider credential during the wait is the leak.
Refresh tokens stay on the same key as the access tokens they replace. Writing the new access token onto the app, instead of onto the user, undoes the split at the moment the old token expires. Name the user on the refresh record, then decide which facts may leave that user for a workspace store.
When may a workspace share memory?
Shared workspace memory is allowed when the people in the room can already open the fact, and it is forbidden when sharing would widen that audience. A company procedure that is already visible inside the workspace can live in a tenant store and be read in any channel of that tenant.
This figure is a framework. It sorts facts by who may receive them, and it reports no frequencies, because a leakage rate would be a number we do not have.
| Fact | Direct message with the owner | Channel where it was posted | Another user, same tenant | Another tenant |
|---|---|---|---|---|
| Company procedure already visible in the workspace | Yes | Yes | Yes, inside the tenant | No |
| Note posted in this channel | Only as a pointer back to the channel | Yes | Only if they are in the channel | No |
| Personal preference | Yes, for that user | No | No | No |
| Fact the user stored from a direct message | Yes, for that user | No, unless they explicitly file it into the channel | No | No |
| OAuth token or refresh token | Never placed in the prompt | Never | Never | Never |
The row teams argue over is the channel note, which stays in that channel even when the tenant id matches a person outside it. A summary posted into a larger room is a new publication, and the agent may publish it only when a member of the source channel files it under their name.
A preference looks small until it pulls a document the room should not see, such as an unreleased pricing sheet on a later question. Store it on the user, and still ask the source whether this room may see the sheet before anything from it enters the prompt. The preference does not outrank the source.
Every row stops at the tenant boundary, including the rows that look harmless inside one company. An identifier that matches across two customers refers to two people, and the prefix on the key keeps their stores apart. The prefix is checked before the read, not after a result comes back.
Refuse a request with no tenant, and treat a demo workspace as a real tenant rather than a shortcut. A demo token copied into production is one way a test user's notes show up as a customer's context. A refusal from the matrix has to be explicit, and someone has to be allowed to see that it happened.
Who may see stored memory, and when does the agent refuse?
A fact the matrix refuses is not dropped quietly, because a silent skip leaves the model to invent a softer answer. Who may see the body, who may see only that an entry exists, and who may act on it are different permissions, and they belong in the record before anyone argues about a tool scope.
| Who | What they may see | What they may do |
|---|---|---|
| User the fact is keyed to | The body, in a room the channel rule allows | Read it back, or delete it |
| Tenant administrator | That an entry exists, who wrote it, and when | Delete it. The body stays closed unless a separate grant says otherwise |
| Incident investigator | A redacted record by default | The body only under a named grant that expires and is logged |
| The agent | The body only inside that user's own request | No search across other users to answer one person's question |
The agent may act on a fact only for the user it is keyed to, by loading it into that user's request and nowhere else. A search across users, opened because the question was whether anyone else had Thursday held, builds a directory of private notes and then asks the model to read it.
Refuse the search and say so, because an empty list will be told back as if nothing were scheduled. Return a typed refusal the model can repeat, and keep the two outcomes distinct in the log. Our guide on dual-layer production agent ops shows why an empty success gets investigated by the wrong owner when the trace looks finished.
The audit line names the tenant, the acting user, the conversation, the channel type, the memory key and the credential id. The token stays in a vault, and the log holds a reference, because a token in the log is a second copy for everyone who can read logs. Keep the line in the same observability stack as the other tool logs.
Deletion has to reach the copies, or the fact is still available to the next summary. A note removed from the user store can still sit in a conversation summary, in a channel memory someone filed, and in a trace kept for a later eval. Compaction will copy whatever it can still see.
Delete or redact the copies before the next summary is written, including traces you plan to replay later. The record you keep for a challenge is the read or the refusal, not the body, unless a grant says otherwise. Our guide on how to build an audit trail for AI decisions is the longer account of what a challenged decision must still show. That rule does not loosen when the tool call goes out through an MCP server.
How does this sit in front of MCP tool access?
A tool scope on an MCP server is the credential half of the pattern, and it does not choose the memory key. The authorization specification of 25 November 2025 requires a protected MCP server to behave as an OAuth 2.1 resource server, and to reject a token that was not issued for that server.
The requirement stops a token minted for one server from being spent at another, which is a real boundary and a narrow one. It does not stop your process from attaching one user's token to a different user's request, and it does not choose which memory was loaded before the call.
This figure is a framework. It places identity, memory and the channel rule in your layer, leaves audience validation on the MCP server, and describes that structure without standing in for any vendor's product.
Resolve the acting user and the tenant in your own layer, before a tool is selected, and load memory only on that key. Apply the channel rule, then call the tool with the token stored for that user, so the server validates a token you already chose. The server cannot repair a memory load that happened against the wrong person.
The consolidation guide checks the caller on every request and keeps the upstream credential inside the layer, out of the model's reach. That check still needs the acting user on the request, or it authenticates the installer and reads the installer's data. The governed context guide puts the same policy on every client, which holds only when each call arrives as the person who asked.
A playbook run as a tool, which the playbooks guide covers, has to record the user it ran as and refuse a memory that user does not own. A scope on the playbook tool does not supply that user, any more than a scope on a calendar tool does. Write the user onto the run record before the playbook reads anything.
Prompt injection raises the cost of a wrong key, and it does not change which key is the right one. A channel message can tell the agent to repeat a direct-message note, or to call a tool with a broader token than the asker holds. Both are the same miss, aimed at two stores.
The channel rule lives in code, outside the prompt, because injected text will argue with any instruction sitting beside it. Read the principal from the request, not from a name the model supplied, or the injection chooses the person and the token follows. Guardrails that trust the model for the principal will apply a correct scope to the wrong person.
The failure this arrangement still allows is a server-side session that outlives the person who opened it. The client connects as the installer, caches the tool list and the token, and then serves every person in the workspace through that one session. The cache is convenient, and it is also a shared credential with a long life.
The audience check passes, because the token really was issued for that server, and the acting user never entered the key. Open a session for each acting user, or pass the user on every call and ignore a cached token that belongs to someone else. When the grant is revoked, the next call has to refuse, and that refusal belongs in the infrastructure log rather than a run that continues on the cache. The same split, token in the layer and memory keyed to the acting user, is how we build MCP servers that more than one person can use.
Common questions
Why is a scoped tool token not enough?#
A scope limits which operations a token may call, and it says nothing about which person this request is for or which room will see the result. A narrowly scoped token that belongs to the installer still returns that person's data when someone else asks. Key the token and the memory to the acting user inside the tenant, and decide in code what a group channel is allowed to receive.
What is memory bleed between users?#
Memory bleed is a later request reading a fact that was stored for someone else, or for a different room than the one now asking. It shows up when the store is keyed by the bot, by the workspace, or by a thread more than one person can continue. The tool call can succeed and the reply can be fluent, because the fact is real. The failure is that the owner of the fact is not the person in the room.
Can one bot token serve every user?#
A bot token is the app, not a person. It can see the channels the app was invited into, and a call made with it is checked against the app. Use it for posts that should appear as the bot. When the call reads or changes a person's data, use a token that person granted. If the request has no such token, refuse, and do not fall back to the bot token in order to answer anyway.
When is shared workspace memory acceptable?#
Shared workspace memory is acceptable for a procedure or a fact the audience can already open, and only inside one tenant. A company playbook that is already visible across the workspace can be read in a channel. A note from a direct message, a personal preference, and any credential cannot. A fact posted in one channel stays there unless a person who can see it explicitly files a copy somewhere wider.
Who is allowed to read stored user memory?#
The user the fact is keyed to may read it, in a room the channel rule allows, and may delete it. A tenant administrator may see that an entry exists and may delete it, and does not get the private body by default. An investigator gets a redacted record, or the body under a named grant that expires and is logged. The agent does not search across users to answer a question one person asked.
How should a memory refusal be returned?#
Return a typed status that says this caller may not use that memory or that credential, and say the same thing to the user in plain language. Do not return an empty list. An empty list is easy to retell as "nothing is scheduled" or "no notes exist", which is a different fact. Log the refusal with the user, the tenant, the channel type and the key that was refused, and leave both the secret and the note body out of the log.
Do MCP audience checks replace per-user memory keys?#
No. The server should still reject a token that was issued for a different server, and that check stays. In front of it, your layer has to select the token that belongs to the acting user and load only that user's memory, after the channel rule has run. A session opened by the installer and then shared across a workspace can pass the audience check and still leak. Use one session per acting user, or send the user on every call.
What happens to memory when a user is deactivated?#
Lock or delete their credential immediately, so a refresh cannot mint a new access token for a person who has left. Keep the audit references that say reads and refusals happened. Remove the private body from the user store, from conversation summaries, and from any channel copy they did not explicitly file as workspace knowledge. A deactivated user's token that stays valid until it expires is still a credential someone can call with.
Further reading
- How to scope AI agent permissions. Binding one action to a principal, a task and a target, which still needs the user key this guide describes.
- MCP tool design: fewer, larger tools. Checking the caller on every request, and keeping the upstream credential inside the layer.
- Shared business context for every AI agent. Running each call as the person who asked, with one policy at the entry point.
- Agent playbooks as MCP tools. Recording the user a playbook ran as, once the catalog is exposed as tools.
- Dual-layer production agent ops. Why a denied call that looks like an empty success is handed to the wrong owner.
- How to build an audit trail for AI decisions. What a retained record must still show when a decision is challenged.
- MCP servers. The service line where this keying is part of the build.
- Agent memory, MCP, tool calling and prompt injection.
- Slack, Agent design. Their current guidance that an agent must not use information the invoking user could not open, and that private information stays in a direct message, a private channel or an ephemeral message.
- Slack, Tokens. Bot tokens belong to the app. User tokens carry the access of a specific workspace member.
- Model Context Protocol, Authorization, 25 November 2025. A protected server acts as an OAuth 2.1 resource server and must reject a token that was not issued for it.