Guides

Agentic Systems

MCP tool access in production: scoped tokens, audit, and refusal

How to let an agent call internal tools with a token minted for the run, a check before the call leaves, and a record of both the allow and the refuse.

By Tirth Gajjar · Founder & CTO

16 min

At a glance

An MCP connection is not a place to put the application secret. Mint a short-lived token for the run, check policy before the call leaves, and keep an audit line for every allow and every refusal.

Use this guide to Agentic Systems to review the design choices and checks for your system.

Who this is for
Engineers building AI systems and technical leads reviewing the implementation.
Topics
  • Model Context Protocol (MCP)
  • Tool / Function Calling
  • AI Agent
  • Guardrails
  • LLM Observability & Tracing
  • Prompt Injection

Published

How do we let agents call internal tools safely? At Bigcircle we treat that question as a property of the path a tool call takes, not as a warning written into the prompt. MCP is one way to declare the tools on that path. Any other tool bus, meaning any other way a call leaves the process, has to make the same decisions.

The call uses a token minted for this run, the tool names an operation we have declared, a policy runs before anything upstream is touched, and we keep a record when we allow the call and when we refuse it. A refusal is a result the agent can repeat. It is not an error the model is free to retell as an empty search.

The hard part is that a leaked key and a smoothed denial do not look like failures in the product. The key sits in a prompt or a trace and can perform every operation the application can perform, long after this task has ended. The denial, when the upstream API returns a 403 or an empty body, comes back as text the model already knows how to turn into a sentence about missing data. The user sees a finished answer. If the call never left your process, the log has no line that says anyone refused it.

The second difficulty is a tool that can do anything while still carrying a tidy name. A shell, a raw HTTP client, or a tool whose argument is a SQL string will sit in a catalog and still let the model choose the operation. A policy written against tool names cannot see inside that argument, so the allow you recorded is an allow for a program you did not review. Narrowing the token does not repair an argument that is itself a program.

This guide starts with what goes wrong when the agent holds the application key. It then separates a declared tool from an escape hatch, and covers the token we mint for a run and keep out of the model's context. The policy check has to run before the call, and the audit line has to exist for an allow and for a refuse.

Discovery at a large catalog comes last, because a tool the agent finds at call time still has to pass the same gate. Who the token belongs to, and how large each tool should be, are other guides.

What goes wrong when the agent holds the key?

Direct wiring puts the application's own secret on a path the model can read. The secret might sit in the system prompt, in an error the tool prints, or in a client the model is told to call with a header it can also copy into a message.

Once the secret is in context, a prompt injection can ask for it, a trace can store it, and a later run can reuse it. The reach is the application's reach, because the secret was issued for the application and not for this one task.

A first version of a layer often calls the vendor with a shared key while the auth work is still open. The calls succeed, which is why the shortcut stays. A person who has left keeps reaching the vendor for as long as that key stays in the process, and a log of successful calls cannot say which person the key was standing in for.

Rotating the key closes the copy you found. It does not name the calls that used it, and it does not stop the next prompt from asking the model to print the replacement.

What you wiredWhat the model can do with itWhat you can no longer show
The application key in the prompt or the toolEvery operation that key allows, until someone rotates itWhich person, which task, and which run used it
A client built with that keyThose operations, plus helper methods the client addsThat a denial was a decision, rather than a missing row
HTTP errors and empty bodies as the only denialTurn a 403 into a sentence about data that was not foundThat a policy ran, or that the user heard the truth

Take a refund as the shape of the failure, rather than as an incident we are publishing. The agent is given the payments key and a prompt that says to refund an order when the ticket asks. A line in the ticket says to ignore the cap and refund a different order.

The model calls the API, the key is valid, and the refund is created. A log of HTTP 200s is a record that a request succeeded. It is not a record of a decision anyone reviewed before the money moved.

The payments key is already in the promptIllustrative model. One imagined refund. Not an incident and not a rate.TicketRefund this order, above the cap.Model contextThe payments API key is in the prompt.Payments APIThe key is valid. The refund is created.What you can read afterwardsHTTP 200. No line that a policy ran.A sentence telling the model to be careful does not change which key was sent.

This figure is an illustrative model. It walks one imagined refund through a prompt that already holds the payments key, and the record at the end is a success status with no policy line. It is not an incident we measured, and it states no rate.

A longer prompt that tells the model to be careful with the key does not move the key. The model cannot be the place the secret is kept, because anything it can read it can be asked to repeat, including by text that arrived in the ticket. The secret belongs in a broker the model cannot print.

Each run then receives a token that names the tools this run may call. That token only means something if the tool itself names one operation, which is the next cut.

When is a tool an escape hatch?

A declared tool names one operation, the arguments that operation accepts, and the permission that operation needs. An escape hatch names a program and lets the argument decide the operation. refund_order, with an order id and an amount, is a declared tool. run_sql, http_request, and shell are escape hatches, even when you list them in an MCP catalog and write careful descriptions for them.

The difference shows up at the gate, which is why it belongs here rather than in a note about naming. A policy can allow refund_order under a cap and refuse it above the cap, because the amount is a field the gate can read. A policy cannot allow run_sql without also allowing every statement the database user can run, because the statement is the argument.

If you record an allow on that tool, you have recorded an allow for a class of programs. The audit line cannot say which program ran unless you store the statement, and by then you are reading code the model wrote after it had already run.

We still meet codebases that keep one general tool for the cases the catalog has not covered yet. Those cases are how a catalog grows, and they are real work. Put the general tool on a different credential, with a network allowlist and a short life, and have a person approve the call.

Do not mint that credential on the same token as the declared tools. The moment the model prefers the wide tool, the narrow tools inherit the wide tool's reach, and the scope written on the token no longer describes the call that left.

Our guide on MCP tool design is about how large a declared tool should be, and when several vendor servers should sit behind one layer. This guide assumes that shape and asks an earlier question: can the gate name the operation before it runs. A playbook exposed as a tool, which agent playbooks as MCP tools covers, stays a declared operation when its inputs are typed.

A playbook tool whose input is a free note is an escape hatch with a plainer name. We refuse that input at the gate even when the playbook id itself is on the allow list. The credential those declared calls actually carry is the next decision.

What token does a run actually hold?

Mint a token for this run, this principal, and this set of tools, and drop it when the run ends. The application secret stays in a vault. The tool layer attaches the token on the upstream call. The model receives a result or a typed refusal, and it does not receive the token, the secret, or a header it could copy into the next message.

If a debugger needs to see which credential was used, it sees an id. The vault is where the secret still lives, and a copy in the prompt would make the vault irrelevant.

A token the model can see will end up in a trace, because the trace is how you debug the run and more people can read the trace than can open the vault. We keep the credential id on the audit line for that reason. On Solarpunk the permanent credentials stay in an encrypted vault on the device, and they are injected at the adapter when a tool runs, outside the prompt. We build the same split into MCP servers: a short-lived token for the run, with the application credential kept on our side of the layer.

Scope the token to the tool names this run may call, and bind it to that server as an audience. A token valid for the whole server is a second copy of the server's reach, and a tool added later widens every unexpired token without a new grant.

The MCP authorization specification of 25 November 2025 requires a protected server to reject a token that was not issued for it, which stops a token minted for one server from being spent at another. That check does not choose which tools on this server the run may call. Leaving the tool list off the token means the audience check can pass while the run still reaches an operation nobody granted.

Uber described a production version of this split in their 21 May 2026 post on agent identity. Their security token service mints a JSON Web Token for a single hop, with a specific audience and a lifetime on the order of minutes, instead of handing the agent a long-lived service credential.

A token issued for one hop cannot be replayed at a database or at a different service. We treat that write-up as their published design for their own mesh. The same post says token-exchange latency stayed under 40 milliseconds at the 99th percentile, and that the design had been adopted by thousands of internal agents. Those figures are theirs. They are not a latency budget or a headcount for a system we have not built.

Whose grant you mint from is a separate decision, and a short lifetime does not settle it. Our guide on multi-user agent memory and credential isolation keys the grant to the acting user inside the tenant, and it refuses a fallback to the installer's token. This guide assumes that key is already there.

A token with a tight tool list and a short life, attached to the wrong person, still returns that person's data. Mint the run token from the grant stored for the person who asked, and fail the call when that grant is missing.

How long should the token live? Long enough for the tool calls this run actually makes, and short enough that nobody files it as a standing credential. A run that waits on a person may need a refresh, and the refresh has to mint another token with the same tool list, not a wider one.

If the run has ended, the next call that presents the old token fails closed. A session cache that keeps the first token and serves later runs with it has rebuilt the shared credential the per-run mint was meant to remove. The mint still depends on a check that can say no, which is the piece that has to run first.

Why does the policy run before the call?

A log tells you what already happened. A refund that has been created cannot be undone by a line you read the next morning, and a message that has been sent cannot be pulled back by a dashboard. The policy has to run while the call is still inside your process, on the tool name, the arguments, the principal, the task, and the target. If a check fails, the upstream request is never sent. Reading the log afterwards can explain the miss. It cannot stop the call.

Allow and refuse both stop in your processFramework. The upstream call happens on one branch. The audit line happens on both.CallPrincipal, task, tool, argumentsPolicy gateRuns before any upstream request.AllowRefuseAllowMint a token for this run.Attach it in the layer. Call upstream.Audit the allow. No secret in the log.The model receives the result.RefuseDo not mint. Do not call upstream.Return a typed reason.Audit the refuse. Same run id.The model can repeat the reason.An empty result is a different outcome, and it gets its own status.

This figure is a framework. It shows an allow branch and a refuse branch leaving the same gate, with an audit line on both, and it describes that structure without standing in for a vendor product. It does not state how often either branch fires, because that mix belongs to the system you are running.

The fields of the check are the ones in our guide on how to scope AI agent permissions: a principal, a task, and a target, plus the parameter range and the environment. That guide binds one action so a reviewer can say what was approved. Here the binding is enforced on the bus, on every call, including a call the model found through discovery a moment earlier.

A catalog entry is not an authorization. Finding a tool and being allowed to call it are different decisions, and we keep them in different functions so a change to the catalog cannot widen the grant on its own.

Review after the fact still has a job, and the job is different. It tells you whether the gate is refusing tools people need, whether a policy version is older than the tool, and whether an allowed call had an effect you did not expect. It cannot be the control for a payment, a permission change, or a message sent to a customer.

If the only reason a bad call stops is that someone reads the log, the call has already reached the system the gate was supposed to protect, and the review is a reconstruction.

Return the refusal as structured output the model can repeat, with a reason a person can read, and keep that status distinct from an empty result. An empty list means the query ran and matched nothing. A refusal means the query did not run. Our guide on dual-layer production agent ops shows why those two outcomes get investigated by different owners, and why an empty success sends the wrong team looking. Hand the model a status it can quote, in the shape below, rather than a stack trace it will compress into a story about missing rows.

code
{
  "status": "refused",
  "tool": "refund_order",
  "reason": "amount_above_cap",
  "policy": "refunds-v3"
}

We use that shape so the model and the log name the same check. It is not an export from a client system, and refunds-v3 is a label for the example rather than a policy version we are asking anyone to adopt. The useful part is the split between a reason code and an empty payload. Collapse them and the next sentence the user hears will pick one of the two at random.

Sierra's 22 July 2026 account of their internal MCP gateway draws a related line around data the agent should never receive. Interactive work runs as the user. A scheduled workflow runs as a service account limited to the tools it declared before it started. A read that would cross customers needs an approval written to the audit log, outside the prompt the model can argue with.

They also report that 89 percent of employees were using the gateway to reach 45 services. That adoption figure is theirs, on that date. We are not treating it as a rate another company should expect. The design point is the earlier one: stop the read before the record is in the prompt, because the agent cannot misuse a record it was never given.

Guardrails that only score the reply, after the tool has already run, are reading a sentence about an action that happened. Put the gate on the call, and use the reply check for what the user is about to be told, including whether a refusal was retold as a fact about the world.

A reply check that passes a fluent claim of "nothing was found" has checked the prose. The money or the message may already have moved, which is why the audit line has to exist for the refusal and not only for the call that went out.

What do you record for an allow and for a refuse?

Record both outcomes, or the log is a list of successes with a hole where the policy did its job. A tool that is refused often is either described badly or scoped more tightly than the task needs, and the calls that went through cannot show you that.

A tool that is never refused may mean the policy matches the work, or it may mean the policy is not on the path. The refused line is how you tell those two situations apart. Without it, a gate that is quiet and a gate that is missing look the same on the dashboard you actually watch.

FieldOn an allowOn a refuse
Run id, principal, taskPresentPresent, on the same ids
Tool name and versionThe tool that ranThe tool that was asked for
Argument summaryValidated fields, with personal values removedThe same summary, so the failing argument is visible
Decision and policy versionAllow, and which checks passedRefuse, and which check fired
CredentialAn id, not the tokenAn id, or none if you refused before minting
Upstream resultStatus, latency, and result sizeAbsent, because the call did not leave
What the model was toldA result handle, not the secretThe typed reason, in the words it received

Keep the token and the raw secret out of this record. A log that contains the credential is a second vault, and it usually has a wider set of readers than the first. Rotating the vault does not clear yesterday's log line. Keep the customer's document body out of the record too, unless a separate grant says an investigator may see the body, and make that grant expire and write a line of its own.

The argument summary is enough to see which field failed the cap. The body of an email or a contract is a different disclosure, and it does not belong on every allow.

Our guide on how to build an audit trail for AI decisions is the record you need when a business decision is challenged: which evidence was accepted, who changed it, and whether the later action was confirmed. The line in this section is narrower. It says the bus allowed or refused a tool call under a named policy.

On its own it does not explain the business decision the call belonged to. Keep both records, and link them by the run id. A reviewer who has only the tool trace can see that a call was allowed and still cannot say which terms were the ones accepted.

Write the line into the same observability stack as the rest of the service, and ship it from every instance. A refusal that lives only in a file on one machine will be missing on the day someone asks why a refund did not happen. The infrastructure owner and the agent owner have to be able to open the same line, and that shared line lets the two layers in the ops guide hand a case across instead of each reporting a green dashboard. That line still has to be written when the tool was not in the prompt at the start of the run.

How does discovery stay inside the same gate?

A large catalog cannot sit in the prompt on every turn. On Solarpunk, direct selection held until somewhere around thirty tools and then picked a plausible tool from the wrong system. The registry grew past a hundred tools, and the planner retrieved the definitions a step needed instead of holding the whole catalog.

The count, and the note that this shipped roughly six months before the major labs published the same pattern as a standard, are the ones already on our MCP servers page. We are not adding a new measurement in this guide.

Discovery changes which definitions the model reads on this step. It does not change which calls the model may make. A search that returns three tool definitions is not an allow for those three tools. The allow is the gate above, run against the tool the model then chooses, with the token minted for that run.

If discovery returns a tool the run's token does not include, the call is a refusal, and the refusal names the tool. Dropping the unknown tool with no status teaches the model that the catalog and the permission list are the same list. The next reply will describe the gap as a missing record, which is a different fact.

Uber's published hop, 21 May 2026Cited source. Their design for their mesh. Not a latency target and not a count of ours.1UserOpens a session with an agent.2AgentAsks STS for a token for the next audience.3Security token serviceMints one hop. Audience set. Minutes, not a standing key.4MCP GatewayChecks the tool policy. Can redact before the call.5DownstreamRuns only after that check. Or the gateway stops it.Not in this figureTheir p99 latency.Their agent count.Those stay intheir post.Order we are citing: identity for this hop, then the gateway, then the system.

This figure is a cited source. It redraws the hop Uber published on 21 May 2026, and it leaves their latency and their agent count in that post rather than on the diagram. Nothing in the figure is a measurement from a system we operate.

That post puts the policy on the gateway after a fresh token exchange, so the gateway can see the chain of actors and apply a tool check, plus a redaction, before the downstream service runs. A registry that lets an agent find a tool still ends at that gateway.

We use the order, not the scale they report: identity for this hop, then the check, then the system of record. Discovery in the layers we ship follows that order even when the catalog is one we retrieve from, rather than a mesh of the size described there.

The failure that remains, once the gate is in place, is a session that caches the first tool list and the first token and then serves later callers from the cache. The list goes stale when a tool is removed. The token still names the first principal.

The audience check can pass, because the token really was minted for that server, and the acting grant never gets a second look. Expire the cache with the run. When a grant is revoked, the next call refuses, and the refusal is an audit line rather than a run that continues on memory.

A plan can still be careful and unsafe if the call leaves through a key the plan never had to name, which is why the agents on our AI agents page sit on this path instead of calling systems on their own.

Common questions

Why is an API key in the prompt unsafe even if the prompt says to hide it?#

The key is now text the model can be asked to repeat, and a trace of the prompt is a copy of that key for anyone who can read traces. It also authorizes every operation the application can perform, for as long as the key stays valid, rather than the one operation this run needed.

A sentence in the prompt does not change either fact. Keep the secret in a vault the model cannot print, and mint a token whose tool list and lifetime match the run.

What counts as an escape-hatch tool?#

An escape hatch is a tool whose argument chooses the operation. A shell, a raw HTTP client, and a tool that accepts a SQL string are escape hatches, even when they appear in a catalog with a careful description. A policy can refuse a named operation such as a refund above a cap, because the amount is a field.

It cannot refuse a SQL tool without also refusing every statement the database user is able to run. If you still need a general tool, put it on a separate short-lived credential and require a person to approve the call.

Does a short-lived token replace a per-user grant?#

A short life limits how long a stolen token keeps working, and a tool list limits which operations it can call. Neither fact names the person this request is acting for. Store the grant on the acting user inside the tenant, and mint the run token from that grant.

A narrowly scoped token that belongs to the installer still returns the installer's data when someone else asks. If the request has no grant for the person who asked, refuse the call instead of falling back to a shared key so that an answer still goes out.

Why keep a record when the call never left?#

A log made only of allows cannot show whether the policy ran. A tool that is refused often is either described badly or scoped more tightly than the work requires, and the refused lines are how you tell which problem you have. A tool that is never refused may be correctly open, or it may be skipping the gate.

Write the refusal with the same run id, principal, tool name, and argument summary as an allow. Leave out the upstream result, because the call produced none, and leave out the secret.

Can discovery at call time skip the policy check?#

Discovery chooses which tool definitions the model reads on this step. The policy chooses whether the chosen call may leave the process. A search result is not an allow. If the run's token does not include the tool that came back, refuse the call and record the tool name, so the model cannot describe a permission gap as an empty search. A cache of the first tool list and the first token will serve a later caller as the first principal, so expire that cache when the run ends.

How is this different from an action boundary?#

An action boundary names one action: who it is for, which task it serves, which target it may touch, and which limits apply. This guide is the path that enforces that boundary on a live call, mints the credential for the call, and writes the allow or the refuse.

The boundary can be correct on a page while the agent still holds an application key and calls the API directly. You find out whether the boundary is real by looking at the token the process attached and at the line written when the check said no.

What should the model say when a call is refused?#

It should repeat the typed reason: which tool was asked for, and which check failed, in language a person can act on. It should not say that no records exist, that the operation succeeded, or that it will try a broader tool the token does not include.

An empty payload and a refusal are different outcomes, and the reply has to keep them apart. If the reason itself contains data the caller should not see, return the reason code and a short explanation, and keep the sensitive value in the audit line under the usual grant.

Further reading