How do you stop prompt injection from spreading between agents? Re-check trust at every hop before the next agent or tool treats the text as a task. At Bigcircle we let the parent propose the next step. The application binds it to the same principal, a narrower target, and a fresh expiry.
Each hop looks trusted from the inside, which is why a line in the parent's prompt does not bind the child that receives the text. A tool result from our server, a child called by our orchestrator, and a note written by our own run all look like inside traffic.
Syed Anas Mohiuddin's update on protocol pivoting names the assumption those hops share. Data that crossed the boundary is treated as trusted because it came from inside the system. He describes a tool result shaped like a task for the next agent, which the orchestrator then forwards as ordinary delegation.
This guide walks that chain in order and says what each hop may carry, what the check rejects, and how to test the refusal before another agent depends on the tool. The permissions boundary, the tool credential, the memory key, and the judge question are the same check in four places, and the copy itself is where the chain starts.
Why does a parent's output become the child's instruction?
A parent agent copies text when it delegates, and the child reads that copy as the task it was given. The parent may have been summarizing a page, a ticket, or a file in a repository, and a line in that text can ride along in the message.
The child has its own tools, often with a credential the parent does not hold, so an instruction that could not move the parent's tools can move the child's.
We keep the parent's message in a data field and keep the task in a separate field the application wrote. If those two fields are one string, the line from the page and the task the person asked for are the same tokens, and the child has no way to tell them apart.
Prompt injection is that mix inside one agent, and between agents the mix is copied into the next prompt and run again.
The update gives the concrete shape, with text coming back from an MCP tool, worded as a task for a subagent. The orchestrator passes it on as part of normal delegation, and the subagent runs it because the caller is the orchestrator. If that subagent can make an outbound request, the request leaves from the subagent's network position.
Nothing in that chain has to break for the text to gain reach. Each hop does the job it was built to do, and the line arrives with tools it did not have at the start.
A prompt that tells the parent to ignore untrusted text stops at the parent. The child sees a task from a system it was built to obey, and the next hop has to re-check that task.
What does each hop have to check?
The table is the check we apply when one agent hands work to the next, and the figure repeats it. It is a framework for placing a hop, not a survey of how often any hop is attacked.
| Hop | What crosses | Check before the text is trusted |
|---|---|---|
| Content | A page, a ticket, a repository file, or a tool result | Mark it as data. It does not set the task. |
| Parent | A proposed next step | The task stays the principal's task. A new verb in the text is not a grant. |
| Child | The parent's message | Bind the same principal, a narrower target, and a fresh expiry. Do not reuse the parent's credential. |
| Tool | The argument the model chose | Check a URL, a path, or a ref where it resolves, before a process or a request uses it. |
| Return | The tool output or the child's reply | Keep it as data. It cannot authorize the next hop, and it cannot be stored as a rule. |
This figure is a framework in which each box is one hop, and the line inside the box is the check that hop has to pass before its text is trusted. It states no rates and no incident counts, and on a narrow screen the drawing scrolls sideways inside the column while the table above carries the same checks.
How do you bind a hop to a person?
The principal on the child is still the person, role, or approved process that asked, which is the rule in how to scope AI agent permissions. The parent agent's identity shows which process sent the message. It does not show who asked for the action, or why this hop is allowed to run.
We write that binding again at the hop, because a valid request at the start of the run is otherwise still in force when the child adds a step. The new step can look like a continuation, and the boundary has to be written again before it runs.
The fields are the ones an action boundary already has. They name who asked, the purpose, the target, the parameters, the expiry, and a credential that dies with this hop. A child that may read one case does not receive a credential that can email the case out, even when the parent's message says to send it. The new verb is a new action, and it waits for a new decision.
On the systems we ship, that decision sits outside the model. The child can argue that the send is part of finishing the task. The application compares the verb and the target with the boundary for this hop, and a miss comes back as a typed refusal the child can repeat. A retry that only rephrases the same text is the same action, and it gets the same refusal.
Where do the credential and the memory get checked?
A credential issued for the parent is the wrong credential for the child, even when both agents sit in one product. MCP tool access in production keeps the secret on the server and scopes it to the call.
The hop is a tighter cut, so the child's call gets a session capability for its own target, and the parent's token is not attached to that call. When the parent is the only identity on the call, a line the parent forwarded spends the parent's reach.
The tool still has to check the argument, because the caller being our agent is the assumption that fails. The argument may be text the agent only read.
NVD records CVE-2026-14540 against Google's mcp-toolbox for versions 0.3.0 through 1.4.0, and it scores the issue 8.0 on CVSS 4.0. A crafted path can make the HTTP client follow a redirect to an internal address. The record says the client had no redirect policy and no check on the target address.
Pull request 3448, merged 18 June 2026, checks the address when the connection opens. The pull request describes this as a guard against DNS rebinding, because the name can change between the check and the connection. It also adds allow and block ranges for IP addresses, and it rejects an unsafe base URL when the server starts.
The same update records four more fetch checks the affected teams confirmed and fixed: JPMorgan Chase, Weaviate, France's DINUM, and the Tangerang City government. We cite those as his account of fixes the teams confirmed. The mechanism we take from the set is narrow. A URL the model supplied is checked on the tool, where it resolves, and a prompt upstream of that tool is not the check.
Content in a repository can take the other route, into the argument rather than into a URL. GitHub advisory GHSA-8g28-rj54-p5p2 scores the issue 8.2 and marks security-agent-mcp-server 0.2.0 as the patched release. The advisory's example is content in the scanned repository steering the revision the diff scan passes to git.
NVD lists the same issue as CVE-2026-97662, in versions before 0.2.0. A crafted reference is treated as a command option, and the server can write outside the workspace. Testing MCP servers before agents depend on them pins that class of argument on the pull request, where the file has already become the argument.
A reply is a hop too, and a line that says to store this as a rule is still data. Writing that line into shared memory makes it the instruction the next agent reads.
Multi-user agent memory and credential isolation keys a fact to the person who asked to remember it. A tool result stays on the conversation until the thread ends, and a later run does not load it as policy.
How do a judge and a SQL call stay on this rule?
A judge and a SQL tool are two more hops, and both of them read text the previous hop produced. Introducing Jev into agentic workflows keeps the question text fixed and puts tool arguments, pages, and model output in named fields. If the parent's message is concatenated into the question, the injected line becomes the question, and a high score approves the injection.
The SQL path has the same cut. Building production-grade analytics agents keeps the user's question, the plan, and the SQL in named state, with the question text versioned and fixed.
A clause that arrived from a table, or from the parent agent's summary, stays in a field the parser and the EXPLAIN gate can see. Folded into the question, that clause can rewrite which query the gate is judging, and the gate will then enforce the rewritten question.
The application refuses to build the question from untrusted fields. A hop that wants a different decision has to change the versioned question, which is a release rather than a tool result. We do not ask the judge to notice the injection, because the question it sees was already fixed before the hop ran.
How do you test a hop before another agent depends on it?
Those checks fail open when the only exercise is a model completing a normal task. The model sends the arguments that finish the task. It does not send a path that redirects inward, and it does not send a revision a file suggested so the process would read it as an option. The run stays green, and the hop was never attempted.
Testing MCP servers before agents depend on them splits the work into a contract on every pull request and an agent eval on a schedule, or when the tool list changes.
The contract calls the server with no model and requires the refusal: a URL that resolves to an internal address, a ref the process would read as an option, a session the server did not issue to this caller. The agent eval scores whether a model stops when it sees that refusal. Rubric design for that set lives in how to build LLM evals.
We add a contract case when the server should never have produced the reply, and a frozen task when the server was right and the model used the reply as a new task. The second case is the spread, where the tool did its job and the next hop treated the output as instructions.
Both have to be able to fail a change, or the chain is only tested on the path a normal task happens to take. A change that only the happy path exercises will stay green while the hop is open.
Common questions
How do you stop prompt injection from spreading between agents?#
Re-check trust at every hop, in the application, before the next agent or tool treats the text as a task. Bind the child to the same principal, a narrower target, and a fresh expiry, and validate tool arguments where they resolve. A line in the parent's prompt that says to ignore untrusted text does not bind the child, because the child never sees that prompt.
Why does the check have to run again on the child?#
The parent can be told to treat a page as data and still copy a line from that page into the task it sends. The child trusts the parent because the parent is the orchestrator, and the child may hold a credential the parent does not. The task, the credential, and the tool arguments are checked again on the child, or the copied line spends the child's reach.
What should a delegated task include?#
Name who asked, the purpose, the target, the parameters, the expiry, and the credential for this hop only. A summary that adds a new verb, a new destination, or a new recipient is a new action, and it needs a new decision. The parent's identity shows which process sent the message. It does not replace the principal the boundary was written for.
Does a result from our own server count as trusted?#
A result our server returned can still carry text from a page, a ticket, or a repository the server read. Check URLs, paths, and refs on the server before a process or an outbound request uses them. Keep the result as data for this conversation. Do not write it into memory as a rule the next run will follow.
How do you test that a hop refused the injected text?#
Call the server with no model and send the value a normal task will not send: a URL that resolves to an internal address, a ref a process would read as an option, a session the server did not issue to this caller. Require a typed refusal, and score whether a model stops when it sees that refusal in a separate eval, on a schedule or when the tool list changes.
How is this different from scoping one agent's permissions?#
A permission boundary states what one action may change, for whom, and until when. This guide follows that output when it becomes the next agent's input. You still write the boundary. You also apply it again at the next hop, because the text that arrives there was not part of the original grant.
Further reading
- How to scope AI agent permissions. The action boundary the hop re-applies: principal, target, parameters, and expiry.
- MCP tool access in production. The credential stays on the server, scoped to the call, and the parent's token is not reused.
- Multi-user agent memory and credential isolation. A fact is stored for the person who asked, and a tool result is not written as policy.
- Introducing Jev into agentic workflows. The judge question stays fixed, and untrusted text stays in named fields.
- Building production-grade analytics agents. SQL stays in named state, under the parser and the EXPLAIN gate.
- How to build LLM evals. Rubrics for the agent eval that scores a stop on refusal.
- Testing MCP servers before agents depend on them. The contract that refuses the bad argument on the pull request.
- Prompt injection, AI agents, MCP, and multi-agent orchestration.
- Syed Anas Mohiuddin, Protocol Pivoting, four months later. Vendor-confirmed SSRF fixes at Google, JPMorgan Chase, Weaviate, France's DINUM, and the Tangerang City government, and the case where a tool result is forwarded as the next agent's task.
- NVD, CVE-2026-14540, published 31 July 2026. SSRF in Google mcp-toolbox 0.3.0 through 1.4.0, CVSS 4.0 base score 8.0. The patch is pull request 3448, merged 18 June 2026.
- GitHub advisory GHSA-8g28-rj54-p5p2, 1 October 2026, patched in security-agent-mcp-server 0.2.0 and scored 8.2 there. Content in a scanned repository can steer the diff-scan revision. NVD CVE-2026-97662.