How do you test an MCP server before an agent depends on it? At Bigcircle we split that into two layers, and a demo where the model called a tool once is neither of them.
A contract test runs on every pull request, with no model in the loop. It pins the catalogue, the arguments, the session, and the shape of an error. An agent-in-the-loop eval runs on a schedule, or when that catalogue changes, and it asks whether a model actually uses the server.
The hard part is that the demo and the contract fail in different places. A model on the happy path will not send a ref that starts with a dash, and it will not present another tenant's session id.
The call returns, the trace looks finished, and the argument handling and the session binding were never exercised. A contract test can prove the server refused a foreign session. It still cannot prove that a model will stop when it sees that refusal.
This guide takes the two layers in that order. It says what the contract pins, what the agent eval adds, and which failure each layer is built to see.
Security bugs in the server belong in the contract. The later sections use two public advisories as the shape of those tests, then keep the model call off the pull request.
What does the contract test pin on every pull request?
The contract test speaks the protocol the way a client will, and it fails the pull request when the reply drifts. There is no model, no prompt, and no score that can be talked into a pass. We keep a snapshot of the catalogue in the repo, call the server through a client, and diff the reply against that snapshot.
A catalogue drift is a renamed tool, a required field that became optional, or a description that now tells the model to do a different job. The agent will follow the new description on the next run, and the pull request that changed it can still be green if nobody pinned the old text. Pin the names, the input schema, and the description. A description is part of the contract, because it is the instruction the model will read.
Arguments are the next pin, and they are where a tidy tool name still hides a program. MCP tool access in production separates a declared tool from an escape hatch, and the contract test is how that separation stays true after the next edit. Require the fields the schema requires, reject a type the schema does not name, and reject a value that a process would read as an option rather than as data.
The official TypeScript SDK client is one way to make those calls, over stdio or in-process, with the server under test and nothing else attached. The official Inspector is the visual client for the same conversation, and a person drives it when they want to see one failing call.
We do not put the Inspector on the pull request. A person cannot be the gate, and a click-through that passed on the last build does not pin the schema in this diff.
Session binding is a contract, not a prompt. The server issues the session id, stores it against the caller and the endpoint, and refuses a later call whose header names a session it did not issue to that caller. A client-supplied id, on its own, is not a credential. If the test can present another caller's id and still run a tool, the pull request fails, whether or not an agent has ever connected.
The error shape is the last pin on this layer. A denial and an empty result have to be different payloads, because the model will narrate them the same way if they look the same. Return a typed refusal the model can repeat, which is the structured output the access guide already asks for. The contract asserts the status, the tool name, and the reason code, and it fails if a denied call comes back as an empty list with a success flag.
What does an agent-in-the-loop eval add?
The agent eval asks a question the contract cannot: given a task, does a model pick this tool, recover when the tool returns the typed error, and stop when the tool refuses. Those are properties of the model plus the catalogue, so they move when either one moves. A schema test that passed this morning says nothing about a model that now prefers a different tool with a similar description.
We run that eval off the pull request, on a schedule, and again when the catalogue snapshot changes. The schedule catches a model swap or a prompt change that never touched the server.
The catalogue trigger catches a description edit that the contract accepted, because the new text matched the snapshot, while the model no longer does the task. Both runs use a frozen set of tasks, the same cases each time, so a failure is a diff against the last green run rather than a new anecdote.
Rubric design for that set lives in how to build LLM evals. This guide does not repeat it. What belongs here is the cut between cases: tool choice, a typed recovery, or a stop on refusal, and the run may call only the server under test. A general agent eval that can reach other systems will pass for reasons that have nothing to do with this server.
The runner can be whatever already scores your other evals. The layer is the frozen tasks and the cadence, not a product name. A single successful tool call in a vendor walkthrough leaves the gap open, and a second walkthrough does not close it.
If the score can move because the model sampled differently, keep the run off the pull request. Keep a person who can read a disagreement before the catalogue is treated as ready for other agents.
Agent playbooks as MCP tools and MCP tool design: fewer, larger tools decide what is in the catalogue. Once that shape is chosen, the agent eval is how you learn whether a model can still follow it. A larger tool with a vague description will pass a schema test and fail the task set, which is the result you want before an agent in production depends on the server.
Which failure does which layer see?
Each failure has a layer that is built to see it, and a layer that will stay green while the failure sits there. One run that tries to do both jobs is a slow pull request that still misses the argument bug, or a fast schema test that never notices the wrong tool.
The table is the assignment we use. The figure repeats it. It is a framework for placing a failure, not a survey of how often any cell happens.
| Failure | Contract test, every pull request | Agent eval, on a schedule or when the catalogue changes |
|---|---|---|
| Catalogue drift | Pin tools/list. A renamed field or a rewritten description fails the diff. | Sees the change only if a frozen task calls that tool. |
| Argument handling | Reject the value before a process starts. A leading dash is not a revision. | A normal task never sends the bad value, so the run stays green. |
| Session binding | A session id the server did not issue to this caller is refused. | One tenant's demo still passes. The other tenant's id was never sent. |
| Error shape | A denial is a typed refusal, not an empty list with a success flag. | Score whether the model repeats the reason and stops. |
| Tool choice | No model is in the run, so the test cannot see which tool would be picked. | Score the choice, the recovery from the typed error, and the stop. |
This figure is a framework. The rows are failure types and the columns are the two layers, and a cell says whether that layer is built to see the failure. It states no rates and no incident counts. On a narrow screen the drawing scrolls sideways inside the column, and the table above carries the same assignment in text.
Where do security regressions belong?
Security bugs in the server are contract failures, and they belong on the pull request for the same reason an argument bug does. An agent that is doing the task will not explore them. If the only test is a model completing a normal task, the bug ships, and the first caller who is not doing the task finds it.
AWS security bulletin 2026-121-AWS, published 1 October 2026, assigns CVE-2026-97662 to the open-source security-agent-mcp-server in the awslabs/mcp repository. NVD describes an argument injection in the diff scan. A crafted reference is read as a command-line option rather than a revision, and versions before 0.2.0 can write files outside the workspace.
The current source documents the command as git diff, then the caller-supplied ref, then a double dash. The separator is present, and it sits after the ref, so a value that starts with a dash is still an option. The 0.2.0 fix rejects an empty ref, and a ref that starts with a dash, before git runs.
A contract test for that class of bug does two checks. Looking for a double dash in the argument list is not one of them, and that check would have passed on the vulnerable command.
Assert that a ref starting with a dash is rejected before any process is spawned, and assert that an ordinary ref is passed as a revision. The agent eval will not send the dashed ref on its own. A case that sends it is a contract test, and attaching a model adds nothing.
MetaMCP, the open-source gateway from metatool-ai, is the session half of the same cut. NVD's record for CVE-2026-79537 says that through version 2.4.22 the session store is keyed only by the client-supplied mcp-session-id header. There is no owner, namespace, or endpoint binding, so another tenant's id can run that tenant's tools.
CVE-2026-79538, also through 2.4.22, is code execution on the inspector proxy, where the stdio handler takes process parameters from the request. The project's GitHub releases still listed 2.4.22, published 19 December 2025, as the latest release when checked on 7 October 2026. CyberSecureToday reported no patched release as of 4 October 2026.
The session test issues a session as one caller, presents that id as a different caller, and requires a refusal. A second test refuses a launch whose command comes from the request body rather than from a server record the caller owns. Neither test needs a model. An agent signed in as one tenant, doing one of that tenant's tasks, will not present the other tenant's session id, so the eval stays green while the binding is missing.
Dual-layer production agent ops covers the time after something is in production and a trace looks successful for the wrong reason. This section is earlier than that. The regression should fail the pull request that introduces it, which is the contract, and the agent eval is the check that a model still behaves once the contract holds.
How do you keep the model out of the pull request?
Model calls stay off the pull request because they are slow, they vary between runs, and they miss the failures the contract is there to catch. A red agent eval blocks a merge for a reason the author cannot reproduce from the diff. A green one lets an argument bug through.
We run the contract on every change to the server. We run the agent eval on the schedule, and again when the catalogue snapshot in the pull request differs from the one on the main branch.
That trigger is a diff of the pinned tools/list, not a judgment that the change might affect tools. A description edit is a catalogue change even when the schema is identical, because the model reads the description. A comment in the server that does not change the snapshot does not need a model, and the contract still runs on that pull request.
When the agent eval fails, add a contract case if the server should never have produced that reply. Add a frozen task if the server replied correctly and the model used it badly. A prompt change that hides a schema bug will pass until the next model, and then the same bug is back.
Observability on the eval run keeps the tool result, so you can tell which of the two fixes you need without calling the model again to remember.
A server that has only been clicked through in the Inspector can still be the right server to build on. It stays a demo until the contract has failed a bad change on a pull request and someone has read the catalogue eval. That is the point at which other agents may depend on it. MCP is the protocol those calls use, and tool calling is the moment the choice leaves the model.
Common questions
Why not test the server by letting an agent call it?#
An agent on a normal task sends the arguments that complete the task. It does not send a ref that starts with a dash, and it does not present a session id that belongs to someone else. The call succeeds, and the two failures the contract is for were never attempted. Use the agent eval for tool choice and for stopping on a refusal. Use the contract for everything the server must do the same way on every call.
What should a contract test assert about arguments?#
Assert the schema, and assert the values a downstream process would misread. A required field that is missing fails. A type the schema does not name fails. A ref that starts with a dash fails before any process is spawned, because a separator that comes after the value does not stop the process from treating the value as an option. A test that only checks for the separator would have passed the vulnerable diff scan.
Does a passing agent eval replace the contract test?#
A passing agent eval leaves several contract holes untouched. The catalogue can drift for a tool no task calls, a foreign session id can still be accepted, and a denial can still look like an empty result. The contract is the pin for those replies. The agent eval answers a different question, which is whether a model uses the server that the contract describes.
How do you test session binding?#
Issue a session as one caller and present that same id as a different caller, on a different endpoint if the server has more than one. The call must be refused, and the refusal must name the check rather than return another caller's data. Also refuse an id the client invented and an id the server has already expired. If any of those calls runs a tool, the contract failed.
When should the agent-in-the-loop eval run?#
Run it on a schedule, so a model or prompt change is caught even when the server did not change. Run it again when the pinned catalogue in the pull request differs from main, including a description edit with the same schema. Keep it off the other pull requests. A contract failure is enough to block those, and a model call on them adds time without covering the bug.
How is this different from production agent ops?#
Production ops watches traces after users are on the system, and it separates a bad answer from a failed call. This guide is the gate before an agent is allowed to depend on the server. The contract fails the change that broke the server. The agent eval fails the change that left the server correct and made the model use it badly. You still need that ops split once both of those have passed and real traffic arrives.
Further reading
- MCP tool access in production. The token, the pre-call check, and the allow or refuse record that the contract is pinning.
- MCP tool design: fewer, larger tools. What goes in the catalogue before you pin it.
- Agent playbooks as MCP tools. A playbook stays a declared tool when its inputs are typed, and the contract treats it as one catalogue entry.
- How to build LLM evals. Rubrics and a frozen set for the agent-in-the-loop layer.
- Dual-layer production agent ops. Quality and infrastructure after the server is already a dependency.
- MCP servers. Scoped tokens, a typed tool contract, and an invocation record. This guide is how that contract gets tested before an agent depends on the server.
- MCP, tool calling, evals, and guardrails.
- AWS Security Bulletin 2026-121-AWS, 1 October 2026, and NVD CVE-2026-97662. Argument injection in security-agent-mcp-server before 0.2.0.
- NVD, CVE-2026-79537 and CVE-2026-79538. MetaMCP through 2.4.22. Session dispatch keyed only by the client-supplied mcp-session-id header, and command execution via the inspector proxy.
- CyberSecureToday, MetaMCP unpatched RCE and session hijack, checked 4 October 2026. No patched release as of that date. The GitHub releases list still showed 2.4.22 as the latest release on 7 October 2026.
- Model Context Protocol Inspector, the official visual client, and the TypeScript SDK, one way to drive the contract with no model attached.