Guides

Agentic Systems

What a production AI agent costs to build and run

How the build, the run, and the keep-alive each move the bill, and how a cap in the request path stops a loop before the invoice does.

By Tirth Gajjar · Founder & CTO

14 min

At a glance

A production agent costs the build, the run, and the keep-alive. The model line moves after launch, and a cap that only sends an email has already paid for the calls in between.

Use this guide to Agentic Systems to review the design choices and checks for your system.

Who this is for
Engineers building AI systems and technical leads reviewing the implementation.
Topics
  • AI Agent
  • Evals
  • Model Context Protocol (MCP)
  • Tool / Function Calling
  • Guardrails
  • LLM Observability & Tracing

Published

What does a production AI agent cost to build and run? At Bigcircle we split the answer into the build, the run, and the keep-alive, because one number hides which of the three is about to move. The build is the system that can take a real task, and the run is the model calls, tool calls, and child tasks that task then spends.

The keep-alive is the work that remains after the first week that looked finished. A demo that answers once pays for a prompt and a happy path, while an agent other people rely on also pays for declared tools, a scoped credential, and a stop before the next paid call. If a script or one query can do the job, when not to build an AI agent is the prior cut, and this bill does not start.

The sections below take the three buckets in that order, then the stop we put on the request path, where a hard cap covers turns, tool calls, and fan-out, and a breach ends the run before the next paid call. The ledger sits where the agent cannot delete it, and our own meter gets checked against the provider bill.

Who staffs the build is the engagement guide, and the loop itself is how to build production AI agents, so read those before you price a token, because the staffing choice changes which bucket you are even quoting. Some quotes name only the model. Those quotes have already skipped the build.

What does the build include?

The build is the work that has to be true before the first real task leaves the team. It includes the tools the job actually needs, permission to call them, an eval set, and a path for the step that fails. Choosing a model is one decision inside that work. It does not stand in for the rest.

On the systems we ship, the long part of the build is the boundary with the rest of the company, and a cloud change has to read the live account before it plans, which is the shape of the DevOps agent we shipped. A chief-of-staff task has to cross mail, calendar, documents, and the systems of record, which is the shape of the executive operations agent.

Each integration is a tool contract, an auth path, and a failure the agent will meet on a weekday, and permission and evals sit in the same build, or the first week of traffic teaches them as incidents. MCP is one way to declare the tools, and the production pattern mints a token for the run.

A policy check runs before the call leaves, which MCP tool access in production covers on its own, and evals are how you know a change still does the job, with how to build LLM evals as that set. A build quote that lists model access and skips both will be rewritten once a real task fails.

The hire page prints three ways in, an embedded engineer, a pod, or a fixed outcome, and neither that page nor the engagement guide publishes a rate. This guide will not invent one, because headcount and price are set on the call, and what you can compare before then is which bucket the quote names.

What makes a run grow past one reply?

The run costs more than one reply because the agent chooses how many steps to take, and each step sends the history again, so a retry, a tool result, and a child agent each add calls the first estimate left out. The price per token can fall while the run still grows, because the number of calls grew faster than the price fell.

A workflow pays a known number of model calls, one per step you wrote. An agent pays a number that depends on the trace, and tool calling is where that number leaves the model and starts touching other systems. A tool may call a model of its own, and fan-out starts several children on that same fresh-counter mistake.

The reasoning models page walks one call where thinking tokens are billed with the output, so the visible answer costs more when the model deliberates. An agent repeats that shape on every turn, with a longer transcript each time. That growth is why a single-reply estimate misses the run.

Prompt caching is the discount on a repeated prefix, and a retry that misses the cache pays the input again on the full transcript. Count failed calls in the same total, because the provider already counted them, and a meter that drops them sits under the invoice.

What does keep-alive keep costing?

Keep-alive is the work after the first correct week, once the launch itself is no longer the job. It covers the eval set when a tool or a prompt changes, and a person who can act when a run stops. Both stay on the calendar after that week. They show up later as an incident.

The same work includes a trace you can still read when someone asks what happened, and a launch date with nobody named for that work leaves it on whoever next opens the invoice. We name that person before launch. A stop with no owner is a log line.

Observability on an agent is the trace of the run, stored with the run, rather than a second product you buy when finance asks. Dual-layer production agent ops separates the quality of the answer from the health of the path that produced it. Both layers keep taking attention after launch, because a model swap or a tool schema change can move the result while the code stays still.

A new document in the corpus can move the result the same way. The cases we publish report operational results from those builds, and neither those pages nor this guide publishes a keep-alive price. Use the pages for the shape of a shipped agent, then price upkeep from your volume, your eval cadence, and the person on the stop.

How do the three buckets compare?

Read the bucket before you compare two quotes, or a missing line will look like a saving. Two proposals can name the same agent and still omit a different bucket, and the shorter page often left the keep-alive or the stop off the list. The table is the comparison we use in a scoping conversation, and the cells name drivers you can scope against.

BucketWhat moves itWhat a thin quote skips
BuildTools, credentials, evals, and the failed stepA prompt and a happy path
RunTurns, tool calls, retries, fan-out, and a transcript that growsA price per message
Keep-aliveEval upkeep, a person on the stop, and traces you can still readA launch date with nobody named
Build, run, and keep-aliveFramework. Drivers for a scoping conversation. No prices and no shares.BuildToolsThe systems the task touches.CredentialsA token for this run.EvalsA set before the first week.RunTurns and retriesHistory goes out again.Tool callsEach call can spend.Fan-outChildren share one cap.Keep-aliveEval upkeepAfter each real change.A person on the stopSomeone can act.Traces that stayReadable next month.On the request path, before the next provider callCheck the capsTurns, tools, children.Still insideThe call may leave.BreachEnd the run.Ledger written outside the tool list.The agent credential cannot delete or rewrite it.

The figure is a framework and an illustrative model of a scoping conversation. It is not survey data, and it states no prices and no shares. We use the drivers in that figure when we scope a build, before anyone asks for a single number.

Green is the build you pay before the first real task, and amber is the run, which moves with how many calls a task makes. Blue is the keep-alive that remains after that task looks finished. The row under the columns is the check before the next provider call.

The caps still hold and the call may leave, or a breach ends the run. The ledger under that row is written outside the tool list, where the agent credential cannot delete it. That ledger is the record you reconcile against the provider bill.

How do you enforce a budget in the request path?

A budget holds only when the check runs before the next provider call and a breach ends the run, and a climbing dashboard and a later email have already paid for every call between the threshold and the reader. The cap is code on that path, and it limits turns, tool calls, and how many child agents a run may start.

Fan-out is the case a per-call limit misses, because each child looks small on its own. Together they spend the parent's allowance a second time if each child is handed a fresh counter. The child spends what the parent has left, and the parent refuses another child once the count or the allowance is gone.

The same rule covers retries, because a failed call still counted once the provider counted it, and guardrails that live only in the prompt can be talked past, while a counter the model cannot reset cannot. We put the counter in the request path, outside the prompt, so a polite request to continue cannot raise it.

On 3 October 2026, Simon Willison argued that pay-by-usage services need a default hard cap. Past a configured monthly amount, that cap cuts the service off and returns errors, and a warning email does not do the job. He is describing a backstop you set before the month starts, so the invoice is not the first time the total shows up.

He points at AWS spending limits, from the vendor announcement of 16 September 2026, which pause a project when usage reaches the limit. He also points at Google Cloud Spend Caps from July 2026, a monthly cap on specific services inside a project. We treat both notes as his reading of those announcements, and a provider cap that pauses the account is only a backstop.

The account-level pause arrives after many runs have already spent. It does not know that this run is a retry storm, so the loop still needs its own check before the next provider call leaves.

Jatin Bansal, writing on 26 May 2026, gives two accounts of what that gap costs when the only signal is a dashboard. In his telling, four LangChain agents fell into a loop between an analyzer and a verifier in November 2025. The bill reached $47,000 over eleven days, with observability in place and no stop.

He also describes a 35-engineer company whose April 2026 bill was $87,000, of which one developer's weekend of autonomous refactoring was $4,200. Those figures are his account of those incidents, author-reported, and they were not audited here. They are not Bigcircle measurements, and this guide does not treat them as a market rate.

What we take from the write-up is the order of the check. The request path tests a turn cap, a tool-call cap, a token or spend ceiling, and an external stop before the next call. Any one of those ends the run before another provider call leaves, and the partial result is kept.

He also writes that common SDK defaults are a turn cap of 10 in the OpenAI Agents SDK and 25 in LangGraph. The Vercel AI SDK default he reports is 20, and that set is his report of the defaults. Treat the numbers as his report. Copy a default only after your own traces.

A default is a calibration for someone else's median task, so set the cap from runs you have already finished. Leave room above the runs that completed, so a stuck loop hits the cap and a normal task does not. A cap set on a guess will either cut good tasks or let a loop run.

The ledger has to outlive the agent, or the loop can erase the only record of what it spent, so write the run id, the tool, the allow or the refusal, and the token counts to that store. Add the stop reason, and keep the store where the agent's credential cannot delete or rewrite it.

A tool that can clear logs can clear the evidence of the loop you are trying to bound. Include the calls that failed and the calls a child made, or the meter will sit under the invoice on the next close. You will otherwise argue about a gap that was a missing row.

Reconciling the meter with the provider bill is a close you run on a schedule, and TechnoLynx describes it as two sums. One sum is billed spend from the invoice, and the other is instrumented cost from your own traces. Retries, timeouts, and rejected outputs stay in the second sum, or the gap is only the failures you dropped.

A gap that remains after failed attempts are included is work you are not tracing, such as an eval job or a warm-up call. A second service on the same key opens the same kind of gap. Fix that gap before you treat either number as the cost of a successful task.

The three write-ups above are operator- and author-reported accounts, or a method for comparing two ledgers. None of them is a Bigcircle data set, and we cite them for the check and the comparison. The dollars stay with the author. They belong to the person who reported them.

Common questions

What does a production AI agent cost to build and run?#

The cost is the build, the run, and the keep-alive, and each one moves for a different reason. The build is tools, credentials, evals, and the failed step, while the run is turns, tool calls, retries, fan-out, and a transcript that grows. A price per token covers one line inside the run and leaves the build and the keep-alive unpriced.

Why does the model bill stay the smaller line?#

The model bill moves with how many calls a task makes, while the build is the integrations and the eval set you pay first. Teams that price only the token rate meet the build when the first real task fails. They meet the keep-alive when nobody owns the next change, and failed calls still belong in the run because the provider already counted them.

How do you stop a runaway agent before the invoice?#

Put a hard cap in the request path, and check it before the next provider call. Cap the turns, the tool calls, and the number of child agents, and make each child spend the parent's remaining allowance. When any cap breaks, end the run and keep the partial result. An alert after the threshold has already paid for the calls made in the gap.

Are published incident totals a budget you should copy?#

Treat a published total as the author's account of that incident, including a four-agent loop billed at $47,000 over eleven days. A separate April 2026 bill of $87,000 came from a 35-engineer company, with $4,200 of it from one weekend. Use them to see that a dashboard without a stop can run for days.

How do you reconcile your meter with the provider bill?#

Sum what the provider invoiced for the period, and sum what your traces recorded for every attempt, including retries, timeouts, and outputs you rejected. Compare the two sums before you call either one the cost of a task that succeeded. A remaining gap is work you are not tracing, such as an eval job or a second service on the same key.

When should you spend on a workflow instead?#

Spend on a workflow when someone can write the steps, and spend on retrieval when one query can find the passage. The agent loop costs more to trace on a job those two shapes already do. Make that cut before you staff the build, or you will fund the keep-alive for a sequence the code could have walked.

Further reading