Guides

Agentic Systems

When not to build an AI agent

How to tell a scripted workflow, a retrieval app, and an agent loop apart before you spend the build on the wrong one.

By Tirth Gajjar · Founder & CTO

14 min

At a glance

A workflow is enough when the steps are already written down. Retrieval is enough when one query can find the passage. An agent earns the loop only when the next tool depends on a result you could not name in advance.

Use this guide to Agentic Systems to review the design choices and checks for your system.

Who this is for
Engineers building AI systems and technical leads reviewing the implementation.
Topics
  • AI Agent
  • Retrieval-Augmented Generation (RAG)
  • Tool / Function Calling
  • Guardrails
  • Evals
  • LLM Observability & Tracing

Published

Do we need an agent, or is a workflow or a RAG app enough? At Bigcircle we answer from the path the work takes on systems we ship, not from the label on the ticket. A scripted workflow is enough when you can write the steps and the branches.

A RAG app is enough when one query finds the passage and the answer has to stay inside it. An agent earns the loop when the next tool depends on a result you could not list before the run started.

The hard part is that a demo hides the difference. Each shape can call a model, show a tool, and return a sentence that sounds finished. The split shows up when a step fails or a record is missing.

A workflow stops on the branch you wrote, and retrieval says the passage is absent. An agent has to choose another tool, or stop, from evidence it just received. If you cannot point at a case where that choice is real, the loop is overhead you will spend the next quarter tracing.

This guide takes the three shapes in the order a team usually reaches for them: a scripted workflow, a retrieval app, then an agent loop, and then places five signals in one matrix. The last two sections are the signals teams waive most often: a write that can hurt before any gate exists, and a success check only the model can see. How to build the loop, how to raise one action's authority, and how retrieval compares with fine-tuning are other guides.

When is a scripted workflow enough?

A scripted workflow is the right shape when someone who knows the job can write the steps, including the branches. Invoice extraction, support triage with a fixed set of labels, and document classification usually sit here, which is also how we treat them in how to build production AI agents. The code walks classify, retrieve, draft, check, and send, and the step after a model call stays the one you named. A model may fill one step.

Put an agent on that path and the model spends the run re-deciding a sequence the code already had. A failure then looks like a choice, so you open the trace, when the missing piece was a branch nobody wrote. We keep the model inside the step that needs language, and we leave the order of steps in code. A reviewer can then point at the branch that fired, which is the record you want when the output is wrong.

A branch you can state as a rule is still a workflow. If the amount is under the cap and the order is inside the return window, draft the refund and wait for a person to approve it.

The model can draft the note. Leave the cap in the code that checks the amount. When the rule needs a number, a date, or an account id, compute that value in code and pass it in. A fact you have to look up is a different problem from a value you can compute, and that problem is the next cut.

When is retrieval enough on its own?

Retrieval is enough when the failure is a missing or stale fact and one query can find the passage that holds it. A refund window that changed in March, a policy clause, or a filing the model was never trained on are RAG problems, or plain search problems.

An agent does not repair a weak index. Our guide on RAG vs fine-tuning vs agentic search walks those layers in the order they cost. This section only asks whether the task has left that table too early.

If the question maps to a query, build the query. Hybrid search so an identifier still matches, reranking so the right passage reaches the window, and a permission filter on the query itself are the work that changes the answer.

A loop that searches until it feels sure will call the same weak index several times and then write a fluent sentence over thin evidence. You pay seconds, and several model calls, for a recall problem one better query would have shown in the log.

Agentic search is the case that survives that test, because a question about which suppliers mentioned tariff exposure in their last two filings has no single query. You need the supplier list, then the filings, then a read, then a decision that you have enough.

That is a loop, with a stop condition and a check that the cited span was in the passages that came back. If you cannot name the second query the first result forces, you do not have this case yet, and the workflow or the single query is still the build.

When does the next tool have to be chosen?

An agent earns its complexity when the next action depends on an observation you could not encode as a branch. Debugging a failing integration, investigating a metrics anomaly, and reconciling a messy data room are the shapes we keep meeting. The model selects a tool, reads the result, and either continues or stops. How to build production AI agents is that loop: observe, reason, act, verify, and this guide is the decision to build it.

Tool recovery is the signal that gets skipped, because a timeout on a known call, followed by the same call again, is still a retry inside the workflow. A 404 that means "try the archive, and if that is empty ask for the account id" is a choice among tools.

If every recovery you care about is already a row in a table, keep it in the table. The moment the error has to pick the tool, the workflow has run out of rows. The loop starts to pay for the tracing it adds at that moment, and not before.

Inside the loop a miss gets longer than it would in a workflow. A wrong branch fails one step. A wrong choice can take several more tool calls before anyone sees a bad final state.

The loop needs a budget, a stop rule, and a verifier that does not trust the model's claim that the work is done. Without those three, you have taken on the cost of an agent and kept none of the control. The five signals below are how we place a task before that cost is committed.

How do five signals place one task?

Five signals are enough to place a task, and one product can sit in different columns on different rows. Read down the row for the task in front of you before you pick a column for the whole product.

Triage can be a workflow, the policy answer can be retrieval, and the ticket whose next system depends on the last lookup can be an agent. A single label on the product hides which of the three you shipped, and the next failure goes to the wrong owner.

SignalScripted workflowRAG or searchAgent
Fixed pathUse it. The code walks the steps you listed.Use it when the path is one query and one answer.Skip it. The model would re-choose steps you already listed.
Multi-step judgmentUse it only when every branch is a rule you can write.A retrieved passage does not choose the next action.Use it when the next judgment depends on the last result.
Tool recovery neededUse it when the retry is the same step, or a known fallback.A search miss is a retrieval problem. Fix the index.Use it when the error has to choose a different tool.
High blast radius, no gateKeep the write out. The read path can stay in code.Keep a permission filter on the query.Do not start. The loop does not contain a bad write.
Measurable successScore the output of each step.Score recall, and whether the answer stays inside the passage.Score the stop, and a final state outside the model's claim.
Which shape the signal actually fitsFramework. Fit judgments for one task. No frequencies, because the mix is yours.Scripted workflowRAG or searchAgentFixed pathUse the scriptCode walks the steps.One queryThen one answer.Skip the loopSteps are already listed.Multi-stepjudgmentOnly with rulesEach branch written down.Does not chooseA passage is not a plan.Use the loopLast result decides.Tool recoveryneededKnown fallbackSame step, or a row.Fix retrievalA miss is the index.Use the loopThe error picks the tool.Blast radiusno gate yetKeep the write outRead path can stay.Filter the queryWho may read the passage.Do not startThe loop is not a gate.MeasurablesuccessScore each stepOutput you can check.Score the citationRecall, and the span.Score the recordStop, and final state.Green is a fit. Amber needs the condition in the cell. Red means stop and build the other shape.

The figure is a framework for the same five signals, set against a scripted workflow, a RAG or search app, and an agent. The cells are fit judgments you can argue with on a real task, and the figure states no frequency for any cell.

Green marks a fit for that signal. Amber means the shape survives only when the condition in the cell is true. Red means you stop and build the other shape, because this one adds a way for the bad outcome to happen.

A task that is green in the workflow column and red in the agent column should stay a workflow even if the roadmap says "agent." The word on the roadmap does not change which tool the error is allowed to pick. When a row is red because a write can hurt and no gate exists yet, read the next section before you add a loop.

What happens when a bad write has no gate?

A write that can move money, access, a customer message, or production state has a blast radius, and a loop does not shrink it. If you do not yet have a gate that can refuse the call, a credential scoped to that action, and a check on the final state outside the model, keep the write out.

The read path can still be a workflow or a retrieval app. Adding an agent so the model can be careful puts the carefulness in the prompt, which is the place a later instruction can talk to.

That bundle is the harness: the gate, the scoped credential, and the check on the record after the call. High blast radius without it is a reason to stay on the read path, not a reason to pick a more autonomous shape.

Our guide on when your AI product needs write access is the product boundary, where answering can stay at read and the first write is one action class. How to roll out AI agent autonomy in five levels is how that action moves up a level after the evidence for it exists. Both assume you have already decided an agent is the right shape for the path.

Deel's description of Akai, as reported by Progressive Robot on 28 September 2026, draws the same line for exact fields. A total, a tax rate, an account number, or a payment reference runs on rules. The model is reserved for a judgment a rule cannot state, and a write waits for a named person. We treat that as their published design for their own operations. The revenue and hours figures in the same coverage are theirs, and this guide does not use them.

The same constraint shows up once a company is already counting agents in the thousands. On 30 September 2026, in an interview Algorithmine updated on 5 October, an anonymized platform lead at a Fortune 500 services company described a fleet of more than 10,000 governed units.

The lead counts a unit that has an owner, a declared tool list, a model policy, a service level objective, and a version. A unit with no named owner cannot be promoted, and the interview says the figures are self-reported and were not audited. What we take from it is the order. The owner and the gate come before the count, so a fixed sequence stays a workflow beside that fleet.

What do you score once the shape is chosen?

Success is a state something other than the model can check, or you cannot tell the three shapes apart after launch. A workflow scores the output of each step: the label, the drafted field, the branch that fired. Retrieval scores whether the right passage was in the set and whether the answer stays inside it.

An agent scores the trace and the final record. Count the tools, the recovery, the approval, and whether the record matches. Evals are how that set stays stable from run to run. Until the cases exist, a demo that felt like an agent is not evidence that you needed one.

If the only signal is the model's closing sentence, every shape will report done. The workflow reports done when a branch was skipped, and retrieval reports done when the passage was missing. The agent reports done after a tool error it folded into the prose. Name the check before you pick the shape. Then build the shape that can produce that check with less machinery, and leave the others until a real task forces them.

Observability on the loop is the trace of that check, not a second product. Store the prompt, the tool list, and the stop reason with the run, so a wrong final record can be tied to the choice that produced it. A workflow needs less of that trace, because the branch is in the code. Spend the trace on the shape that chooses its next tool, and keep the workflow's record as the branch id in the code.

Common questions

Do we need an agent, or is a workflow enough?#

Build a workflow when a person can write the steps and the branches, including retries that always repeat the same call. Build an agent when the next tool depends on a result you could not list before the run. At Bigcircle the model stays inside the step. A loop that rediscovers the order costs more to trace and fails on a longer path.

When is a RAG or search app enough on its own?#

A RAG or search app is enough when one query can find the passage and the answer has to stay inside that passage. Fix the index, the identifier match, the reranker, and the permission filter on the query before you add a loop. Add agentic search only when the first result forces a second query you can name, and give that loop a stop condition and a check on the cited span.

Can one product use a workflow, retrieval, and an agent?#

One product can use all three, and the split is per task. Triage can be a fixed workflow, the policy answer can be retrieval, and the ticket that has to pick the next system from the last lookup can be an agent. A single agent label on the product hides which of those three you shipped, and it sends the next failure to the wrong owner.

Should a high-risk write start as an agent?#

A high-risk write should not start as an agent. The call can move money, access, a customer message, or production state, and it needs a gate before any loop may request it. If those controls are missing, keep the write out and ship the read path as a workflow or a retrieval app. Authority for one action class is a later release, after the path has earned the loop.

How do you know the run succeeded?#

Check a state outside the model's closing sentence, so a workflow shows the branch that fired and the field it produced. Retrieval should show the passage and whether the answer stayed inside it. An agent should show the tools, any recovery, any approval, and a final record that matches the result you expected. If you cannot name that check, you cannot yet tell which shape you need.

Does a large fleet of agents change the choice?#

The choice stays with the task. A fleet still needs an owner on each unit, a declared tool list, and a version you can roll back. It does not turn a fixed sequence into an agent. Count the units you can name and govern, and leave a scripted workflow as a workflow even when it sits beside agents.

Further reading