Should you embed an AI engineer, hire full-time, or use an agency? Decide on three things: how fast you need production depth inside the repository, who will still be able to change that code after launch, and how large a hole a bad match leaves in the roadmap. Day rate does not answer any of those, and a lower rate on a handoff you cannot operate is the expensive choice.
An embedded engineer wins when you already have direction and a team, and the missing piece is someone who has shipped retrieval, agents, and evals into a live codebase. A full-time hire wins when AI stays on the critical path for years and a manager on your staff can review that person.
An agency wins when the output is a bounded system and you can score it. It fails when you need ongoing ownership inside your systems, because the people who learned the failure modes leave with the statement of work.
| Embed an engineer | Full-time hire | Agency | |
|---|---|---|---|
| What you are buying | A production-proven engineer in your repo. Employment sits with us | A permanent member of your staff | A bounded deliverable |
| Speed | Shortlist in two to five business days, in the repo inside two weeks | The months-long sourcing cycle our hire page is written against, then a ramp on your stack | A kickoff can be quick. Context inside your systems is the slow part |
| Who owns the code | You do, under your IP terms, while the work is in your repositories | You do, including the employment | Whatever the contract assigns. A repo assignment is not the same as the judgement |
| A bad match | Interview, optional paid trial, then a replacement at no extra cost | The seat, for as long as the exit takes, and then another search | A system your team did not learn how to run |
| Wins when | You have direction and a team, and you need AI depth now | AI is a multi-year competency and you can manage it | The outcome is bounded and the acceptance tests are already written |
| Fails when | Nobody can set priorities week to week | You need depth this quarter, or nobody can review the work | You need the same people still in the system after acceptance |
The timings, the IP line, and the replacement policy in that table are the ones published on our hire page. They describe our embed, not every staffing firm and not a survey of agencies. The agency column is the structural contrast, and this site does not publish agency rates or agency success rates.
Where should each option sit?
Place the three options on the two bets that survive a rate card, before you argue about the third. One bet is speed to production depth inside your repository. The other is whether the judgement that built the system is still available after the work. A permanent seat and a fast start are different goods, and most writeups treat them as one.
This figure is a framework, and the dots are not a dataset. Height is how fast production depth shows up in your repository. The horizontal position is how long that operating knowledge stays available to you. The embed sits high and right of center because our published path has someone in the repo inside two weeks, while the seat itself is not a permanent headcount bet.
The agency sits to the left because acceptance is the end of the people's involvement, even when the kickoff was quick. The full-time hire sits low and far right because the search is the slow part and the seat, once filled, stays.
Read the empty corner as a warning rather than as a fourth product. You end up in the slow and gone corner when an agency also takes months to learn your system and then leaves.
Fast and permanent is the combination teams ask for on a slide, and it is not a staffing model. You can approach it by embedding while a full-time search runs, which keeps the quarter from going empty, and that sequence is a plan rather than a point on the chart.
What does a bad match actually cost?
The chart above has no place for the loss you still carry after you notice the miss, and that loss is the third bet. Teams compare salaries because the salary is a number and the hole in the roadmap is a story. That hole slips a launch. A bad embedded match, a bad hire, and a bad agency engagement do not leave the same hole, so a single "cost of a mistake" line on a spreadsheet is already the wrong model.
This figure is a framework. It orders the three misses by who still carries them, and it does not contain a failure rate, a salary multiple, or an industry average. On an embed with us, the published remedy is concrete: you interview, you can run an optional paid trial on a real task from your backlog, and if the match is still wrong we replace the engineer at no extra cost.
The hole is a sprint, and the employment risk sits with us. That is the policy on the hire page, including the line that there is no renegotiation and no argument about fault.
An agency miss is deferred, which is why it feels safe in the statement of work and expensive later. If you wrote acceptance tests, you can reject a deliverable that fails them, and that is a real control. What you cannot reject your way back to is the judgement. The people who learned which documents break retrieval, which hallucination your users actually hit, and which tool call must never fire, leave when the work is accepted.
Your team inherits a repository. The next model upgrade falls to whoever is still employed, and if that person was never in the reviews, the system gets worse quietly.
A full-time miss sits on your books. You own the employment relationship, so the seat stays occupied while notice, a slow exit, or a second search runs, and the roadmap waits on that seat. We will not invent a duration or a probability for that. The structural point is enough, and it is the correct downside when the role is truly core. It is a bad downside when the role was a workaround for one system you needed this quarter.
When does an embedded engineer win?
Speed to production depth favors an embed when someone on your side can already direct the work. You have a roadmap and leads who can say which system matters this month.
What you do not have is a person who has taken retrieval, an agent, or a fine-tune through the failures that only show up on real documents and real users. Opening a req for that person means running a search your team has not run before, then reviewing a specialty your existing leads may not be able to score.
On our side the mechanics are published, not negotiated in a footnote. You see a shortlist within two to five business days, because the match is drawn from engineers who have already cleared vetting. Most clients have someone embedded inside two weeks. The engineer is full-time in your timezone window, in your repositories, your standups, and your sprint cadence, and is not split across other clients. You set priorities.
You own the code and the IP. We handle employment, payroll, and compliance, so you are not opening an entity or defending a contractor classification.
The work they are there to do is the part under the demo. A production system needs evals that can fail a release, guardrails on what may leave the building, observability when a run goes wrong, and tool calling scoped so the model cannot do more than the task allows.
If a proposal cannot point at those, the staffing model is beside the point and you are buying a demo. Our production AI engineering page is the list of that layer, and the generative AI engineer role is the person who ships a feature through it.
Embed fails when there is no direction to embed into. An engineer in your repo still needs someone to say which failure matters this week. If the AI roadmap is a list of ideas and nobody owns the priority, you will get activity and not a production system.
In that case the missing role is leadership of the build. That is a different engagement, covered in embed vs CTO-led pod vs fixed-scope, and adding a second engineer will not create the missing lead.
When does a full-time hire win?
A permanent seat is the right instrument when the work will still be yours in three years and someone on your staff can manage it. AI as a feature you ship once is a weak reason to open a req.
AI as the way the product works, with a backlog that keeps producing new failure modes, is a reason to have the skill on payroll. The person learns your domain in a way a rotating contractor will not, sits in planning, and can hire the second person when the team grows.
You need the management capacity before you need the req. A staff engineer who has never reviewed a retrieval eval cannot tell whether the new hire is improving the system or polishing a demo.
The practical version of that review is in how to build LLM evals: a failing score has to point at a cause, or the manager is approving anecdotes. If you cannot spare that review, the hire drifts, and a customer is often the one who discovers it.
The blast radius does not appear on the offer letter. You own the employment. A miss costs the seat until you can exit it, and then you search again while the roadmap waits. There is no published number on this site for how often that happens, and we will not invent one. Carry that downside when the role is core. Do not carry it to fill a one-quarter gap.
Full-time also loses on speed when the system has to be in production this quarter. The hire page describes our embed as the way out of a months-long sourcing cycle.
If your date sits inside that cycle, the req is the wrong instrument even when the long-term answer is still a hire. A path we use with teams is to embed first, learn which skills the system actually needs, and open the permanent req against that evidence rather than against a job description written in a workshop.
When does an agency win?
An agency fits a system you are allowed to finish. Discovery, build, evals, and deployment against a written acceptance test is a coherent purchase. "Make our AI better" is not a purchase, because nobody can tell when it has been delivered, and the engagement slides into open-ended staff work without an owner in your repo. Write the tests in the currency of the system before the kickoff, or you will be paying someone to invent the definition of done.
For a retrieval product, the test is whether the right passage comes back and whether a citation resolves, which is the failure catalogue in why RAG systems fail in production. For an agent, the test is a termination condition and a check that the work happened, which is the subject of why AI agents report unfinished work as done.
For anything a customer will see, the test includes what the system must refuse. If you cannot write those tests, you are not ready to buy a bounded deliverable, and the more honest next step is to spend a week on the test rather than on a vendor selection.
The failure mode is ownership after acceptance. An agency team learns your edge cases and your data quirks, and then the statement of work ends. Use an agency when the system is allowed to be finished: a migration, a one-off extraction pipeline, a proof with a kill date. Do not use one when the system is a product you will still be changing next year. That work needs a named owner inside the repo.
An embed supplies that owner while you keep the management. A full-time hire supplies that owner permanently. An agency supplies a handoff.
One warning before you treat "agency" and "fixed-scope build" as the same row. A fixed-scope engagement with us is still our engineers, under the IP terms on the hire page, with discovery, build, evals, and deployment under one point of accountability. A general agency is a different market, and we are not publishing a comparison of agency brands.
The question in this section is the shape of the buy: a finished system versus ongoing ownership. Which of our own models matches that shape is the subject of the engagement guide.
How does the rubric score them?
Read the bars as the words High, Med, and Low, assigned by rules this page states, and not as measurements of companies. The same three sections you just read are the rules. Speed to depth is High when a vetted person is in your repository inside the two weeks our hire page publishes, Med when work can start without a req but the people who learn the system are not staying, and Low when you are still inside a months-long search.
Code ownership is High when pull requests are in your repositories under your IP terms, and Med for a general agency because the assignment depends on a contract this site does not publish. None of the three is scored Low on code ownership, and that is deliberate: we are not claiming agencies take the IP, only that you have to read the clause.
Judgement after launch is High when the person is your employee, Med while an embed continues and the seat is not permanent, and Low when the statement of work ends and the people leave. Small blast radius is High when the published remedy is a trial and a free replacement, Med when written tests can reject a deliverable and the judgement still leaves, and Low when you own the employment.
Multi-year competency is High for the permanent hire, Med for an embed used as a bridge while you learn the real skill mix, and Low for an agency, because the competency does not stay.
This figure is a framework, and the bar length repeats the word beside it. A longer bar is not a measured advantage and not a percentage of teams. If you disagree with a cell, disagree with the rule in the paragraph above and change the word. Do not read the picture as data we collected.
We did not collect it. The picture exists so a buyer can see, in one pass, that the option which wins on speed is not the option which wins on a permanent competency, and that no column wins every row.
What does production depth refer to?
Shipped systems make the phrase concrete, because the hard part is still in the repository after the demo. These case studies are not a controlled comparison of staffing models. They are evidence of the work the phrase "production depth" points at when we use it. If a proposal cannot talk about evals, citations, and an owner for the week after launch, you are not yet choosing between embed, hire, and agency.
Our trademark drafting platform has been in production since 2023. It runs a fine-tune, citation-grounded retrieval, and schema-constrained decoding in the same request, and a citation from outside the firm's approved index is not possible. That constraint is architectural.
Whoever owns the system has to keep the authority index, the eval harness, and the decoding schema aligned when the base model changes. A handoff that stops at "the demo drafts" does not include that upkeep, and the attorney still has to trust the draft on Monday.
Prospect intelligence is our own pipeline, which is why the numbers are ours to publish. One observed run cut 420 collected sources to 6 pieces of grounded evidence, and cache-first execution cut external API calls by an estimated 50 to 60 percent. The ongoing work is freshness and evidence locks, not the first prompt. An outside team can build the first pipeline. The question is who notices when a cached fact has gone stale before a sales call.
ESG filing intelligence compressed an analysis cycle from weeks to hours, with lineage back to the page on every metric. That cycle-time change is an internal working result, not yet reconfirmed on current volumes, and the case study says so. The structural lesson is the one to copy into an acceptance test: a number you cannot source is a number you cannot use. Lineage belongs in the definition of done, not in a screenshot of a dashboard.
How does ownership move over time?
The shapes below are a model of who can still change the system, not a measured ramp and not a forecast. The vertical position has no unit. Higher means more of the operating knowledge is available to you. We drew three shapes so the handoff problem is visible as a drop, and so the full-time search is visible as a long flat stretch before anyone is in the repo.
This figure is an illustrative model. One mark on it is a published duration, and the caption inside the picture says which: our embed is in the repository inside two weeks, after a shortlist in two to five business days. The full-time line stays low across the opening because the hire page contrasts that path with a months-long sourcing cycle, and we have not published a month count for a typical search.
The agency line rises while the work runs and drops at acceptance. We have not published agency durations, so that drop is a shape, not a calendar.
The embed line also drops after the engagement, on purpose. You still own the code, which is the IP term on the hire page, and the person may leave when the engagement ends. Those are different facts. Teams conflate them and then feel surprised that a repository is not the same thing as a teammate.
If the system will keep changing, plan the next owner before the drop: a full-time hire you have already started sourcing, or an extension of the embed. If the system is allowed to be finished, the drop is the point of the agency buy, and you should have the acceptance tests in hand before it happens.
How do you decide, in order?
Stop at the first honest yes.
- Can you write what done means, as a test your team can rerun? If not, do not sign an agency scope yet. Spend the time on the test. The derisking guide covers why pilots die when nobody can say what a correct output looks like.
- Will this system still be changing inside your product a year from now? If it will not, and the test from the first question exists, an agency or a fixed-scope build is the fit.
- Do you have a lead who can set weekly priorities and review the work? If you do, and you need depth in the repo now, embed. If you do, and AI is the competency you intend to keep, open the full-time req, and consider embedding while that search runs so the quarter is not empty.
- If you do not have that lead, an embedded engineer will not create one. You need someone to run the build. That is the pod question in the next guide, not a reason to hire a senior generalist and hope.
- Is the downside of a bad permanent hire acceptable this year? If losing the seat for the length of an exit would stall a launch, do not open the req yet. Buy the smaller blast radius, learn the real job, and hire against what you learned.
The order matters because teams start at step 5, with a compensation band, and never write the test in step 1. The test decides whether the agency column is available. The lead in step 3 decides whether an embed can work. Without those two, the rate comparison is a conversation about a purchase you have not specified.
Questions buyers ask
Should we embed an AI engineer or hire full-time?#
Embed when you need production depth in the repository now and you do not want a permanent headcount bet yet. Hire full-time when AI will remain a core competency for years and a manager on your team can review the work.
The two are often a sequence: embed while you learn the real skill mix, then open the req against that evidence. They are a poor substitute for each other when the permanent hire is filling a one-quarter gap, or when the embed is a way to avoid ever building the skill you already know you will need.
When is an agency the better buy?#
When the deliverable is bounded and you can write acceptance tests before kickoff. A migration, a one-off pipeline, and a system with a kill date all qualify.
The agency is the wrong buy when you need the same people still owning the system after the statement of work, inside your permissions, your on-call, and your model upgrades. If you cannot write the tests, you are not ready for this column, and signing the scope anyway means the vendor will define done while billing for the definition.
Who owns the code and the IP?#
On an embed with us, you do. Engineers work in your repositories under your IP and confidentiality terms, and we handle employment on the back end. That sentence is the hire FAQ, not a slide. On a full-time hire, you own the employment and the code.
On an agency engagement, read the contract: ownership of the repository is not the same as ownership of the operational knowledge, and only the first is usually assigned in writing. This site does not publish a standard agency IP clause, so the FAQ answer stops where our own engagements stop.
How fast can an embedded engineer start?#
For our embed, most clients see a shortlist within two to five business days and have someone in the repo inside two weeks. There is a 30 minute call on day zero, your own interview or a paid trial in the first week, and onboarding after that. The speed is possible because the pool is already vetted, and historically fewer than three in a hundred applicants make it through.
A full-time search is the months-long cycle that page is written to avoid. An agency kickoff can be faster than a hire and still leave you, in month four, with nobody who can change the system.
What is the blast radius of a bad hire?#
We will not invent a salary multiple or a failure rate. The costs that are structural are these. A full-time miss occupies a headcount slot until you can exit it, and then you search again. An embed miss, on our terms, is an interview, an optional paid trial, and a replacement at no extra cost if it is still wrong.
An agency miss is a deliverable your team cannot operate, which is a cost you pay after the invoice, when the next change has no owner. Compare those three holes, not the day rates, when you decide which risk you are willing to hold.
What does production-proven mean on this site?#
Every engineer we present has shipped to production. The public description of the filter is a five-stage vetting: an applied generative-AI build, a live system-design review, a communication screen, and a frontier-knowledge check, on top of resume screening. Historically fewer than three in a hundred applicants make it through, and the hire page states a bench behind that filter of more than 200 production systems.
It is a filter on people. It is not a guarantee about your project, and a vetted engineer still needs a lead, a repository, and a definition of done.
Where to start this week
Write three lines before you ask anyone for a rate. Name the system you need in production, name the person who will still be allowed to change it after launch, and name the test that decides whether it shipped. If the second line is blank, you are not choosing between rates.
You are choosing whether to rent judgement for a while, employ it, or accept a handoff. Bring those three lines to a conversation if you want them mapped onto an embed, or read the engagement guide if the open question is which of our three models fits a team that has already decided to work with us.
Further reading
- Hire AI engineers. Timelines, ownership, replacement, and the vetting behind the filter.
- Embed vs CTO-led pod vs fixed-scope. Once you have decided to work with us, which engagement fits.
- Why enterprise AI pilots fail. Build, buy, or partner, one level up from staffing.
- How to build LLM evals. The acceptance tests an agency scope and a manager's review both depend on.
- RAG vs fine-tuning vs agentic search. The technical choice the engineer, the hire, or the agency will still have to make.
- Production AI engineering. The layer under the demo.
- Case studies: trademark drafting, prospect intelligence, ESG filings.