Hire prompt and eval engineers who make quality measurable.
Specialists in the discipline most teams skip: designing prompts that hold up, and building the evals that turn "it feels better" into a number you can ship against.
At a glance
The people who turn AI quality into a number you can ship against.
Review the work Prompt & Evaluation Engineers can own, the skills required, and the systems they can help you build.
- Who this is for
- Engineering leaders who need an engineer to work within their existing team.
- Topics
- LLM-as-judge
- Custom metrics
- Golden sets
- Versioning
- Templating
- Structured outputs
What they own.
- Design and version prompts that stay reliable in production
- Build evaluation suites that catch regressions automatically
- Set up LLM-as-judge and human review workflows
- Turn quality from a gut feeling into a tracked metric
- Red-team prompts for safety and edge cases
- Tune for accuracy, tone, and format together
What you can ship with them.
Tools they reach for in production.
- Eval
- LLM-as-judgeCustom metricsGolden sets
- Prompt
- VersioningTemplatingStructured outputs
- Tools
- LangfusePromptfooTracingRagasLangSmithBraintrust
- Stack
- PythonClaude / GPT
Seniority: Engineers who treat evals as core engineering, not an afterthought.
Five stages.
The top 3% remain.
Every stage asks the same question: can they keep AI running once real customers are using it? Getting something started is the easy part, and it is not what we screen for.
400 applicants
3%of applicants reach
your shortlist
What removes them
We start with something they built
100% → 20%One real system, pressed hard. How much traffic did it take? What broke first? Who got the call when it did?
We break something and watch them fix it
20% → 9%A working system with a bug hidden inside it. Anyone can build a demo in a weekend. Fixing code you have never seen is the actual job.
How will you know it is working?
9% → 5%Before they write anything, they have to tell us how they would test it, and what they would do when it gets an answer wrong.
Make it fast without running up the bill
5% → 4%We give them a speed target and a budget, then ask them to explain the tradeoffs they made to hit both.
It is late and the AI got it wrong
4% → 3%What do you do first? How do you find out what happened, undo it, and explain it to the customer in plain words?
Often hired together.
LLM Engineers
Model specialists who fine-tune, quantise, and serve the model itself.
Generative AI Engineers
All-rounders who ship a whole AI feature end to end, or anchor a pod of the specialists below.
AI Agent Engineers
Builders of agents that plan and take real actions in your systems, safely.
Find the people to accelerate your roadmap.
You don’t need more resumes. You need proven AI engineers embedded in your workflow and ready to build from day one. Tell us what’s missing and we’ll line up a shortlist.