All services

Stop paying frontier prices for routine work.

Most systems send every request to one model because choosing is work. We measure each class of request on cost, latency and quality, then route it to the smallest model that still passes, with fallbacks when a provider degrades and caching where the same question keeps arriving. The routing table is configuration you own, so when a lower-cost model gets good enough you change a value rather than rebuild.

Built for production, not the demo.

01 / ROUTING

Routing across models by task, for teams paying frontier prices for routine work.

For teams paying frontier prices for routine work. Each request class measured on cost, latency and quality, then routed to the smallest model that holds.

Usually shipped with

  • Production AI engineering
  • AI observability
  • Self-hosted model deployment

Not a bundle to buy. Whichever you start from, the engagement covers what the build actually needs.

02 / Scope

What we build.

  • Request classes measured on cost, latency and quality before routing
  • Routing across frontier, fast and open models by class
  • Fallback paths when a provider degrades or rate-limits
  • Prompt and response caching where traffic repeats
  • A routing table you own and can re-measure as providers change
  • Built on LiteLLM, OpenRouter or a gateway you already run, rather than a wrapper only we understand

03 / Outcomes

What you can ship.

  • Inference cost reduced without a quality drop
  • Latency held under load through fallbacks
  • Provider changes absorbed by configuration, not a rebuild

04 / Deliverables

Artefacts, not activities.

  • A measured routing tablePer-class cost, latency and quality for each candidate model, with the chosen route and the reasoning.
  • The measurement harnessRe-runnable when a new model lands, so the decision stays current.
  • Fallback and caching layerProvider degradation, rate limits and repeat traffic, handled before they reach a user.

05 / Stack

What it is built on.

Models
Frontier / Fast tiers / Open models
Routing
Class-based routing / Fallback chains / LiteLLM / OpenRouter / Portkey
Caching
Prompt caching / Response caching
Measurement
Cost/latency/quality harness

06 / Why us

The routing table is configuration, not code

When a lower-cost model gets good enough you change a value and re-run the harness. Teams that hard-code the choice rebuild instead.

Quality is measured per class, not overall

An average quality score hides the class where the small model fails. Routing decisions made on averages are how cost savings turn into complaints.

We re-measure as providers move

Model quality and price change monthly. A routing decision made once is stale within a quarter, so the harness ships with the table.

A path from your problem to production.

  1. Week 1

    Classify the traffic

    Requests are not uniform. We separate them into classes by what they actually demand, because routing an easy class to a frontier model is where the money goes.

  2. Week 1-2

    Measure each class three ways

    Cost, latency and quality per class, per candidate model. The measurement is the work; the routing table falls out of it.

  3. Week 2-3

    Add fallbacks and caching

    A provider degrades, a rate limit hits, the same question arrives twice. Each has a less expensive answer than a retry against the same model.

  4. Week 3-5

    Hand over the table

    The routing configuration is yours, with the measurement harness, so you can re-run it when a new model lands.

One model for every request is the expensive default.

Each request is classified by task, length and whether it carries personal data, and a router picks a fast, frontier or local model against cost, latency and failover rules. Responses merge through a schema check with retry before they reach your app, and every call writes what it cost and how long it took. Routing is only worth doing when you can prove which lane a request took.

YOUR REQUESTIDENTITY & QUOTAwho · how much leftCLASSIFYtask · length · PIICARRIES PII?CACHEanswered before a modelROUTERcost · latency · failoverFASTFRONTIERLOCALMERGE & CHECKschema · retryYOUR APPlocal lane onlyon error → routertimeout gives upCOST PER CALLfast · frontier · localWHICH LANE, AND WHAT IT COST
Intersecting translucent blue planes suggesting alternate paths

Production-proven

Built by engineers who've already shipped this in production.

The questions buyers actually ask.

How much does routing usually save?

It depends entirely on your traffic mix, and anyone quoting a percentage before measuring is selling. The harness produces your number in the first fortnight.

Does quality drop?

Not if quality is measured per class, which is the whole method. A class routes to a smaller model only when it passes the same bar.

Are we locked to your routing layer?

No. It is configuration in your codebase, and the harness is yours. The point is that you can change providers without us.

What about prompt caching?

Included where traffic repeats. It is usually the largest single saving and the least disruptive, so we look there first.

Let's scope your model routing and cost control build.

Tell us where you are and what you're trying to ship. We'll come back with a concrete plan, the right engineers, and a path to production, not a generic pitch.