Nine layers, one production-proven path.

From the foundations under a model to fine-tuning, serving, and running it at scale. Pick a layer and read it whole: its competencies, the concepts behind them, the technologies that implement them, and the production practice that separates a shipped system from a demo. 252 items, each with a note, and most with a written guide and its sources.

of applicants make it through
Top 3%
have shipped to production
100%
production AI systems shipped
200+
from foundations to scale
9 layers

At a glance

Plan a learning path from AI foundations to production systems. Open a stage to review its concepts, technologies, and engineering practices.

Who this is for
Engineers developing AI skills and leads planning team development.
Topics
  • Foundations
  • Model training
  • Inference
  • Production systems

Start here

Stage 01

Foundations

The engineering bedrock under every AI system. Before the model, the fundamentals that keep it honest in production.

Programming and asyncData handlingML fundamentalsEngineering hygiene

Read stage 01

Stage 02

LLM applications

Turning a raw model into a dependable product surface, with output you can actually build on.

Model APIsPrompt engineeringStructured outputTool and function calling

Read stage 02

Grounding and retrieval

Stage 03

Retrieval and RAG

Grounding answers in real data, with citations, so the system says what is true rather than what is plausible.

Embeddings and chunkingVector storesRetrieval qualityRAG patterns

Read stage 03

Stage 04

AI agents

Systems that plan and act across real tools, check their own work, and recover when a step goes wrong.

Planning and controlTool useMemory and stateOrchestration

Read stage 04

Beyond text

Stage 05

Multimodal and voice

Vision, generation, and real-time voice, built to stay fast and grounded on a live call.

Vision and documentsImage generationSpeechRealtime voice

Read stage 05

Measure and improve

Stage 06

Evaluation and error analysis

The core production discipline. Not generic metrics, but looking at your own data, naming failures, and measuring the things that actually break.

Error analysisCustom evaluatorsLLM-as-judgeOffline CI and online monitoringTracing

Read stage 06

Make the model yours

Stage 07

Fine-tuning and adaptation

When prompting and retrieval run out, reshape the model itself. Done right, a small fine-tune can match a frontier model on your task at a fraction of the cost.

When to adaptSupervised fine-tuningPreference optimizationDistillation

Read stage 07

The systems layer

Stage 08

Serving and inference

What it takes to run a model yourself, fast and affordably. The two-phase nature of inference governs your real cost and tail latency, and it is where most teams have no depth at all.

Self-host vs APIInference enginesThroughput and latencyCompression and scale

Read stage 08

Ship and operate

Stage 09

Production and LLMOps

Running it for real, inside a latency and cost budget, observable, defended against hostile input, and reliable enough that customers feel it.

DeploymentObservability and LLMOpsGuardrails and securityCost and FinOpsReliability

Read stage 09

Nine layers deep, and every one of them already shipped.

Put this whole roadmap on your team.

Every layer above is someone you can hire, production-proven and embedded in your team in days. Tell us what you are building and we will line up a shortlist.