Skip to main content
Content Hub · AI Engineering

AI Engineering

The full professional track for building AI products for production — from a first prompt to a stable, safe and measurable system. Organized into 3 levels: beginner, advanced and expert.

AI Engineering is the field of building real products on top of language models — not just "talking to ChatGPT," but designing, building, evaluating and maintaining AI systems that run in production for real users. It's a blend of software engineering, working with LLMs, and infrastructure. This track takes you step by step — pick your level and start.

What this discipline actually is

AI engineering is what sits between a model that works in a notebook and a system people rely on. Very little of it is about the model. The model is a component you mostly do not control — it is hosted somewhere else, it changes under you, and its behaviour is statistical. The engineering is everything you build around that component so the unpredictability stays contained.

In practice that means: deciding what context reaches the model and what does not, validating what comes back before anything downstream trusts it, measuring quality with something better than reading a few outputs and nodding, and knowing what happens when the provider is slow, expensive, or returns something you did not anticipate.

The order things usually go wrong

Projects tend to fail in a recognisable sequence, and the sequence is worth knowing because it tells you which problem you are allowed to skip for now.

  • First: the demo works and nothing else does. The prompt that produced a beautiful answer on the example you picked produces something unusable on the third real input. This is almost always a context problem rather than a model problem.
  • Then: nobody can tell whether a change helped. Without evaluation you are tuning by vibe, and every improvement is also a regression somewhere you did not check. This is the point at which most teams should stop adding features.
  • Then: the output is right but the shape is wrong. Free text cannot be consumed by code. Structured outputs, schema validation and a defined behaviour for "the model returned something invalid" are what make a model callable from a program.
  • Finally: it works and the bill arrives. Cost and latency become the constraint once correctness is handled — and they are much easier to fix when the evaluation harness already exists to prove the cheaper version is still good enough.

Evaluation is the load-bearing part

It is also the part that gets skipped, because it feels like overhead before it feels like leverage. The argument for doing it early is simple: every other decision in the list above — a cheaper model, a shorter prompt, retrieval instead of a stuffed context, a smaller cache — is a trade of quality against cost that you cannot evaluate without a way to measure quality.

It does not have to be elaborate to be useful. Thirty real inputs with known-good outputs, run on every change, catches more than an elegant framework that nobody maintains. The bar is being able to answer "did that change make it better or worse" with something other than an opinion. More in evaluating LLM systems.

Retrieval before fine-tuning, almost always

The common instinct when a model does not know something is to train it on your data. That is usually the expensive answer to the wrong question. Fine-tuning adjusts behaviour — tone, format, adherence to a convention. It is a poor way to install knowledge, because knowledge changes and a fine-tuned model is a frozen artefact you have to rebuild.

Retrieval solves the knowledge problem directly: keep the facts in a store you can update, fetch the relevant part at call time, and the answer changes the moment the source does. Start there, confirm the remaining failures are about style rather than substance, and only then consider fine-tuning. See RAG and fine-tuning.

The three tiers above, and who they are for

  • Beginner — prompting, context, structured output, tool use. This is the whole job for a large number of real systems, and it is where the most value per hour of reading sits.
  • Advanced — retrieval, embeddings, reranking, memory, frameworks. Read this when you have a working system and specific failures you can name.
  • Expert — evaluation, observability, guardrails, cost, injection defence. The tier that decides whether the thing survives contact with real users, and the one most likely to be reached too late.

The tiers are a reading order, not a ranking of difficulty. Plenty of engineers who are comfortable with vector databases have never built an evaluation set, which is roughly the equivalent of shipping a service with no tests and a very good cache.

Starting from scratch?

Start with the "What is AI Engineering" guide — it explains the full picture and the right learning order.