AI Agents
The central guide — from the basic concepts to building an agent that analyzes, decides and acts on its own: fundamentals, RAG, frameworks, MCP and security. Start here.
An AI agent isn't just a chat — it's a system that takes a goal, plans steps, uses tools (search, API, code), reads from your knowledge sources (RAG), and performs actions until the result is achieved. This page gathers everything you need to understand, build and secure a real AI agent — step by step.
AI agent fundamentals
What an agent is, how it plans and acts, and how to connect tools and capabilities.
Knowledge & context (RAG)
How to ground an agent in your own information, so its answers come from your documents rather than from the model’s own recall.
Frameworks & building
The tools and libraries you actually build agents with — and ready-to-run templates.
Security & reliability
How to keep the agent safe, predictable and measurable — before you ship to production.
What makes something an agent
The word is used for everything from a chatbot with a nicer name to a system that runs unattended for hours, which makes it nearly useless as a category. The distinction worth keeping is narrow: an agent decides its own next step. A workflow has its steps written down in advance; an agent is given a goal and a set of tools, and chooses which tool to call, in what order, and when it is finished.
That single property is where all the power and all the difficulty come from. A fixed workflow either works or fails visibly. An agent can take a path nobody anticipated, succeed at something adjacent to what you asked, or loop quietly until it runs out of budget — and each run can differ from the last with identical inputs.
When you do not need one
Most tasks described as needing an agent need a workflow with one model call in it. If you can draw the steps on paper and they do not change, write them down and run them; that system will be cheaper, faster, easier to debug, and will behave the same way tomorrow.
Reach for an agent when the number of steps genuinely depends on what is found along the way — researching a question where each answer determines the next search, triaging an issue where the diagnosis decides which tool to reach for, working through a codebase where the shape of the problem is not known until you look. If you can enumerate the branches in advance, they are branches, not autonomy.
The three parts that actually decide quality
- The tools. An agent is only as capable as what it can call, and tool descriptions matter more than most people expect — they are the documentation the model reads before deciding. A vague description produces a tool used at the wrong moment with the wrong arguments. Narrow, well-named tools with explicit parameters beat one flexible tool that does everything.
- The context. Every step appends to the conversation, so by step twenty the model is reasoning over a transcript largely made of its own earlier output. What you keep, what you summarise and what you drop between steps is a design decision, not an implementation detail — and it is usually the difference between an agent that stays coherent and one that drifts.
- The stopping condition. Knowing when to stop is a genuine problem, not an afterthought. Agents declare success early, retry the same failing call, or keep going long past the point of usefulness. A hard ceiling on steps, on wall-clock time and on spend is the minimum, and it belongs in the first version rather than the version you write after the first surprise.
Cost behaves differently than people expect
The instinct is to multiply: ten steps, so roughly ten times a single call. The real figure is higher, because each step carries the whole accumulated history. Step one sends a small context; step ten sends everything that came before it. The bill is the sum of the contexts, not the count of the steps, and it grows faster than linearly with the length of the run.
The practical consequences: a run that goes off course is expensive as well as wrong, caching what is stable across steps is worth real money, and an agent left without a spend ceiling will eventually find a loop and stay in it. See cutting LLM costs for the arithmetic.
Security is not a later chapter
The moment an agent reads something it did not write — a web page, an email, a file, an API response — that content is inside its context, and the model does not natively distinguish your instructions from text it encountered along the way. This is the whole of prompt injection, and it is a design problem rather than a bug with a patch.
The defences that hold up are structural: give the agent the narrowest permissions the task needs rather than the ones that are convenient, keep anything irreversible behind a human approval, treat retrieved content as untrusted input, and log what the agent did so that a bad run can be reconstructed afterwards. The agent security guide covers this properly.
Reading the clusters above
- Fundamentals — start here, including if you have already built something. Most agent problems come from skipping the part about what the loop is actually doing.
- Knowledge and context — how an agent gets at information it was not given. Relevant the moment your agent needs to know anything specific to your organisation.
- Frameworks and building — worth reading after you have written a simple loop by hand. Frameworks are much easier to judge once you know what they are hiding.
- Security and reliability — the cluster people read last and should read second, particularly before an agent touches anything it can break.
The guides here are written around what fails in production rather than around demos. A demo agent needs to work once; a production agent needs to fail safely on the run you did not anticipate, which is a different engineering problem and a much less photogenic one.
Want to build an AI agent for your business?
We'll take you from idea to an agent running in production — with RAG, tools and security done right.