AI Guides › Playbooks

The Five-Rung AI Engineering Ladder

By Nigel Guy · 7 min read

Most people who want to become AI engineers start at the top. They install an agent framework, copy a demo that "books flights", and discover three weeks later that they cannot say why it fails on Tuesdays. It feels like progress because something runs. It is failing you because every layer underneath is unlearned, so every bug is a mystery.

The rule: climb one rung at a time, and do not step up until you can show a small piece of working evidence that you have finished the rung below.

Nothing here is a guarantee of a job or a timeline. The five rungs below are a sensible order of work, and the "roughly five months" figure is a planning estimate for someone studying part-time, not a measured outcome. Your pace will differ, and that is fine.

The Ladder

Rung You build Evidence you can step up
1. Groundwork Python scripts that read files, call web APIs and handle errors A small command-line tool in its own virtual environment, with a README, that someone else can run
2. Building blocks Calls to a model API with structured output, retrieval over your own documents, and a first eval A question-answering script over 20 or more of your own documents, with a written test set and a score
3. Agents A model that chooses between tools in a loop A workflow first, then an agent, with a log of every tool call and a hard step limit
4. Shipping The system behind a web endpoint, with config, cost limits, logging and tests Deployed somewhere you can reach it, with a budget cap and a way to roll back
5. Safety Permissions, input and output handling, human approval and an incident plan A written threat list checked against your build, with at least one test per item

Safety is rung five in the table because it is the last thing you finish, not the first thing you think about. Start the habits at rung one (no keys in code) and carry them up.

Rung 1 — The groundwork

Learn Python well enough to read someone else's script without fear. The official Python tutorial is written for people who already programme and covers the pieces you need: modules, errors and exceptions, file input and output, and a section on virtual environments and pip. Do those sections rather than a whole course.

Then learn four plain skills: calling a web API with an HTTP library and reading the JSON that comes back, storing secrets in environment variables rather than in the file, using Git so you can undo mistakes, and writing a test for a function. These are dull and they are the job.

Evidence to step up: a small tool that fetches something, transforms it, and fails politely when the network or the input is bad.

Rung 2 — The building blocks

Now add the model. Read the API docs of the provider you pick (Anthropic and OpenAI both publish theirs) and work through, in this order: a single request and response, a system prompt, structured output so the reply arrives as validated JSON rather than prose, and then the idea of tokens, context limits and cost per call. Claude's features overview lists structured outputs, citations, prompt caching and token counting as separate features, which tells you where the cost and reliability levers live.

Next is retrieval: splitting your own documents into chunks, finding the relevant ones for a question, and giving them to the model with an instruction to answer only from them. Ask for citations so you can check the source. A plain keyword search is a legitimate first retriever; add embeddings when the plain version demonstrably misses things.

Last on this rung, and the one most people skip: an eval. Anthropic's testing guidance says to write specific, measurable success criteria, build test cases that mirror real use including awkward inputs (empty, overlong, ambiguous), and prefer many cheaply graded cases over a few hand-graded ones. Start with 20 questions and the answers you expect. Rerun them whenever you change a prompt.

Evidence to step up: a score on your test set that you can say out loud, and a list of the questions it gets wrong.

Rung 3 — Agents, later than you think

An agent is a model that decides its own next step, calls tools, reads the results and continues until it decides it is done. Anthropic's "Building effective agents" draws the useful line: workflows follow code paths you wrote; agents direct their own process. Its advice is to start with the simplest thing and add complexity only when simpler solutions fall short, reserving agents for open-ended problems where you cannot hard-code the steps.

So build the task as a fixed workflow first. If it works, you are done and it is cheaper and easier to debug. If it fails because the steps genuinely cannot be known in advance, make the model choose between two or three tools.

Treat tool descriptions as interface design: say what each tool does, what the inputs look like, and give an example and the edge cases. Then add three controls from day one: a maximum number of steps, a full log of every tool call and result, and a rule that anything which changes the outside world (sending, deleting, paying) needs a human yes.

Evidence to step up: a log you can read afterwards to explain every decision the agent made.

Rung 4 — Shipping it for real

A notebook is not a service. Here you put the thing behind an endpoint and then deal with what changes when strangers use it:

Pin the model version where the provider lets you, and recheck the provider's deprecation notices, because models are retired on a schedule.

Rung 5 — Where safety fits

Safety is a set of ordinary engineering controls, not a separate department. Use the OWASP Top 10 for LLM Applications (2025 edition) as your checklist. The entries that bite builders first:

OWASP item What it means for your build
Prompt injection Text in a web page, email or document can carry instructions. Treat retrieved content as data, never as commands.
Sensitive information disclosure Do not put in the context what the user is not allowed to see.
Improper output handling Validate and sanitise model output before it reaches a database, a shell or a browser.
Excessive agency Give each tool the narrowest permission it needs. Prefer read-only.
Unbounded consumption Rate limits, token limits and step limits.

For a wider management frame, NIST's AI Risk Management Framework is voluntary guidance organised around Govern, Map, Measure and Manage, with a Generative AI Profile published in July 2024. NIST says the framework is currently under revision, so check the current version if you cite it.

Worked example (hypothetical)

Suppose you want an assistant that answers questions about a small firm's policy PDFs. Rung 1: a script that extracts text from the PDFs. Rung 2: chunk, retrieve, answer with citations, and score it on 25 real staff questions. Rung 3: probably skip it. A fixed retrieve-then-answer workflow does the job, and saying so is a result. Rung 4: put it behind a login with a daily cost cap and log every question. Rung 5: confirm staff can only retrieve documents they already have access to, and that a poisoned PDF saying "ignore your instructions" changes nothing. Notice that the agent rung disappeared. That is the ladder working.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo