AI Guides › Playbooks
By Nigel Guy · 7 min read
Most people who want to become AI engineers start at the top. They install an agent framework, copy a demo that "books flights", and discover three weeks later that they cannot say why it fails on Tuesdays. It feels like progress because something runs. It is failing you because every layer underneath is unlearned, so every bug is a mystery.
The rule: climb one rung at a time, and do not step up until you can show a small piece of working evidence that you have finished the rung below.
Nothing here is a guarantee of a job or a timeline. The five rungs below are a sensible order of work, and the "roughly five months" figure is a planning estimate for someone studying part-time, not a measured outcome. Your pace will differ, and that is fine.
| Rung | You build | Evidence you can step up |
|---|---|---|
| 1. Groundwork | Python scripts that read files, call web APIs and handle errors | A small command-line tool in its own virtual environment, with a README, that someone else can run |
| 2. Building blocks | Calls to a model API with structured output, retrieval over your own documents, and a first eval | A question-answering script over 20 or more of your own documents, with a written test set and a score |
| 3. Agents | A model that chooses between tools in a loop | A workflow first, then an agent, with a log of every tool call and a hard step limit |
| 4. Shipping | The system behind a web endpoint, with config, cost limits, logging and tests | Deployed somewhere you can reach it, with a budget cap and a way to roll back |
| 5. Safety | Permissions, input and output handling, human approval and an incident plan | A written threat list checked against your build, with at least one test per item |
Safety is rung five in the table because it is the last thing you finish, not the first thing you think about. Start the habits at rung one (no keys in code) and carry them up.
Learn Python well enough to read someone else's script without fear. The official Python tutorial is written for people who already programme and covers the pieces you need: modules, errors and exceptions, file input and output, and a section on virtual environments and pip. Do those sections rather than a whole course.
Then learn four plain skills: calling a web API with an HTTP library and reading the JSON that comes back, storing secrets in environment variables rather than in the file, using Git so you can undo mistakes, and writing a test for a function. These are dull and they are the job.
Evidence to step up: a small tool that fetches something, transforms it, and fails politely when the network or the input is bad.
Now add the model. Read the API docs of the provider you pick (Anthropic and OpenAI both publish theirs) and work through, in this order: a single request and response, a system prompt, structured output so the reply arrives as validated JSON rather than prose, and then the idea of tokens, context limits and cost per call. Claude's features overview lists structured outputs, citations, prompt caching and token counting as separate features, which tells you where the cost and reliability levers live.
Next is retrieval: splitting your own documents into chunks, finding the relevant ones for a question, and giving them to the model with an instruction to answer only from them. Ask for citations so you can check the source. A plain keyword search is a legitimate first retriever; add embeddings when the plain version demonstrably misses things.
Last on this rung, and the one most people skip: an eval. Anthropic's testing guidance says to write specific, measurable success criteria, build test cases that mirror real use including awkward inputs (empty, overlong, ambiguous), and prefer many cheaply graded cases over a few hand-graded ones. Start with 20 questions and the answers you expect. Rerun them whenever you change a prompt.
Evidence to step up: a score on your test set that you can say out loud, and a list of the questions it gets wrong.
An agent is a model that decides its own next step, calls tools, reads the results and continues until it decides it is done. Anthropic's "Building effective agents" draws the useful line: workflows follow code paths you wrote; agents direct their own process. Its advice is to start with the simplest thing and add complexity only when simpler solutions fall short, reserving agents for open-ended problems where you cannot hard-code the steps.
So build the task as a fixed workflow first. If it works, you are done and it is cheaper and easier to debug. If it fails because the steps genuinely cannot be known in advance, make the model choose between two or three tools.
Treat tool descriptions as interface design: say what each tool does, what the inputs look like, and give an example and the edge cases. Then add three controls from day one: a maximum number of steps, a full log of every tool call and result, and a rule that anything which changes the outside world (sending, deleting, paying) needs a human yes.
Evidence to step up: a log you can read afterwards to explain every decision the agent made.
A notebook is not a service. Here you put the thing behind an endpoint and then deal with what changes when strangers use it:
Pin the model version where the provider lets you, and recheck the provider's deprecation notices, because models are retired on a schedule.
Safety is a set of ordinary engineering controls, not a separate department. Use the OWASP Top 10 for LLM Applications (2025 edition) as your checklist. The entries that bite builders first:
| OWASP item | What it means for your build |
|---|---|
| Prompt injection | Text in a web page, email or document can carry instructions. Treat retrieved content as data, never as commands. |
| Sensitive information disclosure | Do not put in the context what the user is not allowed to see. |
| Improper output handling | Validate and sanitise model output before it reaches a database, a shell or a browser. |
| Excessive agency | Give each tool the narrowest permission it needs. Prefer read-only. |
| Unbounded consumption | Rate limits, token limits and step limits. |
For a wider management frame, NIST's AI Risk Management Framework is voluntary guidance organised around Govern, Map, Measure and Manage, with a Generative AI Profile published in July 2024. NIST says the framework is currently under revision, so check the current version if you cite it.
Suppose you want an assistant that answers questions about a small firm's policy PDFs. Rung 1: a script that extracts text from the PDFs. Rung 2: chunk, retrieve, answer with citations, and score it on 25 real staff questions. Rung 3: probably skip it. A fixed retrieve-then-answer workflow does the job, and saying so is a result. Rung 4: put it behind a login with a daily cost cap and log every question. Rung 5: confirm staff can only retrieve documents they already have access to, and that a poisoned PDF saying "ignore your instructions" changes nothing. Notice that the agent rung disappeared. That is the ladder working.