AI Guides › Playbooks

The Eight-Rung AI Engineer Ladder

By Nigel Guy · 8 min read

Most "learn AI engineering" roadmaps start at agents, because agents are what people are talking about. You build one in a weekend, it works on the demo input, and you have no idea why it fails on the next one. That is not a talent problem. It is a missing floor: nobody told you that an agent is a loop around model calls, tools and data, and that each of those has to behave before the loop does.

The rule: climb in order, and do not move up a rung until you can show one small thing you built or measured at the current one.

The mechanism: the Eight-Rung Ladder

This ladder is Juliemango's own ordering, not an industry standard. There is no official AI engineer syllabus, and job titles vary a lot. What can be verified is the shape of the dependencies. Anthropic's engineering write-up "Building effective agents" separates workflows (LLMs and tools orchestrated through predefined code paths) from agents (LLMs directing their own process and tool use), and advises finding the simplest solution possible and adding complexity only when needed. The ladder follows that advice: simple things first, agents late.

Each rung has five things to learn and one proof: an artefact you can show. If you cannot produce the proof, you are still on that rung.

Rung Name Proof
1 Code fluency A script you wrote, in git, that someone else can run
2 Data and the web A script that pulls from a web API and stores clean results
3 Calling a model A command-line tool that sends a prompt and prints cost
4 Prompts and structured output A function that returns validated JSON, with failures handled
5 Context and retrieval A question-answering tool over your own documents, with sources
6 Tools and workflows A multi-step workflow where the model calls tools you wrote
7 Evals and ML basics A test set, a score, and a before-and-after comparison
8 Production Logging, limits, secrets and a security review on something real

Rungs 1 to 3: the foundations

Rung 1, code fluency. Learn: Python syntax and functions; reading a stack trace without panic; virtual environments and installing packages; the command line; git basics (commit, branch, revert). The official Python tutorial is aimed at programmers new to Python, not people new to programming, so if you have never coded, expect to need a gentler first step.

Rung 2, data and the web. Learn: JSON in and out; HTTP requests and status codes; calling a REST API; basic SQL (select, filter, join); keeping API keys in environment variables, never in code. This rung is dull and it is where most later bugs live.

Rung 3, calling a model. Learn: the request and response shape of one provider's API (system prompt, messages, parameters); what tokens are and why they drive cost; streaming; rate-limit and error handling; reading the pricing page. Anthropic's features overview is a fair map: it groups the API into model capabilities, tools, tool infrastructure, context management and files. Pick one provider, get fluent, then compare a second.

Do this now, however early you are: from rung 3, keep ten real example inputs and what you would call a good answer for each in a plain file. That file becomes your rung-7 test set.

Rungs 4 to 6: where it becomes engineering

Rung 4, prompts and structured output. Learn: giving a role, context and a clear output format; few-shot examples; asking the model to say when information is missing; requesting JSON against a schema; validating the result in your own code and retrying or failing cleanly. Providers now offer schema-conformance features (Anthropic documents "structured outputs" with JSON outputs and strict tool use), but validate anyway, because your code is the last line of defence.

Rung 5, context and retrieval. Learn: what fits in a context window and what that costs; chunking documents; embeddings and similarity search; retrieving, then answering with citations to the source passage; trimming and summarising long conversations. The honest test is simple: can the tool say "not in the documents" instead of inventing something?

Rung 6, tools and workflows. Learn: tool (function) calling, where you describe a function and the model asks you to run it; the Model Context Protocol (MCP), an open standard for connecting AI applications to data sources, tools and workflows; the five workflow patterns Anthropic names (prompt chaining, routing, parallelisation, orchestrator-workers, evaluator-optimiser); and only then agents, meaning loops where the model chooses its own next step. Give every tool the least access it needs.

Rungs 7 and 8: the ML layer and production

Rung 7, evals and ML basics. Learn: train, validation and test splits and why you never tune on the test set; precision, recall and accuracy, and when each misleads; embeddings as a concept, not just an API call; evals, meaning tests that check model outputs against criteria you specify (OpenAI's evals guide describes defining the task, running tests, then iterating); and error analysis, which is reading the failures one by one. You do not need to train a neural network from scratch to be useful, but you do need to be able to say whether a change made things better.

Rung 8, production. Learn: logging prompts, outputs, cost and latency per request; caching and batching to cut spend (providers document both, so check the current terms); timeouts, retries and fallbacks; prompt injection and why untrusted text must never be able to trigger a risky tool unchecked; data retention and privacy rules for what you send to a vendor; and keeping a human approval step on anything irreversible. If you handle personal data, check UK GDPR obligations with a qualified person.

A worked example (hypothetical)

Imagine you are a support analyst, comfortable with spreadsheets, who wants to become the person who builds the team's ticket-triage tool. You score yourself: rung 1 yes (some Python), rung 2 half (you have never used SQL), rung 3 no. A common mistake is to jump to "build a triage agent". Using the ladder, your next four weeks are: finish SQL basics and one API call script (rungs 2 to 3), then a function that returns a category and urgency as validated JSON (rung 4), with your ten saved example tickets as the first test set. You have not built an agent, but you now have the parts one needs, and a number that says whether the thing works.

A prompt to place yourself on the ladder

Fill in your background, your goal and what you have actually built; the prompt asks the model to hold you to the proofs.

You are a patient senior engineer coaching someone towards an AI engineering role. Your job is to place me on an eight-rung ladder and set my next two-week plan.

Context:
- My background: [BACKGROUND]
- My goal: [GOAL, e.g. "build internal tools at work" or "get a junior AI engineer job"]
- Hours per week I can spend: [HOURS]
- Things I have built or measured so far: [LIST OF REAL PROJECTS OR "none"]

The ladder (name, then proof of completion):
1 Code fluency: a script in git that others can run
2 Data and the web: a script that pulls from a web API and stores clean results
3 Calling a model: a command-line tool that sends a prompt and prints cost
4 Prompts and structured output: a function returning validated JSON with failures handled
5 Context and retrieval: a question-answering tool over my documents, with sources
6 Tools and workflows: a multi-step workflow where the model calls tools I wrote
7 Evals and ML basics: a test set, a score, and a before-and-after comparison
8 Production: logging, limits, secrets and a security review on something real

Steps:
1. If any detail above is missing or vague, ask me up to five questions before continuing. Do not guess.
2. For each rung, say "proven", "partly proven" or "not proven", citing only what I told you. Do not give credit for courses watched or tools merely used.
3. Name the lowest rung that is not fully proven.
4. Give a two-week plan for that rung: five topics, one small project, and the exact proof I should end with.

Format: a table of the eight rungs with your verdict, then the plan as a numbered list.

Constraints: no invented course names, prices or statistics. If you recommend a resource, name the official documentation type and tell me to check it is current. Before answering, check that your verdicts match the evidence I gave and that the plan fits [HOURS].

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo