An AI engineer builds products on top of foundation models they did not train and do not control. The work is closer to backend engineering than to machine learning research: the model is a remote dependency with a latency, a cost, a failure rate and, unusually, output that varies from call to call. The job is to turn that dependency into a feature that is reliable, measurable and safe.

This roadmap walks through seven stages in the order they build on each other, from the first API call to agents in production. Each stage adds one capability and, just as important, one way to check that it works. Code is in Python with the Anthropic SDK; the concepts carry over to other providers. GPU internals, training and serving infrastructure belong to the ML systems track and are left out here.

Advertisement

The stage map

Resist the temptation to start at agents. Every agent failure you will debug is a failure of an earlier stage: a context window stuffed with irrelevant text, a tool whose errors were swallowed, a retrieval step that returned the wrong document, or a change nobody measured. Evals sit in the middle of the map because from stage five onward every change to a prompt, model, retriever or tool should pass through them.

AI engineer roadmap: each stage adds one capability and one way to measure it1. First API calltokens, context, cost2. Structured outputschemas, prompts3. Toolsloop, errors, limits4. Retrievalchunks, hybrid, ACLs5. Evalsgolden set, graders, CI6. Agentscontext, memory, HITL7. Productiongateway, cost, tracesSecurityinjection, least privilegeevals gate every later change: prompts, models, retrieval and tools
Seven stages plus security. The dashed line is the feedback loop: from stage five on, evaluation gates every change.

Stage 1: the first API call

Start with a plain request and read everything that comes back, not just the text. A model sees text as tokens; you pay per input and output token, the context window limits how many tokens one request can hold, and output length drives latency because tokens are generated one at a time. The stop_reason tells you why generation ended: finished normally, hit your max_tokens cap, asked to use a tool, or declined. Code that ignores it will one day show a user half a sentence.

# pip install anthropic   (reads ANTHROPIC_API_KEY from the environment)
import anthropic

client = anthropic.Anthropic()
MODEL = "claude-opus-5-5"

resp = client.messages.create(
    model=MODEL,
    max_tokens=1024,
    system="You summarise customer support tickets in one sentence.",
    messages=[{"role": "user", "content": "My invoice for September was charged twice..."}],
)
if resp.stop_reason == "refusal":
    raise RuntimeError("model declined")
text = next(b.text for b in resp.content if b.type == "text")
print(text)
print(resp.stop_reason, resp.usage.input_tokens, resp.usage.output_tokens)

Log token usage from day one; cost problems are always discovered late otherwise. To understand what happens inside the call, read how tokenizers work, the transformer, step by step and the KV cache, which explains why long prompts cost time.

Advertisement

Stage 2: prompts and structured output

Free text is hard to route, store or test. As soon as model output feeds code, constrain it to a schema. With the Python SDK, messages.parse takes a Pydantic model and returns a validated instance, so a misspelled category never reaches your routing logic.

from typing import Literal
from pydantic import BaseModel

class Triage(BaseModel):
    category: Literal["billing", "bug", "account", "other"]
    urgency: Literal["low", "normal", "high"]
    summary: str

resp = client.messages.parse(
    model=MODEL,
    max_tokens=1024,
    system=TRIAGE_RULES,              # stable text first: it caches well
    messages=[{"role": "user", "content": ticket_text}],
    output_format=Triage,
)
triage = resp.parsed_output           # a validated Triage instance
route(triage.category, triage.urgency)

Treat the system prompt as code: version it, review changes, and keep stable instructions at the start so provider-side prompt caching can reuse them. Clear delimiters between instructions and untrusted content matter both for quality and for injection resistance. Further reading: structured output patterns, XML delimiters, prompt caching and prompt registries.

Stage 3: tools and the agentic loop

Tools let the model ask your code to do something: fetch an invoice, search a catalogue, run a query. The model never executes anything itself. It returns a tool_use block, your code runs the handler, and you send back a tool_result with the matching id. Write this loop by hand once before adopting a framework, because every framework is a variation of it.

TOOLS = [{
    "name": "get_invoice",
    "description": "Fetch one invoice for the signed-in customer by invoice id.",
    "input_schema": {
        "type": "object",
        "properties": {"invoice_id": {"type": "string"}},
        "required": ["invoice_id"],
        "additionalProperties": False,
    },
}]

def run(user_text, customer_id, max_turns=8):
    messages = [{"role": "user", "content": user_text}]
    for _ in range(max_turns):
        resp = client.messages.create(model=MODEL, max_tokens=4096,
                                      tools=TOOLS, messages=messages)
        if resp.stop_reason in ("refusal", "max_tokens"):
            return escalate(resp.stop_reason)
        messages.append({"role": "assistant", "content": resp.content})
        if resp.stop_reason != "tool_use":
            return next(b.text for b in resp.content if b.type == "text")
        results = []
        for block in resp.content:
            if block.type != "tool_use":
                continue
            try:
                # customer_id comes from the session, never from the model
                out = HANDLERS[block.name](customer_id=customer_id, **block.input)
                results.append({"type": "tool_result", "tool_use_id": block.id,
                                "content": json.dumps(out)})
            except Exception as e:
                results.append({"type": "tool_result", "tool_use_id": block.id,
                                "content": f"error: {e}", "is_error": True})
        messages.append({"role": "user", "content": results})   # all results, one message
    return escalate("turn limit")

Four details in that loop prevent most production incidents. The turn limit stops runaway loops. Tool failures are returned with is_error rather than raised, so the model can recover or explain. All results from one turn go back in a single message. And identity (customer_id) comes from the authenticated session, never from model output, so a manipulated prompt cannot fetch another customer's invoice. Further reading: tool use in depth, tool selection, tool retries, the Model Context Protocol and MCP tools.

Stage 4: retrieval

Models do not know your documents, and their knowledge stops at a training cutoff. Retrieval fetches relevant passages at request time and puts them in the context. The pipeline has more failure points than it first appears: documents are split into chunks, each chunk is embedded, queries are matched by vector similarity and usually by keyword search too, results are reranked, and the top few are inserted into the prompt.

Most retrieval quality problems are chunking and query problems, not model problems. Measure retrieval separately from generation: for a set of questions, did the right passage appear in the top five at all? And enforce access control at retrieval time, filtering by the caller's permissions before ranking, because a model will happily summarise a document the user was never allowed to see. Further reading: RAG pipelines at scale, ACL-aware retrieval, incremental embedding pipelines, vector search infrastructure and RAG evaluation.

Stage 5: evaluation

Evaluation is the skill that most separates an AI engineer from someone who has called an API. Because output varies, you cannot assert one example and move on; you need a set of representative cases, a way to grade them, and a baseline to compare against. Start with 50 to 200 real cases, labelled by hand, covering common inputs and the hard ones you have already seen fail.

# evals/triage.jsonl: {"ticket": "...", "category": "billing", "urgency": "high"}
import json, collections

def run_eval(path, classify):
    per_cat = collections.defaultdict(lambda: [0, 0])
    for line in open(path, encoding="utf-8"):
        case = json.loads(line)
        got = classify(case["ticket"])
        ok = got.category == case["category"] and got.urgency == case["urgency"]
        per_cat[case["category"]][0] += ok
        per_cat[case["category"]][1] += 1
    for cat, (ok, n) in sorted(per_cat.items()):
        print(f"{cat:10s} {ok}/{n}")
    total = sum(v[0] for v in per_cat.values()) / sum(v[1] for v in per_cat.values())
    return total

# CI: fail if any category drops more than 2 points below the stored baseline

Use exact checks wherever the output is structured, and a model-as-judge only for qualities that cannot be checked mechanically, calibrated against human labels. Report results per category, because an average hides a category that collapsed. Run the set in CI on every prompt, model or retrieval change. Further reading: golden datasets, eval harnesses and regression tests, LLM-as-judge calibration and agent evaluation at scale.

Stage 6: agents

An agent is the tool loop given a larger goal and more autonomy: it plans, calls many tools, and decides when it is done. Use one only when the task genuinely needs open-ended, multi-step exploration and when errors can be caught and reversed. A fixed workflow in which code calls the model at defined points is cheaper, faster and easier to test, and covers most product features.

When you do build an agent, the hard problems are context and control. Long runs fill the context window, so you need compaction or memory. Irreversible actions need a human approval step. Code execution needs a sandbox. And the agent must be evaluated on whole trajectories, not single replies. Further reading: the ReAct pattern, context engineering, context compaction, human in the loop, sandboxing and agents as state machines.

Stage 7: production

Production adds the concerns every backend service has, plus a few specific to models. Route calls through a gateway that centralises credentials, rate limits, retries, fallbacks and cost accounting per feature and per tenant. Set deadlines on every call and stream long responses. Cache stable prompt prefixes. Trace each request end to end, including every tool call and retrieved document, so that a bad answer can be explained. Roll out prompt and model changes behind flags, with the eval set as the gate.

Security deserves its own review. Any text the model reads, whether a web page, an email or a retrieved document, can contain instructions. Treat model output as untrusted input to your own systems, give tools the narrowest permissions possible, and restrict where data can be sent. Further reading: LLM gateway architecture, cost and latency budgets, agent tracing, prompt injection, the confused deputy, egress filtering and output handling.

Worked example: a support triage assistant in ten weeks

A team receives about two thousand support tickets a week and wants them categorised, prioritised and, for billing questions, answered with the customer's actual invoice data. Here is how the stages map onto delivery.

  1. Weeks 1-2, stages 1 and 2. The Triage schema and a system prompt. Two hundred historical tickets are labelled by the support leads; this becomes the eval set before any prompt tuning starts.
  2. Week 3, stage 5. The harness runs in CI. The first prompt scores well on bugs and poorly on account issues, which reveals that the category definitions were ambiguous. The fix is in the rules, not the model.
  3. Weeks 4-5, stage 3. A read-only get_invoice tool, scoped to the session's customer. Tool errors are returned to the model, and an adversarial test tries to fetch another customer's invoice through the ticket text.
  4. Weeks 6-7, stage 4. Retrieval over the help centre so replies cite current policy. Retrieval recall is measured separately; chunking by heading instead of fixed length lifts it noticeably.
  5. Weeks 8-10, stage 7. Gateway, per-ticket cost logging, traces, and a shadow period in which drafts are shown to agents but not sent. Only then are low-urgency billing replies sent automatically, with a one-click human override.

The team never needed a full autonomous agent. A workflow with one tool and retrieval covered the requirement, and every stage had a measurement attached before the next one began.

Failure modes

  • Vibes-based iteration. Tuning prompts against the three examples you remember. Without an eval set, every fix breaks something you are not looking at.
  • Unchecked stop reasons. Truncated or declined responses reaching users or parsers.
  • Identity from the model. Letting the model supply user ids or tenant ids to tools, which turns prompt injection into data access.
  • Context stuffing. Pasting everything into the prompt. Cost and latency rise, and irrelevant text degrades answers.
  • Agent by default. Building an autonomous loop for a job a fixed workflow would do more cheaply and predictably.
  • No cost attribution. One feature quietly consuming most of the budget, discovered on the invoice.

Trade-offs

Framework or hand-written loop. Frameworks speed up the first demo and hide the loop you will need to debug. Write it by hand first; adopt a framework when you know which parts it saves you. Bigger or smaller model. Choose per route with your eval set, measuring cost per completed task rather than per request. Retrieval or fine-tuning. Retrieval handles knowledge that changes and needs citations; fine-tuning shapes format and style. Most product teams need retrieval first. See quality, cost and latency frontiers and LoRA and QLoRA.

What to do next

  • Make one API call, print the stop reason and token usage, and write down what each field means.
  • Convert one free-text feature to a schema with messages.parse.
  • Hand-write the tool loop with a turn limit, is_error results and session-derived identity.
  • Label 50 real cases and build an eval harness that reports per category; add it to CI.
  • Add retrieval to one feature and measure retrieval recall separately from answer quality.
  • Put all model calls behind a gateway with deadlines, cost logging and tracing.
  • Run an adversarial test: hide an instruction in a document and check that no tool is misused.
Key takeaway: Become an AI engineer in stages: make and fully read a first API call, constrain output to schemas, hand-write a correct tool loop, add retrieval with access control, and build an evaluation set that gates every later change. Reach for agents only when a task needs open-ended exploration and errors can be reversed, and run everything through production disciplines: gateways, deadlines, cost attribution, tracing, least-privilege tools and a standing assumption that any text the model reads may be hostile.