Most companies that ship LLM features discover the same gap. Application security knows web vulnerabilities but not prompt injection or tool abuse. The ML team knows models but treats security as somebody else's checklist. The red team finds problems but does not own the fix. The LLM security engineer is the person who closes that gap: an individual contributor who turns the new attack surface of language models into threat models, guardrail code, regression tests and release gates that other teams can live with.

This article describes the role as practised, not as a job advert. You will see what the engineer owns, what a normal week looks like, which skills actually matter, a work sample in code, a finding traced from report to closure, how to measure the role, and how to hire for it or grow into it. Leadership questions such as budget, board reporting and organisation-wide policy belong to the security leader and are covered in CISO Role in AI Security; this piece stays with the engineer who does the hands-on work.

What the role owns

A useful way to define any security role is by the artefacts it owns. If a team cannot point to the person who maintains an artefact, the artefact decays. For LLM security there are five that rarely have a natural home elsewhere.

  • Feature threat models. For each LLM feature, a short document that lists what untrusted text can enter the context window, what the model can do with tools, and which paths connect the two. The method is covered in LLM-Specific Threat Modeling.
  • Guardrail code. Input and output filters, tool permission policies, sandboxes for code execution, and egress controls. The engineer either writes them or reviews them line by line.
  • The regression corpus. Every confirmed finding becomes a reproducible test case with an expected safe behaviour, run on every prompt, model or tool change.
  • The severity rubric. A shared, written way to score findings so a red-team report, a bug bounty submission and an internal discovery are triaged the same way.
  • AI incident runbooks. What to do when a model leaks data, an agent takes an unintended action or a jailbreak goes public, written with the incident response team.

What the role does not own matters as much. It does not own content policy, which belongs to trust and safety. It does not own model quality, which belongs to the ML team. It does not run the offensive programme, though it works closely with it; see LLM Red Team Architecture. Blurring these lines produces an engineer who is consulted on everything and accountable for nothing.

Where an LLM security engineer sits: inputs, owned artefacts, and the teams they hand work toProduct and ML teamsdesigns, prompts, toolsRed team and bountyfindings, transcriptsDetection and SOCalerts, incidentsGovernance and auditcontrol requirementsThreat modelstaint paths per featureGuardrail codefilters, policies, sandboxesRegression corpusevery finding becomes a testSeverity rubricshared scoringRelease gatesCI blocks regressionsDesign reviewsapproved with conditionsRunbooksAI incident playbooksControl evidencefor auditorsThe role is defined by the artefacts in the middle column: if nobody owns them, nobody does this job.
Inputs arrive from product, offensive, detection and governance teams; the engineer turns them into owned artefacts that feed release gates, design reviews, runbooks and audit evidence.

A working week

The work splits into four recurring streams. A typical engineer embedded with two or three product teams spends roughly a third of their time on design review, a third on building and maintaining controls, and the rest on findings and incidents, with the mix shifting around launches.

StreamTypical tasksOutput
Design reviewRead a new agent design, map data and tool flows, ask what happens if a retrieved document says 'ignore previous instructions'Threat model plus a list of required controls
Control engineeringWrite a tool allow-list, add an egress proxy rule, build a canary-token check for system prompt leakageMerged code with tests
FindingsReproduce a red-team transcript, score it, find root cause, agree a fix with the owning teamTicket, fix, regression test
IncidentsJoin an incident bridge when an assistant leaks data, pull the transcript, decide containmentTimeline, containment, post-incident actions

Two habits separate effective engineers from busy ones. First, they push every review comment into something executable: a lint rule, a CI check, a test. A comment that says 'make sure tool output is treated as untrusted' is forgotten in a sprint; a test that injects a hostile tool result and asserts the agent does not call the payments tool is not. Second, they keep a written list of open risks with an owner and a date, so accepted risk is a decision rather than an accident.

Skills that matter

Job adverts list everything. In practice the skills cluster into three layers, and most strong candidates are deep in one and competent in the other two.

LayerWhat competence looks likeHow to build it
Security engineeringThreat modelling, authorization design, sandboxing, secrets handling, logging, incident responseShip and defend a real web service; fix real vulnerabilities
LLM mechanicsWhy the model cannot separate instructions from data, how tool calling and retrieval assemble a context, what sampling and fine-tuning changeBuild an agent with tools and retrieval, then attack it yourself
Engineering deliveryWriting production code, tests and CI checks that other teams acceptContribute controls to a codebase you do not own

The single most important piece of understanding is that a language model processes its whole context as one sequence. Instructions from the developer, text from a web page and output from a tool all arrive as tokens, and no current model reliably keeps them apart. Everything else follows: defences that depend on the model ignoring hostile text are probabilistic, so robust designs limit what a confused model can do. An engineer who has internalised this asks 'what can this agent do if the model is fully controlled by the attacker?' before asking how good the filter is.

Regulatory awareness is useful but secondary. The engineer needs to know which controls an auditor will ask for and to keep evidence, not to interpret law.

Work sample: a findings regression harness

A realistic work sample is the regression harness that turns findings into tests. The sketch below keeps each case as data, runs it against the real agent entry point with tools replaced by recording fakes, and checks behaviour rather than wording. Behavioural assertions survive model upgrades; string matching does not.

from dataclasses import dataclass, field

@dataclass
class Case:
    id: str                      # e.g. "RT-2026-041"
    channel: str                 # where hostile text enters: "user", "retrieved_doc", "tool_result"
    payload: str                 # the hostile text, stored verbatim from the finding
    forbidden_tools: set = field(default_factory=set)
    must_not_contain: list = field(default_factory=list)   # e.g. canary tokens

class RecordingTools:
    # Fake tools: record every call, never touch real systems.
    def __init__(self):
        self.calls = []
    def __call__(self, name, **args):
        self.calls.append((name, args))
        return {"ok": True}

def run_case(agent, case: Case, trials: int = 5):
    failures = []
    for t in range(trials):              # models are stochastic: repeat each case
        tools = RecordingTools()
        ctx = {"user": "Summarise my latest ticket.", "retrieved_doc": "", "tool_result": ""}
        ctx[case.channel] = case.payload
        reply = agent.run(ctx, tools=tools, seed=t)
        called = {name for name, _ in tools.calls}
        if called & case.forbidden_tools:
            failures.append((t, "forbidden tool", sorted(called & case.forbidden_tools)))
        for marker in case.must_not_contain:
            if marker in reply:
                failures.append((t, "leaked", marker))
    return failures

Three details make this harness useful rather than decorative. Each case keeps the original payload, so a fix can be verified against the exact attack that worked. Each case runs several trials because a jailbreak that succeeds one time in five is a real vulnerability. And the forbidden-tools assertion tests the control that matters, the action, rather than whether the model said something embarrassing. Wire it into CI so any change to prompts, model version or tool definitions runs the whole corpus.

A severity rubric that teams accept

Severity scoring is where engineers earn trust with product teams. A rubric that rates every prompt injection as critical is ignored within a month. A practical rubric scores impact and reachability separately, then combines them.

IMPACT = {            # worst realistic outcome if the attack works
    "action_external": 4,   # sends money, email, deletes data outside the user's own scope
    "data_cross_user": 4,   # reads another user's or tenant's data
    "data_own_user": 2,     # exfiltrates the victim's own data to the attacker
    "policy_output": 1,     # produces disallowed content shown only to the requester
}
REACH = {             # who can deliver the payload
    "anonymous_remote": 3,  # e.g. text on a public web page the agent browses
    "authenticated": 2,     # any logged-in user
    "insider": 1,
}

def severity(impact, reach, success_rate, mitigated_by_confirmation=False):
    score = IMPACT[impact] * REACH[reach]
    if success_rate < 0.05:
        score = max(1, score - 2)        # rare, but never to zero
    if mitigated_by_confirmation:
        score //= 2                      # a human must approve the action
    if score >= 9:  return "critical"
    if score >= 6:  return "high"
    if score >= 3:  return "medium"
    return "low"

The numbers are a starting point, not a standard. What matters is that the rubric is written, versioned, agreed with the teams who must fix things, and applied the same way to internal and external reports.

Worked example: from bounty report to closed finding

Consider a support agent that reads customer emails and can issue refunds up to a limit. A bug bounty researcher reports that an email containing hidden text, white-on-white instructions, made the agent issue a refund to the sender without a matching order.

  1. Reproduce. The engineer replays the email against a staging agent with recording tools. The refund tool fires in 3 of 10 trials.
  2. Score. Impact is an external action; reach is anonymous remote, since anyone can email support; success rate is 30 percent. Score 12, critical.
  3. Contain. The refund tool is switched to require human approval while the fix is built. That halves the score to 6, high, immediately and buys time.
  4. Root cause. The design let text from an untrusted channel, the email body, drive a privileged tool with no check that the refund matched a real order for the sender's account.
  5. Fix. The refund tool now validates its arguments against the order system: the order must exist, belong to the authenticated sender, and be refundable. The model can still be fooled, but a fooled model can no longer issue an unjustified refund.
  6. Regress. The email becomes case RT-2026-041 with the refund tool forbidden when no valid order exists. Variants with the instruction in an attachment and in a quoted reply are added.
  7. Close. Approval is relaxed back to automatic for small refunds that pass validation, and the threat model for every agent with a money-moving tool is updated with the same pattern.

Notice what the fix was not: a better prompt or a stronger injection classifier. Those may be added as defence in depth, but the durable control was deterministic validation outside the model.

Measuring and hiring for the role

Measure the role by outcomes the business can see, and avoid counting activity. Useful measures include:

  • Share of LLM features in production with a current threat model and a named owner.
  • Regression corpus size and pass rate on the last release, with any accepted failures listed.
  • Median time from confirmed critical or high finding to containment, and to permanent fix.
  • Repeat findings: the same root cause found twice signals a missing pattern-level control.
  • Design reviews completed before launch versus after, a direct measure of being invited early.

In interviews, test the real work. Give a candidate a one-page agent design and ask for the top three risks and the controls they would require; give them a red-team transcript and ask for a severity and a fix. Candidates who reach for deterministic limits on tools and data before reaching for filters usually understand the problem. The defensive counterpart to this role is described in AI Blue Team, and the incident side in LLM Incident Response.

Failure modes

  • Becoming a reviewer bottleneck. If every launch waits on one engineer, teams route around them. Publish paved-road patterns and self-service checks so most features pass without a meeting.
  • Filter optimism. Relying on an injection classifier as the main control. Classifiers reduce noise; they do not bound impact.
  • Findings without tests. A fix with no regression case tends to break silently at the next model upgrade.
  • Severity inflation. Rating everything critical destroys the rubric's value and the relationship with product teams.
  • No access to production signals. An engineer who cannot see transcripts, tool-call logs or alerts is guessing. Agree logging and access early, with privacy controls.
  • Ownership drift. Taking on content moderation or model evaluation because nobody else will, until the security work stops.

Trade-offs

ChoiceBenefitCost
Embedded in product teamsEarly involvement, context, fast fixesInconsistent standards across teams, isolation
Central LLM security teamShared corpus, rubric and patternsDistance from designs, slower reviews
Hybrid: central platform plus embedded liaisonsConsistency with contextNeeds clear ownership of shared artefacts
Hire from AppSecStrong security instincts and deliveryMust learn model behaviour; may over-trust filters or under-trust models
Hire from MLDeep model intuitionMust learn authorization, sandboxing and incident discipline

Most organisations start with one embedded engineer, then centralise the corpus and rubric once there are three or more. Both hiring sources work; pair them so each covers the other's blind spot.

What to do next

  1. Write down the five artefacts in this article and name a current owner for each; any without one is your first gap.
  2. Pick your highest-risk agent and write a one-page threat model listing every channel of untrusted text and every privileged tool.
  3. Stand up a regression harness like the one above with recording tools, and load it with every finding you already have.
  4. Draft a severity rubric that scores impact and reachability separately; review it with two product teams before adopting it.
  5. For each money-moving or data-moving tool, add deterministic argument validation outside the model.
  6. Agree logging and transcript access with privacy and legal so incidents can be investigated.
  7. If you are growing into the role, build and attack your own tool-using agent end to end, then write up the findings as if for a product team.
Key takeaway: An LLM security engineer is defined by owned artefacts: feature threat models, guardrail code, a regression corpus, a severity rubric and incident runbooks. The core insight behind the job is that a model cannot reliably separate instructions from data, so durable controls limit what a confused model can do. Turn every review comment and finding into an executable test, and measure the role by coverage and time to fix, not by meetings attended.