Most companies that ship LLM features discover the same gap. Application security knows web vulnerabilities but not prompt injection or tool abuse. The ML team knows models but treats security as somebody else's checklist. The red team finds problems but does not own the fix. The LLM security engineer is the person who closes that gap: an individual contributor who turns the new attack surface of language models into threat models, guardrail code, regression tests and release gates that other teams can live with.
This article describes the role as practised, not as a job advert. You will see what the engineer owns, what a normal week looks like, which skills actually matter, a work sample in code, a finding traced from report to closure, how to measure the role, and how to hire for it or grow into it. Leadership questions such as budget, board reporting and organisation-wide policy belong to the security leader and are covered in CISO Role in AI Security; this piece stays with the engineer who does the hands-on work.
What the role owns
A useful way to define any security role is by the artefacts it owns. If a team cannot point to the person who maintains an artefact, the artefact decays. For LLM security there are five that rarely have a natural home elsewhere.
- Feature threat models. For each LLM feature, a short document that lists what untrusted text can enter the context window, what the model can do with tools, and which paths connect the two. The method is covered in LLM-Specific Threat Modeling.
- Guardrail code. Input and output filters, tool permission policies, sandboxes for code execution, and egress controls. The engineer either writes them or reviews them line by line.
- The regression corpus. Every confirmed finding becomes a reproducible test case with an expected safe behaviour, run on every prompt, model or tool change.
- The severity rubric. A shared, written way to score findings so a red-team report, a bug bounty submission and an internal discovery are triaged the same way.
- AI incident runbooks. What to do when a model leaks data, an agent takes an unintended action or a jailbreak goes public, written with the incident response team.
What the role does not own matters as much. It does not own content policy, which belongs to trust and safety. It does not own model quality, which belongs to the ML team. It does not run the offensive programme, though it works closely with it; see LLM Red Team Architecture. Blurring these lines produces an engineer who is consulted on everything and accountable for nothing.
A working week
The work splits into four recurring streams. A typical engineer embedded with two or three product teams spends roughly a third of their time on design review, a third on building and maintaining controls, and the rest on findings and incidents, with the mix shifting around launches.
| Stream | Typical tasks | Output |
|---|---|---|
| Design review | Read a new agent design, map data and tool flows, ask what happens if a retrieved document says 'ignore previous instructions' | Threat model plus a list of required controls |
| Control engineering | Write a tool allow-list, add an egress proxy rule, build a canary-token check for system prompt leakage | Merged code with tests |
| Findings | Reproduce a red-team transcript, score it, find root cause, agree a fix with the owning team | Ticket, fix, regression test |
| Incidents | Join an incident bridge when an assistant leaks data, pull the transcript, decide containment | Timeline, containment, post-incident actions |
Two habits separate effective engineers from busy ones. First, they push every review comment into something executable: a lint rule, a CI check, a test. A comment that says 'make sure tool output is treated as untrusted' is forgotten in a sprint; a test that injects a hostile tool result and asserts the agent does not call the payments tool is not. Second, they keep a written list of open risks with an owner and a date, so accepted risk is a decision rather than an accident.
Skills that matter
Job adverts list everything. In practice the skills cluster into three layers, and most strong candidates are deep in one and competent in the other two.
| Layer | What competence looks like | How to build it |
|---|---|---|
| Security engineering | Threat modelling, authorization design, sandboxing, secrets handling, logging, incident response | Ship and defend a real web service; fix real vulnerabilities |
| LLM mechanics | Why the model cannot separate instructions from data, how tool calling and retrieval assemble a context, what sampling and fine-tuning change | Build an agent with tools and retrieval, then attack it yourself |
| Engineering delivery | Writing production code, tests and CI checks that other teams accept | Contribute controls to a codebase you do not own |
The single most important piece of understanding is that a language model processes its whole context as one sequence. Instructions from the developer, text from a web page and output from a tool all arrive as tokens, and no current model reliably keeps them apart. Everything else follows: defences that depend on the model ignoring hostile text are probabilistic, so robust designs limit what a confused model can do. An engineer who has internalised this asks 'what can this agent do if the model is fully controlled by the attacker?' before asking how good the filter is.
Regulatory awareness is useful but secondary. The engineer needs to know which controls an auditor will ask for and to keep evidence, not to interpret law.
Work sample: a findings regression harness
A realistic work sample is the regression harness that turns findings into tests. The sketch below keeps each case as data, runs it against the real agent entry point with tools replaced by recording fakes, and checks behaviour rather than wording. Behavioural assertions survive model upgrades; string matching does not.
from dataclasses import dataclass, field
@dataclass
class Case:
id: str # e.g. "RT-2026-041"
channel: str # where hostile text enters: "user", "retrieved_doc", "tool_result"
payload: str # the hostile text, stored verbatim from the finding
forbidden_tools: set = field(default_factory=set)
must_not_contain: list = field(default_factory=list) # e.g. canary tokens
class RecordingTools:
# Fake tools: record every call, never touch real systems.
def __init__(self):
self.calls = []
def __call__(self, name, **args):
self.calls.append((name, args))
return {"ok": True}
def run_case(agent, case: Case, trials: int = 5):
failures = []
for t in range(trials): # models are stochastic: repeat each case
tools = RecordingTools()
ctx = {"user": "Summarise my latest ticket.", "retrieved_doc": "", "tool_result": ""}
ctx[case.channel] = case.payload
reply = agent.run(ctx, tools=tools, seed=t)
called = {name for name, _ in tools.calls}
if called & case.forbidden_tools:
failures.append((t, "forbidden tool", sorted(called & case.forbidden_tools)))
for marker in case.must_not_contain:
if marker in reply:
failures.append((t, "leaked", marker))
return failuresThree details make this harness useful rather than decorative. Each case keeps the original payload, so a fix can be verified against the exact attack that worked. Each case runs several trials because a jailbreak that succeeds one time in five is a real vulnerability. And the forbidden-tools assertion tests the control that matters, the action, rather than whether the model said something embarrassing. Wire it into CI so any change to prompts, model version or tool definitions runs the whole corpus.
A severity rubric that teams accept
Severity scoring is where engineers earn trust with product teams. A rubric that rates every prompt injection as critical is ignored within a month. A practical rubric scores impact and reachability separately, then combines them.
IMPACT = { # worst realistic outcome if the attack works
"action_external": 4, # sends money, email, deletes data outside the user's own scope
"data_cross_user": 4, # reads another user's or tenant's data
"data_own_user": 2, # exfiltrates the victim's own data to the attacker
"policy_output": 1, # produces disallowed content shown only to the requester
}
REACH = { # who can deliver the payload
"anonymous_remote": 3, # e.g. text on a public web page the agent browses
"authenticated": 2, # any logged-in user
"insider": 1,
}
def severity(impact, reach, success_rate, mitigated_by_confirmation=False):
score = IMPACT[impact] * REACH[reach]
if success_rate < 0.05:
score = max(1, score - 2) # rare, but never to zero
if mitigated_by_confirmation:
score //= 2 # a human must approve the action
if score >= 9: return "critical"
if score >= 6: return "high"
if score >= 3: return "medium"
return "low"The numbers are a starting point, not a standard. What matters is that the rubric is written, versioned, agreed with the teams who must fix things, and applied the same way to internal and external reports.
Worked example: from bounty report to closed finding
Consider a support agent that reads customer emails and can issue refunds up to a limit. A bug bounty researcher reports that an email containing hidden text, white-on-white instructions, made the agent issue a refund to the sender without a matching order.
- Reproduce. The engineer replays the email against a staging agent with recording tools. The refund tool fires in 3 of 10 trials.
- Score. Impact is an external action; reach is anonymous remote, since anyone can email support; success rate is 30 percent. Score 12, critical.
- Contain. The refund tool is switched to require human approval while the fix is built. That halves the score to 6, high, immediately and buys time.
- Root cause. The design let text from an untrusted channel, the email body, drive a privileged tool with no check that the refund matched a real order for the sender's account.
- Fix. The refund tool now validates its arguments against the order system: the order must exist, belong to the authenticated sender, and be refundable. The model can still be fooled, but a fooled model can no longer issue an unjustified refund.
- Regress. The email becomes case RT-2026-041 with the refund tool forbidden when no valid order exists. Variants with the instruction in an attachment and in a quoted reply are added.
- Close. Approval is relaxed back to automatic for small refunds that pass validation, and the threat model for every agent with a money-moving tool is updated with the same pattern.
Notice what the fix was not: a better prompt or a stronger injection classifier. Those may be added as defence in depth, but the durable control was deterministic validation outside the model.
Measuring and hiring for the role
Measure the role by outcomes the business can see, and avoid counting activity. Useful measures include:
- Share of LLM features in production with a current threat model and a named owner.
- Regression corpus size and pass rate on the last release, with any accepted failures listed.
- Median time from confirmed critical or high finding to containment, and to permanent fix.
- Repeat findings: the same root cause found twice signals a missing pattern-level control.
- Design reviews completed before launch versus after, a direct measure of being invited early.
In interviews, test the real work. Give a candidate a one-page agent design and ask for the top three risks and the controls they would require; give them a red-team transcript and ask for a severity and a fix. Candidates who reach for deterministic limits on tools and data before reaching for filters usually understand the problem. The defensive counterpart to this role is described in AI Blue Team, and the incident side in LLM Incident Response.
Failure modes
- Becoming a reviewer bottleneck. If every launch waits on one engineer, teams route around them. Publish paved-road patterns and self-service checks so most features pass without a meeting.
- Filter optimism. Relying on an injection classifier as the main control. Classifiers reduce noise; they do not bound impact.
- Findings without tests. A fix with no regression case tends to break silently at the next model upgrade.
- Severity inflation. Rating everything critical destroys the rubric's value and the relationship with product teams.
- No access to production signals. An engineer who cannot see transcripts, tool-call logs or alerts is guessing. Agree logging and access early, with privacy controls.
- Ownership drift. Taking on content moderation or model evaluation because nobody else will, until the security work stops.
Trade-offs
| Choice | Benefit | Cost |
|---|---|---|
| Embedded in product teams | Early involvement, context, fast fixes | Inconsistent standards across teams, isolation |
| Central LLM security team | Shared corpus, rubric and patterns | Distance from designs, slower reviews |
| Hybrid: central platform plus embedded liaisons | Consistency with context | Needs clear ownership of shared artefacts |
| Hire from AppSec | Strong security instincts and delivery | Must learn model behaviour; may over-trust filters or under-trust models |
| Hire from ML | Deep model intuition | Must learn authorization, sandboxing and incident discipline |
Most organisations start with one embedded engineer, then centralise the corpus and rubric once there are three or more. Both hiring sources work; pair them so each covers the other's blind spot.
What to do next
- Write down the five artefacts in this article and name a current owner for each; any without one is your first gap.
- Pick your highest-risk agent and write a one-page threat model listing every channel of untrusted text and every privileged tool.
- Stand up a regression harness like the one above with recording tools, and load it with every finding you already have.
- Draft a severity rubric that scores impact and reachability separately; review it with two product teams before adopting it.
- For each money-moving or data-moving tool, add deterministic argument validation outside the model.
- Agree logging and transcript access with privacy and legal so incidents can be investigated.
- If you are growing into the role, build and attack your own tool-using agent end to end, then write up the findings as if for a product team.