All 99 articles, sorted alphabetically
Activation Engineering, in depth: steering vectors, how to extract and apply them, probes as monitors, and what steering means for LLM security
A practical guide to activation steering: the residual stream, contrastive activation addition, PyTorch hook code to extract and apply a steering vect…
Read article →Confused Deputy, in depth: authorizing agent actions with delegated tokens, the intersection rule and designation checks
How an AI agent becomes a confused deputy and how to authorize its actions properly: Hardy's authority-versus-designation split, the user-AND-age…
Read article →Agent Kill Switch, in depth: stop levels, out-of-band control, fail-closed leases, credential revocation and drills
How to build a kill switch for LLM agents that actually stops them: stop levels from soft cancel to network isolation, an out-of-band control plane, f…
Read article →Tool Bombs, in depth: how agents amplify one input into runaway tool chains, the cost arithmetic, and the budget tree that contains them
Tool bombs explained from first principles: fan-out, recursion, sub-agent spawning, loops, retry storms and oversized tool output that turn one input …
Read article →Backdoor Detection in Models, in depth: implementing data, model and input-time detectors and measuring them on planted backdoors
A practitioner's guide to backdoor detection: what each detector decides, spectral signatures and activation clustering with runnable cod…
Read article →Confidential Computing for LLM Inference, in depth: what the end user can verify, with attestation roles, encrypting to the node, OHTTP relays, transparency logs and metadata
Confidential LLM inference from the client's side: the RATS roles, why requests must be encrypted to an attested node rather than to a TLS fr…
Read article →Data Exfiltration via LLM Tools, in depth: every write-capable tool as a sink, a multi-call attack trace, data-flow labels at the tool gateway, covert-channel capacity and plan/execute separation
How tool-using LLM agents leak data through their own tool calls, and how to stop it: the private data, untrusted input and outbound channel condition…
Read article →EU AI Act, in depth: an engineering guide to classification, evidence and the 2026 to 2028 deadlines
The EU AI Act for engineers: roles, the risk tiers, how to classify a system, what high-risk obligations mean as artefacts (risk management, data gove…
Read article →Federated Prompt Engineering
Cross-org sharing of prompt patterns. Community-curated repositories, collective benefit and collective risk. Best practices for publishing, consuming…
Read article →GCG, in depth: how Greedy Coordinate Gradient finds universal adversarial suffixes, why they transfer, and how to defend and test against them
A defender's guide to the GCG attack on aligned language models: the target-string object…
Read article →Jailbreaks, in depth: the DAN lineage and its variants, why hand-written prompts work, and how to test against them
A defender's guide to hand-written LLM jailbreaks: what the DAN lineage and its variants (persona override, fiction a…
Read article →LLM Data Exfiltration
How data leaves an LLM app: injection plus an egress channel. Markdown image rendering, tool calls, side channels, and the controls that stop them.
Read article →LLM Denial of Service in Depth: The Cost of an Inference Request, the Attacks That Exploit It, and a Layered Admission Architecture
Why LLM endpoints fail differently under attack: the prefill, decode and KV-cache cost model, long-input, max-output, reasoning-inflation, agent-loop …
Read article →LLM Deployment Hardening
Hardening LLM deployments in production: container security with distroless images, network isolation with VPC and TLS, secrets with Vault, input vali…
Read article →LLM Hallucination Risk
Hallucination as a security and reliability risk: the harm taxonomy, fabricated identifiers, induced fabrication, and how to measure a hallucination r…
Read article →Indirect Prompt Injection in Depth: Threat Model, a Worked Attack Trace, and Architectural Defenses for Agents
How indirect prompt injection works when LLM applications read web pages, email, documents and tool output: the trust-boundary threat model, a step-by…
Read article →LLM Jailbreaking in Depth: Why Safety Training Fails, How Attacks Are Found, and How to Threat-Model Your Own Product
A first-principles guide to LLM jailbreaking: how jailbreaks differ from prompt injection, the research explaining why safety training fails (competin…
Read article →LLM PII Leakage
Why language models memorize personal data from their training corpus, what drives it (duplication, capacity, context), and the defenses that actually…
Read article →LLM Abuse Detection Architecture in Depth: Signals, Entity Risk Scores, Graduated Enforcement and Review
How to design abuse detection for an LLM product or API: a threat taxonomy, inline and asynchronous detection paths, the entity graph, features and de…
Read article →Agent hijacking, in depth: how untrusted data takes over tool-using agents, and the designs that stop it
Agent hijacking explained from the attack chain to the defences: how injected content redirects an agent&a…
Read article →Agent Tool Permissions: Least Privilege Between a Model and Real Actions
Why an LLM agent&a…
Read article →Agentic AI Boundaries, in depth: the Rule of Two as a session invariant, boundary manifests and tests that keep an agent inside its envelope
How to draw and enforce boundaries for AI agents: why per-tool permissions miss dangerous combinations, the Agents Rule of Two as a session-level inva…
Read article →AI Alignment, in depth: outer and inner alignment, specification gaming, reward overoptimization, sycophancy and deception, and how engineering teams turn them into evals and release gates
A first-principles map of AI alignment for practitioners: the gap between intended, specified, learned and deployed objectives, documented failure cla…
Read article →AI Bill of Rights and Regulation, in depth: turning notice, explanation, human fallback and non-discrimination into features an engineering team can build and prove
What the 2022 Blueprint for an AI Bill of Rights actually was, where its five principles live on in binding rules as of September 2026 (EU AI Act, Col…
Read article →AI Ethics, in depth: turning principles into decisions, controls and evidence for an LLM product
AI ethics as an engineering discipline: why principle lists do not ship, mapping stakeholders and harms, turning value conflicts into written decision…
Read article →AI Forensics, in depth: evidence, chain of custody, timelines, replay and context ablation for LLM and agent incidents
How to investigate an incident in an LLM application or agent: what makes AI evidence different, the evidence inventory and its volatility, preservati…
Read article →The AI Labs Security Ecosystem, in depth: ATLAS, OWASP, open scanners, lab programmes and institutes, and how an application team plugs in
A working map of the AI security ecosystem for engineers: what model labs publish, MITRE ATLAS and OWASP as shared vocabulary, open tooling such as ga…
Read article →AI Safety Frameworks, in depth: NIST AI RMF, ISO/IEC 42001, the EU AI Act and frontier-lab policies, crosswalked into controls, eval gates and evidence an engineering team can run
A practical guide to AI safety and risk frameworks for engineers: what the NIST AI RMF and its generative AI profile, ISO/IEC 42001, the EU AI Act (wi…
Read article →AI Worms, in depth: self-replicating prompts, how they propagate through RAG and agents, and how to keep R0 below 1
A first-principles guide to AI worms: what makes a prompt self-replicating, what the Morris II research demonstrated with GenAI email assistants, an e…
Read article →LLM audit logging architecture
Deep-dive on audit logging for LLM systems: capturing request context, assembled prompts, model I/O and tool effects; redaction to hashed placeholders…
Read article →LLM Audit Logs, in depth: the event schema, content vaults and crypto-shredding, coverage checks, detection queries and retention
How to design the records themselves for an LLM audit trail: what one event must answer, a schema that survives an investigation, keeping prompt conte…
Read article →LLM Audit Preparation, in depth: scoping, control matrices, evidence as code, sampling populations and a mock audit
How to prepare an LLM application for an external audit or assessment: what auditors actually test, scoping the system, building a control matrix for …
Read article →Model Backdoors, in depth: triggers, why safety training misses them, detection, weight supply chain and containment
How backdoors in language models work: trigger types, insertion points from pretraining data to published weights, what the Sleeper Agents and 250-doc…
Read article →AI Bug Bounties, in depth: what programs pay for, scoping an AI program, reproducing nondeterministic findings, severity, triage and reports that get paid
How bug bounties work for AI systems: what OpenAI, Anthropic, Google, Microsoft and AI-focused platforms reward and exclude, a finding taxonomy mapped…
Read article →LLM security canary tokens architecture
Deep-dive on canary tokens for LLM systems: why detection matters where prevention is incomplete, what makes a good canary (uniqueness, plausibility, …
Read article →Confidential Compute for LLM, in depth: building a measured inference server, GPU confidential mode, composite attestation, multi-GPU limits and day-two operations
An engineering guide to running LLM inference under confidential computing: who is protected from whom, the boot and attestation chain from CVM launch…
Read article →Confidential Prompts, in depth: what a prompt can keep secret, classifying its contents, server-side assembly and the leak paths that never touch the model
How to handle confidential prompts properly: why a system prompt cannot be a security boundary, classifying prompt contents into tiers, moving secrets…
Read article →The confused deputy in agentic LLM systems
Deep-dive on the confused-deputy vulnerability as it reappears in tool-using LLM agents: an agent holding ambient authority (OAuth tokens, DB creds, p…
Read article →Content Authentication, in depth: signing and verifying AI-generated media with C2PA manifests, watermarks and fingerprints
How content authentication works for AI-generated images, audio and video: C2PA manifests, hard and soft bindings, COSE signatures and time-stamps, th…
Read article →Context Smuggling, in depth: how untrusted text launders its way into trusted parts of an LLM's context, and how to stop it
A defender's guide to context smuggling i…
Read article →Data poisoning -- corrupting the model through its training data
Deep-dive on data poisoning: corrupting a model via malicious training data, poisoning types (backdoor/bias/degradation), backdoor triggers (hidden be…
Read article →LLM Defense in Depth: The Full Security Stack
A 2500-word walkthrough of a production LLM security architecture: perimeter, identity, rate, input, firewall, model, retrieval, tools, output, egress…
Read article →DP-SGD, in depth: per-example clipping, calibrated noise, privacy accounting and private LLM training
How differentially private SGD works and how to run it: the (epsilon, delta) guarantee and privacy units, Poisson sampling, per-example clipping, Gaus…
Read article →LLM egress filtering architecture
Deep-dive on egress filtering for LLM agents: the forced choke point, destination allowlists, content inspection for secrets and PII, canary tokens, m…
Read article →Encryption in Use for LLM Systems, in depth: mapping where plaintext lives, memory-encrypted VMs and GPUs, attestation-gated keys and shrinking the plaintext window
A practical guide to protecting data while it is being processed by an LLM service: where prompts and outputs exist as plaintext, what CPU and GPU mem…
Read article →EU AI Act, in depth: roles, risk classification, obligations and the post-Omnibus timeline for teams that ship AI
An engineering guide to the EU AI Act (Regulation 2024/1689) as of October 2026: the timeline after the AI Digital Omnibus, provider and deployer role…
Read article →LLM safety evals architecture
Safety evals for LLMs: grading harm rather than refusal strings, public benchmarks and their limits, over-refusal, bias, contamination, release gates.
Read article →AI Fairness, in depth: per-group confusion matrices, why the fairness criteria conflict, mitigation at three stages and counterfactual tests for LLM systems
A practitioner's guide to AI fairness: the kinds of harm, selection rate, TPR, FPR and PPV gaps from per-group co…
Read article →GDPR for LLM Applications, in depth: roles, lawful basis, data subject rights across every store, DPIAs, transfers and retention
How the GDPR applies to applications built on large language models: controller and processor roles with model providers, lawful basis and purpose lim…
Read article →Gemini Safety Filters, in depth: the three layers, where a block shows up in the response, calibrating thresholds and what the filter will never catch for you
How Gemini API safety filters work for an integrator: non-adjustable protections, per-category HarmBlockThreshold settings and the model&a…
Read article →Autonomous Hackbots, in depth: how LLM offensive-security agents work, what the evidence shows, and how to contain one
A defender's guide to autonomous offensive-security agents: the ReAct loop that makes an LLM act as a…
Read article →HIPAA for Healthcare LLM Applications, in depth: mapping PHI flows, business associate chains, minimum necessary prompts, de-identification and audit controls
An engineering guide to building LLM applications that handle protected health information under HIPAA: who is covered, where PHI lands in an LLM pipe…
Read article →Human-in-the-loop approval gates for AI agents
Deep-dive on human-in-the-loop approval gates for agent actions: a risk classifier that routes by risk and reversibility, an auto-execute path for saf…
Read article →LLM Incident Response, in depth: incident classes, severity, detection signals, a containment ladder, the first hour, recovery gates and postmortems for LLM and agent systems
A practitioner's guide to responding to incidents in LLM applications and agents: what counts as an incident, a severity model, detec…
Read article →Ingress Control for LLM APIs, in depth: identity, entitlement, request validation, token-budget admission, ingress screening and streaming limits
How to control what enters an LLM API before it reaches a model: an ordered admission pipeline covering identity, entitlements, strict request validat…
Read article →LLM Jailbreak Defense Architecture in Depth
A 2500-word walkthrough of jailbreak defense architecture: input classifier, safety-tuned model, system prompt hardening, constrained decoding, output…
Read article →JSON Injection Into LLMs, in depth: string-built prompts, poisoned fields, parser differentials and hardened parsing
How JSON injection works in LLM applications and how to stop it: breaking out of string-built JSON requests, instructions hidden in JSON values and ke…
Read article →LLM-Powered Bug Discovery, in depth: a pipeline where models propose vulnerabilities and sanitizers, tests and humans decide
How to use language models to find real vulnerabilities in your own code: what public results from Big Sleep, OSS-Fuzz, CVE-2025-37899 and AIxCC do an…
Read article →LLM Guard, in depth: the scanner chain, the Vault, the API server, verified failure modes and running an archived library safely
How Protect AI's open-source LLM Guard works and how to run it responsibly now that the project is archived: input and output scanner cha…
Read article →LLM Infrastructure Security, in depth: weights, serving endpoints, GPU nodes and the control plane around them
Securing the infrastructure that serves large language models: what is worth stealing, weight integrity and safe formats, inference and model-manageme…
Read article →Membership Inference Defenses, in depth: measuring leakage, choosing defenses and knowing which ones actually hold
A practical guide to defending models, including LLMs, against membership inference: how to measure leakage with TPR at low false-positive rates, LLM …
Read article →Membership-inference defense architecture
Deep-dive on defending LLMs against membership-inference attacks: how the train/test confidence gap left by overfitting lets an attacker read a model&…
Read article →Model extraction
Deep-dive on model extraction attacks and defenses: systematic API querying, response collection, surrogate training, behavior-not-weights theft, dete…
Read article →LLM moderation architecture
Deep-dive on layered LLM moderation: heuristics + ML + policy engine + human review with a training loop that closes on hard examples.
Read article →Moderation Bypass, in depth: how content filters around LLMs are evaded, and how to close each gap
A defender's guide to moderation bypass in LLM applicatio…
Read article →Multi-Turn Attacks, in depth: conversation state as the attack surface, escalation and splitting patterns, trajectory-level detection and red-teaming a dialogue
Why LLM attacks spread across many turns beat per-message filters: who can write conversation history, the escalation, payload-splitting, dilution, fo…
Read article →NIST AI Risk Management Framework, in depth: the four functions, 19 categories, profiles, the generative AI profile, and a risk register you can run
An engineer's guide to the NIST AI RMF (NIST AI 100-1): what it is and is not, the trustworthiness characteristics, the GOVERN, M…
Read article →LLM output handling -- treating model output as untrusted
Deep-dive on LLM output handling: the untrusted-output threat, output flowing into injection sinks (HTML/SQL/shell/code), context-aware encoding and e…
Read article →OWASP LLM Top 10, in depth: the 2026 edition, what moved since 2025, and how to run an assessment against it
How to use the OWASP Top 10 for LLM Applications as an assessment instrument rather than a poster: the 2026 ranking and what changed from 2025, keying…
Read article →OWASP LLM Top 10, in depth: how the risks chain into real attacks, and where to break each chain
The OWASP Top 10 for LLM applications read as a graph of entry points, amplifiers and impacts: three worked attack chains, one policy point in code th…
Read article →PCI-DSS for Payment-Handling LLM Applications, in depth: keeping card data out of prompts, logs and model providers, and what stays in scope
How PCI DSS v4.0.1 applies when an LLM agent takes payments: cardholder data versus sensitive authentication data, every place card numbers leak into …
Read article →LLM Pentesting, in depth: running an authorised assessment of one feature, from rules of engagement to retest
How to run an authorised security assessment of a single LLM feature: getting written authorisation and rules of engagement first, building a test pla…
Read article →LLM PII Detection and Redaction Architecture in Depth
Defensive PII engineering for LLM systems: what counts as PII and why it is jurisdictional, honest detection accuracy for patterns vs NER vs LLM passe…
Read article →Prompt Injection Scanners, in depth: what they can detect, where to put them, how to set thresholds and how to evaluate them honestly
A practical guide to prompt injection scanners: why detection is probabilistic, scanning every untrusted channel, layered normalisation, rules and cla…
Read article →LLM output provenance architecture
Deep-dive on LLM output provenance: source citations, watermark, audit records, verifier tool, provenance UI, policy, and metrics.
Read article →Meta Purple Llama, in depth: Llama Guard, Prompt Guard, CodeShield, LlamaFirewall and CyberSecEval wired into one runtime-defence and evaluation loop
A practical guide to Meta's Purple Llama toolkit: what Llama Guard, Prompt Guard 2, CodeShield, LlamaFirewall and CyberSecEva…
Read article →RAG Defense Architecture in Depth
A 2500-word walkthrough of RAG defense architecture: source allowlist, signing, sanitization, tenant-scoped retrieval, quarantine, citations, hallucin…
Read article →LLM red team architecture
Deep-dive on LLM red team: attack corpus, automated attacks, human red team, judge, coverage, reporting, fix cycle, metrics.
Read article →LLM Security Reference Architecture, in depth: trust zones, component contracts, request flows, policy as data and a control map for the OWASP LLM Top 10
A deployable security blueprint for LLM applications: five trust zones and their boundaries, the contract of each component (AI gateway, context assem…
Read article →AI Regulation Deep Dive, in depth: an applicability engine, a dated obligation register and the 2026 changes as a worked example
How an engineering team turns moving AI law into something it can run: system facts, roles and jurisdictions, a dated obligation register, an applicab…
Read article →Agent tool-execution sandboxing architecture
Deep-dive on sandboxing LLM agent tool calls: policy engines, Firecracker/gVisor/WASM isolation tiers, warm slot pools, secrets brokering, egress prox…
Read article →Secret Scanners, in depth: how detectors find credentials, where to scan LLM prompts, outputs and tool results, streaming redaction, verification and tuning
How secret scanners detect credentials and how to run them on LLM traffic: keyword prefilters, regex, entropy, checksummed token formats and live veri…
Read article →Secrets management architecture for LLM applications
Deep-dive on LLM secrets security: why context windows and prompt logs leak credentials, vault-issued short-lived credentials, opaque secret reference…
Read article →Secure Inference for LLMs, in depth: threat models, confidential GPUs and attestation, homomorphic encryption, MPC and what each one actually protects
How to run LLM inference so the operator cannot read prompts, outputs or weights: defining the threat model, confidential VMs plus GPU confidential co…
Read article →SOC 2 for LLM Applications, in depth: scoping the model provider, mapping the criteria to prompts, models and retrieval, and producing evidence an auditor can sample
How to bring an LLM feature through a SOC 2 examination: what the report is and is not, Type 1 versus Type 2, drawing the system boundary, treating th…
Read article →LLM spotlighting architecture
Deep-dive on spotlighting as a prompt-injection defense: why transformers lack a code-vs-data boundary, how delimiting, datamarking, and encoding mark…
Read article →System prompt leakage architecture
Deep-dive on LLM system prompt leakage: why the system role is not a privilege boundary, extraction via direct ask, roleplay, continuation, encoding, …
Read article →Multi-tenant LLM isolation architecture
Deep-dive on tenant isolation for LLM platforms: unforgeable tenant context, retrieval namespaces vs filters, cache-key design, per-tenant adapters, e…
Read article →Third-Party API Security, in depth: the trust boundary between LLM applications, model vendors and the APIs their tools call
A practical guide to securing the third-party APIs an LLM application depends on, in both directions: inventory and data classification, vendor terms …
Read article →Tool Abuse, in depth: how agents misuse legitimate tools, and the tool-call gateway that stops it
How LLM agents misuse the tools they are legitimately given: scope misuse, argument injection (path traversal, SSRF, SQL), dangerous chains, volume an…
Read article →Unicode smuggling defense architecture
Deep-dive on defending LLMs against Unicode smuggling: hidden Unicode Tag-block payloads, zero-width and bidirectional controls, homoglyph/confusable …
Read article →US AI Executive Orders, in depth: how presidential orders reach your AI system through OMB memos and contract clauses, and how to track them in code
An engineering guide to US AI executive orders from EO 14110 to EO 14365 and the 2026 orders: what an order can and cannot bind, the four paths by whi…
Read article →Membership Inference Attacks, in depth: how the scores work, why most LLM results were measuring the wrong thing, and how to run an honest audit on your own model
Membership inference against language models from first principles: threat models and what counts as a member, loss, reference-model, zlib, neighbourh…
Read article →Model Stealing, in depth: what an LLM's outputs and infrastructure leak, from logit geometry to weight custody
Model stealing beyond distillation: why full logit vectors reveal a transformer's hidden size and final l…
Read article →Output Grounding Verification, in depth: checking every claim in an LLM answer against its evidence, scoring support and failing safely
How to verify that an LLM's answer is supported by the evidence it was given: freezing an evidence ledger, extracting atomic claims, aligning cla…
Read article →OWASP Top 10 for LLM Applications, in depth: each 2025 risk, where it lands in your architecture, the controls that work, and how to test them
A practitioner's walk through the OWASP Top 10 for LLM Applications 2025, LLM01 to LL…
Read article →RAG Document Curation
How to prevent poisoned and malicious documents from entering a RAG system&amp…
Read article →Safe Completion via Planner Layer, in depth: choosing a response strategy before generation, the plan contract, verification, attacks on the planner and how to run it
How to build a planner layer that turns LLM safety from a refuse-or-comply switch into a chosen response strategy: the architecture, a typed response …
Read article →Scheming and Situational Awareness in LLMs, in depth: the threat model, what the evidence shows, and a control layer for deployed agents
What scheming and situational awareness mean for LLM security: the three ingredients of covert misaligned behaviour, what Sleeper Agents, in-context s…
Read article →