Most writing on LLM penetration testing is about prompts: injection, jailbreaks, data pulled out through the model's answers. That work matters, and LLM pentesting in depth covers how to run it against a single feature. But when a team hosts its own models on its own GPUs, the prompt is the least privileged way in. The model server is a network service written in Python and C++; it loads files that can contain executable code; it talks to peers over internal sockets; and it runs in containers that need privileged access to a GPU driver. Every one of those layers has produced critical, publicly disclosed vulnerabilities since 2023.

This article is a defensive test plan for that stack: how to scope an authorised assessment of a self-hosted inference platform, what to examine at each layer, how to check for exposure without causing harm, and how to report findings so the platform team can fix them. It assumes written authorisation from the system owner and a defined window. Everything here is a check against infrastructure you are permitted to test; none of it is an exploit, and the version numbers below are there so you can confirm a system is already patched.

Why GPU inference is a different target

Three properties make GPU inference infrastructure different from an ordinary web service. First, the software is young and moves fast: serving engines ship every few weeks and add distributed features such as disaggregated prefill and KV-cache transfer, many first designed for trusted research clusters. Second, models are code. A checkpoint saved in Python's pickle format can run arbitrary code when loaded, and model repositories can ask the loader to execute custom Python through a trust_remote_code flag. Third, GPUs are expensive, so they are shared between teams, tenants and jobs; the isolation that sharing relies on is weaker and less understood than CPU virtualisation.

The consequence is that one misconfiguration can be worth more to an attacker than any prompt. An unauthenticated model API reachable from the internet gives away inference at your expense and leaks whatever is in its context. An unauthenticated cluster or job endpoint can mean code execution on GPU nodes, which usually hold cloud credentials, weights and training data; a container escape on a shared node gives the host and every tenant on it. A good assessment spends most of its time below the prompt.

Authorisation and rules of engagement, first

Before any traffic, get the scope in writing. Name the exact hosts, clusters, namespaces and model endpoints in scope, and what is explicitly out of scope, such as the cloud provider's control plane. Agree a window, a rate limit, and a named contact who can stop the test, and decide how findings that touch live customer data are handled before you find one. For a shared GPU platform, settle one question early: are you testing as an outside attacker, or as a hostile tenant who already has a legitimate account? The two produce very different findings, and tenant-isolation testing can affect the neighbours, so it usually needs its own window and sign-off.

Capture non-destructive evidence: request and response pairs, version banners, configuration snippets. Agree to demonstrate access rather than exercise it, for example by reading a canary file placed for the test rather than real data. The deliverable is a report the platform team can act on, not a trophy.

Map the stack before you test it

Self-hosted inference stack: the layers an assessment walksClients and appsSDKs, agents, CI jobsGatewayauth, quotas, loggingModel servervLLM, Triton, OllamaSide portsmetrics, admin, clusterModel storeweights, tokenizer, codeInternal fabricZeroMQ, NCCL, KV transferOrchestratorKubernetes, Ray, SlurmloadContainer runtime + GPU drivertoolkit hook, device filesGPU sharing boundarywhole GPU, MIG slice, MPS, time-sliceBlue: request path. Violet: serving. Yellow: inputs the server trusts. Red: isolation the tenant relies on.Prompt-level testing covers the top row; many of the worst findings live below it.
The layers of a typical self-hosted stack. The server trusts everything below the request path, so a thorough assessment walks each layer.

Map the target onto this picture first. The gateway may be an API-management product, an Envoy or NGINX proxy, or nothing at all. The model server is usually vLLM, NVIDIA Triton, Ollama, TGI or a KServe wrapper around one of them. Behind it sit the model store, the internal fabric, the orchestrator, and underneath everything the container runtime, GPU driver and the mechanism that shares a physical GPU. Note which layers are reachable from where: the most common serious finding is that a port meant for the private network, such as a metrics or cluster port, answers from somewhere it should not.

Attack surface: ports and who can reach them

Start with exposure, because it is cheap to check and it is where the real incidents happen. Inventory every listening port on the serving hosts and ask, for each, who is supposed to reach it and who actually can. Model servers expose more than the inference API: an OpenAI-compatible endpoint, a Prometheus metrics port, sometimes a health or admin route, and in distributed mode internal sockets.

  • The inference API. Many servers ship with no auth and rely on the operator to put a gateway in front. Confirm the gateway cannot be bypassed by reaching the server directly on the pod or node IP.
  • Metrics and health. A metrics endpoint leaks model names, request volumes and sometimes prompt lengths; it belongs on an internal interface, not the public one.
  • Distributed sockets. Multi-GPU and multi-host serving open sockets for weight and KV-cache transfer that trust their peers; they must never be reachable from untrusted networks.
  • Cluster and job endpoints. An unauthenticated Ray dashboard, Kubernetes API or notebook server on a GPU node is a direct path to code execution.

The historical record is blunt. Anyscale, which maintains Ray, calls unauthenticated job submission (tracked as CVE-2023-48022) expected behaviour for a cluster meant for a trusted network rather than a vulnerability; the researchers who reported it disagreed, and Oligo later found exposed Ray clusters exploited in the wild in a campaign it named ShadowRay. Either way, treat any cluster control surface as code execution for whoever can reach it, and verify network policy rather than trusting a default.

The model is executable: pickle and remote code

The model-loading path is the layer teams most often forget is attacker-controlled. If your platform lets users upload or select arbitrary models, or pulls them from a public hub, you are running their files. Two mechanisms turn a model into code.

A PyTorch checkpoint saved with torch.save is a Python pickle, and unpickling runs constructors chosen by whoever wrote the file, so loading an untrusted .bin or .pt checkpoint runs its author's code. The safetensors format avoids this: it stores only tensors, with no executable payload. The test is a configuration audit, not an exploit: enumerate how models enter the platform and check that untrusted ones must be safetensors, scanned before load, or sandboxed.

The second mechanism is trust_remote_code=True, which lets a Hugging Face repository ship custom Python that the loader imports and runs. It is sometimes needed for a new architecture, but it means the repo author runs code in your process. Grep your serving code and configs for it; every occurrence on an untrusted model path is a finding.

# Audit, not exploit: find unsafe model-loading patterns in a serving codebase.
import pathlib, re

PATTERNS = {
    "trust_remote_code=True": "executes arbitrary repo code on load",
    r"torch\.load\((?![^)]*weights_only\s*=\s*True)": "pickle load without weights_only=True",
    r"pickle\.load": "raw pickle deserialization",
    r"\.bin['\"]|\.pt['\"]|\.ckpt['\"]": "pickle-format checkpoint accepted",
}

def audit(root):
    for path in pathlib.Path(root).rglob("*.py"):
        text = path.read_text(encoding="utf-8", errors="replace")
        for pat, why in PATTERNS.items():
            for m in re.finditer(pat, text):
                line = text[:m.start()].count("\n") + 1
                print(f"{path}:{line}: {why}")

Since PyTorch 2.6, torch.load defaults to weights_only=True, which blocks the arbitrary-pickle path for code that does not override it; confirm the serving stack is on a recent PyTorch and has not set it back to False for untrusted models.

Known-vulnerability review

Known-vulnerability review is the most productive single activity in an AI-infrastructure assessment, because the stack is young and exposed deployments lag patches. Enumerate exact versions of the serving engine, container toolkit, orchestrator and driver, and compare them against the public record. The table lists representative, already-fixed issues to anchor the review; always check the vendor advisory for the current fixed version rather than treating these as complete.

ComponentRepresentative issueFixed in
vLLM (Mooncake KV transfer)CVE-2025-32444: unsafe pickle over an exposed ZeroMQ socket, unauthenticated RCE (CVSS 10.0)0.8.5
vLLM (V0 multi-host TP)CVE-2025-30165: pickle over a ZeroMQ SUB socket; V0 engine, off by default since 0.8.0not fixed; V1 unaffected
NVIDIA Triton (Python backend)CVE-2025-23319 chain: info leak escalating toward RCE25.07
OllamaCVE-2024-37032 (Probllama): path traversal via a crafted manifest digest, RCE0.1.34
NVIDIA Container ToolkitCVE-2024-0132: TOCTOU container escape to host filesystem (CVSS 9.0)1.16.2
NVIDIA Container ToolkitCVE-2025-23266 (NVIDIAScape): OCI-hook escape to host root1.17.8 (GPU Operator 25.3.1)

Two themes repeat. First, pickle over an unauthenticated socket: vLLM's distributed features serialised Python objects with pickle and, for Mooncake, listened on all interfaces, so anyone who reached the socket could run code; the fix replaced pickle with safetensors. Confirm the engine is current and its distributed sockets are bound to a private interface behind network policy. Second, container-escape bugs in the NVIDIA Container Toolkit let a crafted image cross to the host, which on a shared GPU node compromises every tenant, so the toolkit version is among the highest-value things to check.

Multi-tenant GPU isolation

If the platform is multi-tenant, test the sharing boundary explicitly, because the guarantees differ by mechanism. A whole GPU per tenant is the strongest and simplest. Multi-Instance GPU (MIG) partitions one card into instances with hardware-enforced separation of memory and compute, the right choice when different tenants must share one physical GPU. Time-slicing and MPS share a GPU for efficiency but give little isolation: fine inside one trust domain, not across tenants. Record which mechanism is in use and whether it matches the trust model the platform claims.

Memory hygiene is part of this. The LeftoverLocals research (CVE-2023-4969) showed uninitialised GPU local memory letting one kernel read another's leftover data on several vendors' GPUs (Apple, AMD, Qualcomm, Imagination); NVIDIA was not among the affected vendors. The lesson holds regardless of vendor: on shared hardware, treat GPU memory as a side channel worth testing, and keep drivers current.

A repeatable exposure harness

Make the assessment repeatable. A thin harness that records versions and probes exposure turns a one-off test into a control the platform team can run in CI. It detects and evidences rather than exploits: it reads banners and checks reachability, and it belongs to the platform owner.

import socket, json, urllib.request

def port_open(host, port, timeout=2.0):
    with socket.socket() as s:
        s.settimeout(timeout)
        return s.connect_ex((host, port)) == 0

def http_banner(url, timeout=3.0):
    try:
        req = urllib.request.Request(url, headers={"User-Agent": "authorised-assessment"})
        with urllib.request.urlopen(req, timeout=timeout) as r:
            return r.status, r.read(400).decode("utf-8", "replace")
    except Exception as e:
        return None, str(e)

def assess(host, expected_private_ports):
    report = {"host": host, "findings": []}
    for port in expected_private_ports:
        if port_open(host, port):
            report["findings"].append(
                {"port": port, "severity": "high",
                 "note": "internal port reachable from the test vantage point"})
    status, body = http_banner(f"http://{host}:8000/v1/models")
    if status == 200:
        report["findings"].append(
            {"port": 8000, "severity": "high",
             "note": "inference API answered without authentication", "sample": body[:200]})
    print(json.dumps(report, indent=2))
    return report

Run it from both vantage points: from the public internet to prove internet exposure, then from inside the cluster network to map what a foothold would reach. A port firewalled from outside but wide open to every pod is still a finding, because one compromised workload then reaches it.

Worked example: triaging one platform's findings

Rank findings by blast radius and reachability, not by raw CVSS. Suppose one engagement surfaced three: the inference API on port 8000 answering from the public internet with no authentication; NVIDIA Container Toolkit 1.17.4 on the shared GPU nodes, below the 1.17.8 that fixes NVIDIAScape (CVE-2025-23266); and torch.load used on checkpoints that today come only from an internal, trusted bucket.

FindingImpact and reachSeverityFix
Unauthenticated inference API on the internetAnyone: free inference, context and prompt leakageCriticalAuth at a gateway the server cannot be reached around
Container Toolkit 1.17.4 on shared nodesAny tenant: escape to host root, all tenants compromisedCriticalUpgrade to 1.17.8+
torch.load on internal checkpointsRCE only if an untrusted model is ever loadedLow, latentSet weights_only=True; require safetensors before uploads

The ranking is context, not score. The first two are critical because they are reachable and the blast radius is the whole platform; the third is low only because the input is trusted today, logged with a note that it turns critical the day the platform accepts user-supplied models. Retest after each fix: a re-scan should get a 401 on port 8000, nvidia-ctk --version should report 1.17.8 or later, and the audit script should return clean. A finding is not closed until a retest shows it gone.

Failure modes and pitfalls

  • Testing only the prompt. The critical findings are usually an exposed port or an unpatched dependency. Budget time for the infrastructure layers.
  • Scanning from one vantage point. A service firewalled from the internet can be wide open on the pod network. Test from outside and inside.
  • Running destructive checks on shared GPUs. A fuzzing run or real exploit can take down the neighbours' jobs. Prefer version and configuration checks; demonstrate rather than exercise.
  • Trusting default auth. Several serving engines and cluster tools assume a trusted network and ship with no authentication. Verify network policy; never assume a default is safe.
  • Letting the report rot. The stack changes weekly. Hand over the harness so exposure and version checks run in CI.

What to do next

  1. Get written authorisation and scope, including whether tenant-isolation testing is in bounds and in which window.
  2. Inventory listening ports on every serving host and test reachability from both the internet and the internal network.
  3. Enumerate exact versions of the serving engine, container toolkit, orchestrator and driver against current vendor advisories.
  4. Audit the model-loading path for pickle checkpoints and trust_remote_code on any untrusted model source.
  5. If the platform is multi-tenant, confirm the sharing mechanism matches the trust model; see MIG and GPU-enabled containers.
  6. Write findings as reachability, impact, evidence and a versioned fix; retest and record closure.
  7. Hand the exposure-and-version harness to the platform team to run in CI, and pair it with the application-layer tests in LLM pentesting in depth.
Key takeaway: On a self-hosted GPU inference platform the prompt is the least privileged way in. Scope the engagement in writing, then walk the layers below it: exposed inference, metrics and cluster ports; distributed sockets that trust their peers; pickle and trust_remote_code loading of attacker-supplied models; unpatched serving engines and container toolkits with public RCE and container-escape CVEs; and the GPU sharing boundary. Check versions and configuration rather than exploiting, evidence every finding, retest after the fix, and leave the exposure harness running in CI.