Most writing on LLM penetration testing is about prompts: injection, jailbreaks, data pulled out through the model's answers. That work matters, and LLM pentesting in depth covers how to run it against a single feature. But when a team hosts its own models on its own GPUs, the prompt is the least privileged way in. The model server is a network service written in Python and C++; it loads files that can contain executable code; it talks to peers over internal sockets; and it runs in containers that need privileged access to a GPU driver. Every one of those layers has produced critical, publicly disclosed vulnerabilities since 2023.
This article is a defensive test plan for that stack: how to scope an authorised assessment of a self-hosted inference platform, what to examine at each layer, how to check for exposure without causing harm, and how to report findings so the platform team can fix them. It assumes written authorisation from the system owner and a defined window. Everything here is a check against infrastructure you are permitted to test; none of it is an exploit, and the version numbers below are there so you can confirm a system is already patched.
Why GPU inference is a different target
Three properties make GPU inference infrastructure different from an ordinary web service. First, the software is young and moves fast: serving engines ship every few weeks and add distributed features such as disaggregated prefill and KV-cache transfer, many first designed for trusted research clusters. Second, models are code. A checkpoint saved in Python's pickle format can run arbitrary code when loaded, and model repositories can ask the loader to execute custom Python through a trust_remote_code flag. Third, GPUs are expensive, so they are shared between teams, tenants and jobs; the isolation that sharing relies on is weaker and less understood than CPU virtualisation.
The consequence is that one misconfiguration can be worth more to an attacker than any prompt. An unauthenticated model API reachable from the internet gives away inference at your expense and leaks whatever is in its context. An unauthenticated cluster or job endpoint can mean code execution on GPU nodes, which usually hold cloud credentials, weights and training data; a container escape on a shared node gives the host and every tenant on it. A good assessment spends most of its time below the prompt.
Authorisation and rules of engagement, first
Before any traffic, get the scope in writing. Name the exact hosts, clusters, namespaces and model endpoints in scope, and what is explicitly out of scope, such as the cloud provider's control plane. Agree a window, a rate limit, and a named contact who can stop the test, and decide how findings that touch live customer data are handled before you find one. For a shared GPU platform, settle one question early: are you testing as an outside attacker, or as a hostile tenant who already has a legitimate account? The two produce very different findings, and tenant-isolation testing can affect the neighbours, so it usually needs its own window and sign-off.
Capture non-destructive evidence: request and response pairs, version banners, configuration snippets. Agree to demonstrate access rather than exercise it, for example by reading a canary file placed for the test rather than real data. The deliverable is a report the platform team can act on, not a trophy.
Map the stack before you test it
Map the target onto this picture first. The gateway may be an API-management product, an Envoy or NGINX proxy, or nothing at all. The model server is usually vLLM, NVIDIA Triton, Ollama, TGI or a KServe wrapper around one of them. Behind it sit the model store, the internal fabric, the orchestrator, and underneath everything the container runtime, GPU driver and the mechanism that shares a physical GPU. Note which layers are reachable from where: the most common serious finding is that a port meant for the private network, such as a metrics or cluster port, answers from somewhere it should not.
Attack surface: ports and who can reach them
Start with exposure, because it is cheap to check and it is where the real incidents happen. Inventory every listening port on the serving hosts and ask, for each, who is supposed to reach it and who actually can. Model servers expose more than the inference API: an OpenAI-compatible endpoint, a Prometheus metrics port, sometimes a health or admin route, and in distributed mode internal sockets.
- The inference API. Many servers ship with no auth and rely on the operator to put a gateway in front. Confirm the gateway cannot be bypassed by reaching the server directly on the pod or node IP.
- Metrics and health. A metrics endpoint leaks model names, request volumes and sometimes prompt lengths; it belongs on an internal interface, not the public one.
- Distributed sockets. Multi-GPU and multi-host serving open sockets for weight and KV-cache transfer that trust their peers; they must never be reachable from untrusted networks.
- Cluster and job endpoints. An unauthenticated Ray dashboard, Kubernetes API or notebook server on a GPU node is a direct path to code execution.
The historical record is blunt. Anyscale, which maintains Ray, calls unauthenticated job submission (tracked as CVE-2023-48022) expected behaviour for a cluster meant for a trusted network rather than a vulnerability; the researchers who reported it disagreed, and Oligo later found exposed Ray clusters exploited in the wild in a campaign it named ShadowRay. Either way, treat any cluster control surface as code execution for whoever can reach it, and verify network policy rather than trusting a default.
The model is executable: pickle and remote code
The model-loading path is the layer teams most often forget is attacker-controlled. If your platform lets users upload or select arbitrary models, or pulls them from a public hub, you are running their files. Two mechanisms turn a model into code.
A PyTorch checkpoint saved with torch.save is a Python pickle, and unpickling runs constructors chosen by whoever wrote the file, so loading an untrusted .bin or .pt checkpoint runs its author's code. The safetensors format avoids this: it stores only tensors, with no executable payload. The test is a configuration audit, not an exploit: enumerate how models enter the platform and check that untrusted ones must be safetensors, scanned before load, or sandboxed.
The second mechanism is trust_remote_code=True, which lets a Hugging Face repository ship custom Python that the loader imports and runs. It is sometimes needed for a new architecture, but it means the repo author runs code in your process. Grep your serving code and configs for it; every occurrence on an untrusted model path is a finding.
# Audit, not exploit: find unsafe model-loading patterns in a serving codebase.
import pathlib, re
PATTERNS = {
"trust_remote_code=True": "executes arbitrary repo code on load",
r"torch\.load\((?![^)]*weights_only\s*=\s*True)": "pickle load without weights_only=True",
r"pickle\.load": "raw pickle deserialization",
r"\.bin['\"]|\.pt['\"]|\.ckpt['\"]": "pickle-format checkpoint accepted",
}
def audit(root):
for path in pathlib.Path(root).rglob("*.py"):
text = path.read_text(encoding="utf-8", errors="replace")
for pat, why in PATTERNS.items():
for m in re.finditer(pat, text):
line = text[:m.start()].count("\n") + 1
print(f"{path}:{line}: {why}")Since PyTorch 2.6, torch.load defaults to weights_only=True, which blocks the arbitrary-pickle path for code that does not override it; confirm the serving stack is on a recent PyTorch and has not set it back to False for untrusted models.
Known-vulnerability review
Known-vulnerability review is the most productive single activity in an AI-infrastructure assessment, because the stack is young and exposed deployments lag patches. Enumerate exact versions of the serving engine, container toolkit, orchestrator and driver, and compare them against the public record. The table lists representative, already-fixed issues to anchor the review; always check the vendor advisory for the current fixed version rather than treating these as complete.
| Component | Representative issue | Fixed in |
|---|---|---|
| vLLM (Mooncake KV transfer) | CVE-2025-32444: unsafe pickle over an exposed ZeroMQ socket, unauthenticated RCE (CVSS 10.0) | 0.8.5 |
| vLLM (V0 multi-host TP) | CVE-2025-30165: pickle over a ZeroMQ SUB socket; V0 engine, off by default since 0.8.0 | not fixed; V1 unaffected |
| NVIDIA Triton (Python backend) | CVE-2025-23319 chain: info leak escalating toward RCE | 25.07 |
| Ollama | CVE-2024-37032 (Probllama): path traversal via a crafted manifest digest, RCE | 0.1.34 |
| NVIDIA Container Toolkit | CVE-2024-0132: TOCTOU container escape to host filesystem (CVSS 9.0) | 1.16.2 |
| NVIDIA Container Toolkit | CVE-2025-23266 (NVIDIAScape): OCI-hook escape to host root | 1.17.8 (GPU Operator 25.3.1) |
Two themes repeat. First, pickle over an unauthenticated socket: vLLM's distributed features serialised Python objects with pickle and, for Mooncake, listened on all interfaces, so anyone who reached the socket could run code; the fix replaced pickle with safetensors. Confirm the engine is current and its distributed sockets are bound to a private interface behind network policy. Second, container-escape bugs in the NVIDIA Container Toolkit let a crafted image cross to the host, which on a shared GPU node compromises every tenant, so the toolkit version is among the highest-value things to check.
Multi-tenant GPU isolation
If the platform is multi-tenant, test the sharing boundary explicitly, because the guarantees differ by mechanism. A whole GPU per tenant is the strongest and simplest. Multi-Instance GPU (MIG) partitions one card into instances with hardware-enforced separation of memory and compute, the right choice when different tenants must share one physical GPU. Time-slicing and MPS share a GPU for efficiency but give little isolation: fine inside one trust domain, not across tenants. Record which mechanism is in use and whether it matches the trust model the platform claims.
Memory hygiene is part of this. The LeftoverLocals research (CVE-2023-4969) showed uninitialised GPU local memory letting one kernel read another's leftover data on several vendors' GPUs (Apple, AMD, Qualcomm, Imagination); NVIDIA was not among the affected vendors. The lesson holds regardless of vendor: on shared hardware, treat GPU memory as a side channel worth testing, and keep drivers current.
A repeatable exposure harness
Make the assessment repeatable. A thin harness that records versions and probes exposure turns a one-off test into a control the platform team can run in CI. It detects and evidences rather than exploits: it reads banners and checks reachability, and it belongs to the platform owner.
import socket, json, urllib.request
def port_open(host, port, timeout=2.0):
with socket.socket() as s:
s.settimeout(timeout)
return s.connect_ex((host, port)) == 0
def http_banner(url, timeout=3.0):
try:
req = urllib.request.Request(url, headers={"User-Agent": "authorised-assessment"})
with urllib.request.urlopen(req, timeout=timeout) as r:
return r.status, r.read(400).decode("utf-8", "replace")
except Exception as e:
return None, str(e)
def assess(host, expected_private_ports):
report = {"host": host, "findings": []}
for port in expected_private_ports:
if port_open(host, port):
report["findings"].append(
{"port": port, "severity": "high",
"note": "internal port reachable from the test vantage point"})
status, body = http_banner(f"http://{host}:8000/v1/models")
if status == 200:
report["findings"].append(
{"port": 8000, "severity": "high",
"note": "inference API answered without authentication", "sample": body[:200]})
print(json.dumps(report, indent=2))
return reportRun it from both vantage points: from the public internet to prove internet exposure, then from inside the cluster network to map what a foothold would reach. A port firewalled from outside but wide open to every pod is still a finding, because one compromised workload then reaches it.
Worked example: triaging one platform's findings
Rank findings by blast radius and reachability, not by raw CVSS. Suppose one engagement surfaced three: the inference API on port 8000 answering from the public internet with no authentication; NVIDIA Container Toolkit 1.17.4 on the shared GPU nodes, below the 1.17.8 that fixes NVIDIAScape (CVE-2025-23266); and torch.load used on checkpoints that today come only from an internal, trusted bucket.
| Finding | Impact and reach | Severity | Fix |
|---|---|---|---|
| Unauthenticated inference API on the internet | Anyone: free inference, context and prompt leakage | Critical | Auth at a gateway the server cannot be reached around |
| Container Toolkit 1.17.4 on shared nodes | Any tenant: escape to host root, all tenants compromised | Critical | Upgrade to 1.17.8+ |
torch.load on internal checkpoints | RCE only if an untrusted model is ever loaded | Low, latent | Set weights_only=True; require safetensors before uploads |
The ranking is context, not score. The first two are critical because they are reachable and the blast radius is the whole platform; the third is low only because the input is trusted today, logged with a note that it turns critical the day the platform accepts user-supplied models. Retest after each fix: a re-scan should get a 401 on port 8000, nvidia-ctk --version should report 1.17.8 or later, and the audit script should return clean. A finding is not closed until a retest shows it gone.
Failure modes and pitfalls
- Testing only the prompt. The critical findings are usually an exposed port or an unpatched dependency. Budget time for the infrastructure layers.
- Scanning from one vantage point. A service firewalled from the internet can be wide open on the pod network. Test from outside and inside.
- Running destructive checks on shared GPUs. A fuzzing run or real exploit can take down the neighbours' jobs. Prefer version and configuration checks; demonstrate rather than exercise.
- Trusting default auth. Several serving engines and cluster tools assume a trusted network and ship with no authentication. Verify network policy; never assume a default is safe.
- Letting the report rot. The stack changes weekly. Hand over the harness so exposure and version checks run in CI.
What to do next
- Get written authorisation and scope, including whether tenant-isolation testing is in bounds and in which window.
- Inventory listening ports on every serving host and test reachability from both the internet and the internal network.
- Enumerate exact versions of the serving engine, container toolkit, orchestrator and driver against current vendor advisories.
- Audit the model-loading path for pickle checkpoints and
trust_remote_codeon any untrusted model source. - If the platform is multi-tenant, confirm the sharing mechanism matches the trust model; see MIG and GPU-enabled containers.
- Write findings as reachability, impact, evidence and a versioned fix; retest and record closure.
- Hand the exposure-and-version harness to the platform team to run in CI, and pair it with the application-layer tests in LLM pentesting in depth.