Most writing about LLM security is about prompts: injection, jailbreaks, data leaking through outputs. Those are real, but several of the most damaging incidents around language models so far have had nothing to do with prompts. They were ordinary infrastructure failures in the systems that host models: an unauthenticated job API on a Ray cluster, an inference engine deserialising pickles from any host on the network, a container toolkit that could be tricked into mounting the host filesystem, a model server that wrote attacker-controlled paths while pulling a model.

This article covers the serving plane: model weights and where they live, the inference and management APIs, the internal ports that serving stacks open, GPU nodes, and the orchestration and operator access around them. General container and network hardening is covered in LLM deployment hardening and protecting prompts from the operator with confidential computing in secure inference. The goal here is to know what to protect, where the doors are, and which ones have actually been used.

Advertisement

What is worth stealing or breaking

Start with an asset list, because it decides priorities. Weights and adapters: a fine-tuned model can represent months of GPU time and proprietary data, and a stolen copy can be served, studied for its training data or probed offline for jailbreaks. Prompts and outputs: they sit in GPU memory, KV caches, request logs and traces, and often contain customer data. Credentials: serving hosts carry cloud credentials, registry tokens and API keys. GPU time: attackers who gain code execution on GPU clusters routinely install cryptocurrency miners, which is the cheapest attack to monetise and a common first sign of compromise. Integrity: an attacker who can swap weights or an adapter can plant a backdoor that passes normal evaluation.

The threat actors are the usual ones: internet scanners looking for known ports, insiders or compromised operator accounts, malicious or compromised model publishers, and other tenants sharing hardware.

The LLM serving plane: assets, planes and the ports attackers look forClients / appsvia API gatewayAPI gatewayauthn, quotas, loggingInference serversvLLM, Triton, Ollama, TGIInternal portsKV transfer, metrics, adminOrchestrationKubernetes, Ray, SlurmModel registryweights, adapters, hashesGPU nodesdriver, container toolkitSecrets / KMSkeys, tokensOperators / CIexec, deploy, pullonly path inloadRBACAssets: weights and adapters, prompts and outputs in memory and logs, KV caches, credentials, GPU time.Most real incidents came through the middle column: an internal or admin port reachable from somewhere it should not be.Everything except the gateway should be unreachable from outside the serving network.
The serving plane. The API gateway should be the only path in; internal ports, orchestration dashboards and admin APIs are where most real incidents started.

Weights: integrity, format and access

Model files are code-adjacent. The PyTorch pickle format executes arbitrary Python when loaded, which is why the safetensors format, a JSON header plus raw tensor bytes, has become the default for published models. Since PyTorch 2.6, torch.load defaults to weights_only=True, which restricts unpickling to tensors and simple types, but older code, third-party loaders and trust_remote_code paths in model repositories can still execute code. Treat any model that requires custom code as software to review, not data to load.

Integrity checking belongs in the loader, not in a README. Record a SHA-256 for every file when a model is approved, sign that manifest with your existing artifact-signing tooling, verify the signature in the deploy pipeline, and have the serving process refuse to start if any file differs or an unexpected file appears:

import hashlib, json, pathlib
from safetensors.torch import load_file

def sha256(path, chunk=1 << 24):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        while block := f.read(chunk):
            h.update(block)
    return h.hexdigest()

def load_verified(model_dir, manifest_path):
    """manifest: {"files": {"model-00001-of-00004.safetensors": "<sha256>", ...}}
    The manifest itself is signed and verified by the deploy pipeline before this runs."""
    manifest = json.loads(pathlib.Path(manifest_path).read_text())
    model_dir = pathlib.Path(model_dir)
    present = {p.name for p in model_dir.iterdir() if p.is_file()}
    expected = set(manifest["files"])
    if present - expected:
        raise RuntimeError(f"unexpected files in model dir: {sorted(present - expected)}")
    for name, digest in manifest["files"].items():
        if not name.endswith((".safetensors", ".json", ".model", ".txt")):
            raise RuntimeError(f"refusing non-data format: {name}")   # no .bin, .pt, .pkl
        if sha256(model_dir / name) != digest:
            raise RuntimeError(f"hash mismatch: {name}")
    return {n: load_file(model_dir / n) for n in manifest["files"] if n.endswith(".safetensors")}

For confidentiality, store weights encrypted with a key held in a KMS, grant decrypt permission only to the serving identity, and log every access. Alert on reads of the weight bucket from any other principal and on large egress from serving nodes. A pod that can read the bucket and reach the internet can exfiltrate a model in minutes, which is why the network rules below matter as much as the storage policy.

Advertisement

Inference and model-management APIs

Inference servers are built for throughput first. Many ship with no authentication or an optional static key; vLLM's OpenAI-compatible server, for example, accepts an --api-key option but serves without one by default. Ollama listens on localhost by default, and people change that to share a GPU box, exposing an API with no authentication at all. Treat every inference server as an internal component: put it behind a gateway that authenticates callers, enforces quotas and logs requests, and make the gateway the only thing that can reach it.

Model-management endpoints deserve extra care because they write to disk and change what runs. CVE-2024-37032, nicknamed Probllama, was a path traversal in Ollama's model pull endpoint: insufficient validation of the digest in a model manifest let a malicious registry write files outside the model directory, leading to remote code execution. It was fixed in version 0.1.34 in May 2024, and the lesson is general. Pull, load, unload and delete APIs should be disabled in production or exposed only to the deploy pipeline on a separate listener. In Triton, for example, keep the model control mode at its default of none so models cannot be loaded through the API.

Internal ports: where the big incidents happened

Distributed inference opens ports beyond the API: tensor-parallel communication, disaggregated prefill and KV-cache transfer between nodes, metrics, dashboards and job submission. These are designed for trusted networks and frequently have no authentication.

CaseWhat was exposedLesson
CVE-2023-48022, Ray (ShadowRay)Ray's dashboard and Jobs API accept job submissions without authentication, which is remote code execution by designAnyscale disputes it as a vulnerability, saying Ray must not be exposed to untrusted networks. Scanners do not care: campaigns have used it to run miners and steal credentials from exposed clusters
CVE-2025-32444, vLLM mooncake integrationvLLM 0.6.5 up to 0.8.5 deserialised pickles received over ZeroMQ sockets that listened on all interfacesAny host that could reach the socket got code execution. Fixed in 0.8.5; only deployments using that integration were affected

Both failures have the same shape: a component meant for a private network reachable from a wider one. The fix is structural. Run serving and orchestration in a dedicated network segment, bind internal listeners to specific interfaces rather than 0.0.0.0, and enforce the segment with default-deny policies that allow only the traffic you can name:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: {name: inference-lockdown, namespace: serving}
spec:
  podSelector: {matchLabels: {app: vllm}}
  policyTypes: [Ingress, Egress]
  ingress:
  - from: [{podSelector: {matchLabels: {app: api-gateway}}}]
    ports: [{port: 8000, protocol: TCP}]          # OpenAI-compatible API only
  - from: [{namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: monitoring}}}]
    ports: [{port: 9090, protocol: TCP}]          # metrics, if served separately
  egress:
  - to: [{namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: kube-system}}}]
    ports: [{port: 53, protocol: UDP}, {port: 53, protocol: TCP}]   # DNS only; weights come from a volume

Note the egress rule. Serving pods do not need the internet; weights arrive on a volume prepared by the deploy pipeline. Denying egress turns a remote code execution from a data-exfiltration event into a contained one, and it is covered further in egress filtering.

GPU nodes, drivers and sharing

GPU containers depend on host components: the kernel driver and a container runtime hook that mounts driver libraries and device files into the container. That hook runs with host privileges, which makes it a boundary. CVE-2024-0132 was a time-of-check to time-of-use race in NVIDIA Container Toolkit 1.16.1 and earlier: a crafted image could make the toolkit mount the host's root filesystem into the container, a full escape. It was fixed in toolkit 1.16.2 and GPU Operator 24.6.2, and deployments using the Container Device Interface (CDI) were not affected. The operational lessons are to patch the GPU software stack on the same schedule as the kernel, and never to run untrusted images on nodes that also run sensitive models.

Sharing a GPU is also a trust decision. NVIDIA's Multi-Instance GPU partitions memory and compute into isolated instances in hardware. Time-slicing and MPS interleave or co-schedule processes without that isolation, so they are for workloads that trust each other. Memory residue is the other risk: LeftoverLocals (CVE-2023-4969) showed that on affected Apple, AMD, Qualcomm and Imagination GPUs one process could read GPU local memory left behind by another, enough to reconstruct a co-resident LLM's responses. NVIDIA GPUs were not affected by that bug, but the category is real, so separate tenants by node or MIG instance rather than relying on the driver to clear state. Tenant boundaries above the hardware are covered in multi-tenant LLM isolation.

Orchestration and operator access

Anyone who can exec into a serving pod can read weights from disk or memory and see prompts in flight. In Kubernetes, pods/exec, pods/attach, creating pods in the serving namespace and reading its Secrets are all equivalent to model access, so grant them through just-in-time elevation with an audit trail, not standing roles. Notebook servers and training clusters are the usual soft spots: they mount the same buckets, carry broad credentials and are often reachable from office networks.

Keep credentials short-lived and scoped: workload identity instead of key files, registry tokens that can pull but not push, and a separate identity for the pipeline that promotes models. Detailed patterns for keeping credentials away from the model itself are in secrets management for LLM applications.

Detection signals

  • Sustained GPU utilisation with no matching request volume at the gateway: the classic miner signature.
  • New jobs on a Ray, Slurm or Kubernetes cluster that no pipeline submitted.
  • Reads of the weight bucket or registry by any identity other than the serving and deploy principals.
  • Egress from serving nodes to anything outside the allow-list, especially large transfers.
  • Model file hashes that differ from the approved manifest, checked at load time and periodically.
  • Exec sessions into serving pods, and new listeners on unexpected ports found by internal scans.

A worked assessment

A team runs a 70B model on vLLM across two nodes, orchestrated with Ray, behind an internal load balancer. A one-day review walks the diagram above and finds four issues. The Ray dashboard on port 8265 is reachable from the corporate network, so anyone on the VPN can submit jobs: move it behind an authenticated proxy and restrict it to the operators' bastion. The vLLM API has no key and the load balancer is reachable from every internal subnet: route traffic through the gateway and add a NetworkPolicy so only the gateway can connect. Weights are pulled at startup from a public hub with trust_remote_code enabled: mirror the approved revision into the internal registry, pin hashes and disable remote code. Finally, the nodes use a shared cloud credential that can write the model bucket: split read and write identities.

None of these fixes involves the model's behaviour, and together they close every path used in the incidents above.

Trade-offs

Default-deny networking and hash-pinned weights slow down experimentation, so keep them strict in production and looser in a separate research environment that holds no production data or credentials. MIG isolation costs flexibility and some utilisation compared with time-slicing. Encrypting weights and gating decryption adds startup latency and a KMS dependency to every scale-out. These costs are small compared with a leaked model or a compromised cluster, but they are real, so make the exceptions explicit rather than letting them accumulate.

What to do next

  1. Draw your serving plane on one page and list every listening port on every component, including metrics, dashboards and distributed-communication ports.
  2. Scan from outside the serving segment and confirm that only the gateway answers.
  3. Put every inference server behind an authenticating gateway and disable model-management APIs in production.
  4. Convert models to safetensors, record and sign hashes, and make the loader refuse mismatches.
  5. Check versions of vLLM, Ray, Ollama, the NVIDIA Container Toolkit and GPU Operator against the fixes named above.
  6. Deny egress from serving pods and alert on GPU utilisation without matching traffic.
  7. Remove standing exec access to serving namespaces and split the identities that read and write weights.
Key takeaway: LLM infrastructure is attacked like any other infrastructure, and the damaging incidents so far came through exposed internal components rather than prompts: an unauthenticated Ray job API, pickle deserialisation on an open vLLM socket, a container-toolkit escape and a path traversal in a model pull API. Put every inference server behind an authenticating gateway, keep internal and management ports inside a default-deny network segment, load only hash-verified safetensors, patch the GPU software stack promptly, isolate tenants by node or MIG instance, and treat exec access to serving pods as access to the model itself.