Ray Serve is the model-serving library that runs on a Ray cluster. Its LLM layer, imported as ray.serve.llm, wraps an inference engine (vLLM by default) in Serve deployments and puts an OpenAI-compatible HTTP API in front of them. You describe a model with one LLMConfig object; Ray turns it into replicas, finds GPUs for each replica, starts the engine on them, and adds or removes replicas as traffic changes.

This page is about the GPU side of that story: how one config line becomes a set of GPU reservations, how to size a replica so its KV cache is big enough to be useful, when tensor parallelism beats more replicas, how many LoRA adapters share one base model, and which mistakes leave a cluster with idle GPUs and pending replicas. Routing policy and autoscaling behaviour under load are covered in Ray Serve for LLM systems: routing, autoscaling and failure recovery; general Ray scheduling of tasks and actors is in Ray for GPU workloads. Field names below were checked against the Ray documentation (docs.ray.io, Serve LLM configuration reference) in October 2026. The API is still moving between Ray releases, so pin a version and confirm against it.

How a request reaches a GPU

Ray Serve LLM: one LLMConfig becomes a deployment, each replica owns a placement group of GPUsclientsOpenAI SDK / curlHTTP proxyone per nodeOpenAiIngressmodel id to deploymentcontrollerautoscalerLLMServer replica 1vLLM engine, request router picks itLLMServer replica 2same LLMConfig, own KV cacherouteroutebundle: 1 GPUTP rank 0bundle: 1 GPUTP rank 1bundle: 1 GPUTP rank 0bundle: 1 GPUTP rank 1node A: STRICT_PACK keeps both ranks on one hostnode B: second placement group, second copy of weightsscaleBundles per replica = tensor_parallel_size x pipeline_parallel_size; a replica starts only when every bundle is placed.
Request path and GPU layout for one model served with tensor parallelism of 2 and two replicas. Each replica holds its own copy of the weights and its own KV cache.

Four kinds of process are involved when a request arrives.

  • HTTP proxy. Ray Serve runs a proxy actor on each node (by default) that accepts HTTP and forwards to the ingress deployment.
  • OpenAiIngress. The ingress deployment implements the OpenAI routes such as /v1/chat/completions and /v1/models. It reads the model field of the request and forwards to the deployment that serves that model id.
  • LLMServer replicas. One Serve deployment per LLMConfig. Each replica is a Ray actor that owns one engine instance. With vLLM, the engine runs continuous batching and a paged KV cache inside that replica; Ray does not see individual tokens.
  • Engine workers. When the model is split with tensor or pipeline parallelism, the engine starts one worker per GPU. Ray reserves those GPUs as a placement group before the replica starts.

The important consequence is that a replica is the unit of scaling and the unit of failure. Two replicas never share a KV cache or weights. If you need more throughput, you either make each replica bigger (more GPUs through tensor parallelism, more KV memory) or you add replicas (more copies of the weights).

The LLMConfig fields that decide GPU use

An LLMConfig has a handful of fields that matter for GPUs. The table lists them with what each one controls.

FieldWhat it controlsGPU consequence
model_loading_configmodel_id clients use, model_source (Hugging Face id or a path)Download time is part of every cold start
accelerator_typeGPU type such as A10G, L4, A100, H100Replica is placed only on nodes that advertise that type
engine_kwargsArguments passed straight to the vLLM enginetensor_parallel_size, max_model_len, gpu_memory_utilization decide memory
placement_group_configBundles and strategy for engine workersWhich GPUs, and whether they must share a node
deployment_configServe options: autoscaling_config, max_ongoing_requests, health checksHow many replicas, so how many GPUs in total
lora_configWhere adapters live and how many stay loadedAdapter memory per replica
runtime_envEnvironment variables, packagesFor example a Hugging Face token for gated models

Two points trip people up. First, engine_kwargs is not a Ray setting at all: every key is a vLLM engine argument, so the vLLM documentation for your installed version is the reference for names and defaults. Second, accelerator_type works through Ray custom resources. Ray detects the GPU model on each node and advertises a resource named like accelerator_type:A10G; the replica requests a small amount of it. If the node reports a different string than you request, the replica stays pending forever even though GPUs are idle.

Placement groups for tensor and pipeline parallelism

For a model with tensor_parallel_size T and pipeline_parallel_size P, the engine needs T x P workers, one GPU each. Ray reserves them as a placement group: a list of resource bundles that are granted together or not at all. You can give the full list in bundles or give one bundle_per_worker that Ray repeats T x P times, and choose a strategy.

StrategyMeaningUse it for
STRICT_PACKAll bundles on one node, or nothingTensor parallelism: all-reduce on every layer needs NVLink or PCIe, not the network
PACKPrefer one node, spill over if neededSmall models where a spill is tolerable; risky for TP
SPREADPrefer different nodesPipeline stages across nodes, where traffic is one activation per micro-batch
STRICT_SPREADEvery bundle on a different nodeRarely useful for one replica

Tensor parallelism exchanges activations twice per transformer layer, which is why it belongs inside one host on fast links (see tensor parallelism in depth). Pipeline parallelism sends far less traffic, so the usual layout for a model too big for one node is TP equal to the GPUs per node and PP equal to the number of nodes. A placement group that cannot be satisfied does not fail; it waits. A cluster with eight free GPUs spread one per node can never place a TP=2 STRICT_PACK replica, and the only symptom is a replica stuck in a pending state.

Worked example: sizing an 8B model on 24 GB cards

Sizing starts from memory, because the KV cache, not compute, sets how many requests a replica can hold. Take Llama 3.1 8B in bf16 on NVIDIA A10G cards (24 GB each). The model has 32 layers, 8 key-value heads and a head dimension of 128.

  1. Weights. About 8.0 billion parameters x 2 bytes = roughly 16 GB.
  2. KV per token. 2 (K and V) x 32 layers x 8 heads x 128 x 2 bytes = 131,072 bytes, so 128 KiB per token.
  3. Budget. vLLM claims a fraction of GPU memory set by gpu_memory_utilization (0.9 by default in current vLLM). Of about 22 GB usable per A10G, 0.9 gives about 20 GB. Subtract 16 GB of weights and roughly 1.5 GB for activations and CUDA graphs, and about 2.5 GB is left for KV.
  4. Capacity at TP=1. 2.5 GB / 128 KiB is about 19,000 tokens. At 4,000 tokens per conversation that is fewer than 5 concurrent sequences, and a single 32k-token request cannot fit at all, so vLLM refuses to start unless you lower max_model_len.
  5. Capacity at TP=2. Each GPU now holds 8 GB of weights and half of every token's KV. Budget across both is about 40 GB; minus 16 GB of weights and about 3 GB of overhead leaves about 21 GB, or roughly 160,000 tokens: around 40 concurrent 4k conversations.

Two TP=1 replicas would use the same two GPUs and give 2 x 19,000 = 38,000 tokens of cache, less than a quarter of the TP=2 figure, because each replica stores the 16 GB of weights again. On small-memory cards TP=2 wins clearly. On an 80 GB card the picture flips: weights are a small share of memory, a TP=1 replica already has tens of gigabytes of KV, and two independent replicas avoid the all-reduce cost and fail independently. The rule: split a model across GPUs only until the KV cache is large enough for your concurrency target, then scale out with replicas. The token arithmetic is worked further in KV cache size in depth.

Code: Python and YAML deployment

The Python form, with the sizing decision above written into it:

from ray import serve
from ray.serve.llm import LLMConfig, build_openai_app

llm_config = LLMConfig(
    model_loading_config={
        "model_id": "llama-8b",                         # name clients send
        "model_source": "meta-llama/Llama-3.1-8B-Instruct",
    },
    accelerator_type="A10G",
    engine_kwargs={                                     # vLLM engine arguments
        "tensor_parallel_size": 2,
        "max_model_len": 16384,
        "gpu_memory_utilization": 0.90,
    },
    placement_group_config={
        "bundle_per_worker": {"CPU": 1, "GPU": 1},      # repeated TP x PP times
        "strategy": "STRICT_PACK",
    },
    deployment_config={
        "autoscaling_config": {"min_replicas": 1, "max_replicas": 4},
    },
    runtime_env={"env_vars": {"HF_TOKEN": "<set from your secret store>"}},
)

app = build_openai_app({"llm_configs": [llm_config]})
serve.run(app, blocking=True)

The same thing as a Serve config file, which is what you deploy in production with serve deploy or a KubeRay RayService, so that configuration lives in version control:

applications:
  - name: llm_app
    route_prefix: /
    import_path: ray.serve.llm:build_openai_app
    args:
      llm_configs:
        - model_loading_config:
            model_id: llama-8b
            model_source: meta-llama/Llama-3.1-8B-Instruct
          accelerator_type: A10G
          engine_kwargs:
            tensor_parallel_size: 2
            max_model_len: 16384
          deployment_config:
            autoscaling_config:
              min_replicas: 1
              max_replicas: 4

Clients use any OpenAI SDK with base_url pointed at the Serve endpoint and model="llama-8b". Several LLMConfig objects in one list give several models behind one endpoint, each with its own GPU type and replica count.

Many fine-tunes on one base model with LoRA

Fine-tuned variants of one base model rarely need their own replicas. With lora_config set, each replica keeps the base weights once and loads small adapters on demand.

llm_config = LLMConfig(
    model_loading_config={"model_id": "llama-8b",
                          "model_source": "meta-llama/Llama-3.1-8B-Instruct"},
    accelerator_type="L4",
    engine_kwargs={"enable_lora": True, "max_lora_rank": 32},   # vLLM arguments
    lora_config={
        "dynamic_lora_loading_path": "s3://my-bucket/adapters/llama-8b/",
        "max_num_adapters_per_replica": 16,                     # LRU eviction beyond this
    },
)
# client: model="llama-8b:support-tone-v3" selects adapter "support-tone-v3"

A client names an adapter as <base_model_id>:<adapter_name>. The router prefers a replica that already holds that adapter; if that replica is busy it sends the request to a less loaded one, which downloads and loads the adapter first. That is model multiplexing. Each loaded adapter costs GPU memory roughly proportional to its rank and target modules, so 16 resident adapters of rank 32 take memory from the KV cache. The trade-off and the batching kernels behind it are in multi-LoRA serving. Watch the adapter load rate: if it climbs, requests are bouncing between replicas faster than the cache can keep up, and you need either a larger max_num_adapters_per_replica or fewer adapters per base model.

Routing and autoscaling on GPUs

By default Serve routes with a power-of-two-choices policy: sample two replicas, send to the one with fewer requests in flight. For chat or agent traffic that reuses long system prompts, Ray Serve LLM ships a PrefixCacheAffinityRouter that sends requests with a shared prefix to the replica that already has those tokens in vLLM's prefix cache, and falls back to load balancing when replicas become uneven. You enable it through request_router_config inside deployment_config; check the exact import path for your Ray version. Autoscaling uses autoscaling_config (min_replicas, max_replicas, target_ongoing_requests). Two GPU-specific cautions apply. A new replica is not useful until the weights have downloaded and the engine has warmed up, often one to several minutes for an 8B model and far longer for a 70B one, so set downscale delays generously and keep a warm minimum. And an autoscaler that asks for more replicas than there are free GPU groups only creates pending placement groups, so pair it with a cluster autoscaler or cap max_replicas at real capacity.

Failure modes

SymptomLikely causeFix
Replica pending, GPUs idleaccelerator_type string does not match the node, or STRICT_PACK cannot fit on any single nodeRun ray status and compare resource names; free a whole node or reduce TP
Engine exits at start with a KV cache errormax_model_len needs more KV than the budget leavesLower max_model_len, raise TP, or use a bigger card
CUDA out of memory under loadgpu_memory_utilization too high when another process shares the GPUDo not co-locate; lower the fraction
Slow first request after scale-upWeights downloaded from the hub on each new replicaPre-stage weights on local disk or object storage near the cluster
Hang on startup with TP across nodesCollective setup over a slow or misconfigured networkKeep TP inside a node; use PP across nodes
Latency spikes for adapter trafficAdapter thrash: LRU evicts and reloads constantlyRaise the per-replica limit or pin hot adapters to their own deployment

In every case start with ray status (what resources exist and are used) and the Serve dashboard (replica states). A pending replica almost always means a placement group that cannot be satisfied, not a bug in the model code.

Trade-offs: Ray Serve or a plain engine server

Ray Serve LLM is worth its extra layer when you serve several models or many adapters, need Python logic around the model, already run Ray, or want one autoscaler for mixed GPU pools. For a single model on a fixed set of GPUs, a plain vLLM server behind a load balancer is simpler, has one less moving part to upgrade and exposes the same engine, as described in vLLM on the GPU. The cost of Ray is operational: a head node to keep healthy, version alignment between Ray, vLLM and CUDA, and a second scheduling layer to debug when replicas do not start.

What to do next

  1. Write down the concurrency and context length you must serve, then compute KV bytes per token for your model and the KV budget per GPU.
  2. Pick the smallest tensor_parallel_size whose KV budget meets the target; scale beyond that with replicas.
  3. Use STRICT_PACK for tensor parallel workers, and confirm with ray status that one node can hold a full group.
  4. Check that the accelerator_type you request matches the resource string the nodes actually report.
  5. Put the config in a Serve YAML file under version control and deploy it, rather than calling serve.run from a notebook.
  6. Stage model weights close to the cluster and measure cold-start time before tuning autoscaling delays.
  7. If you serve many fine-tunes of one base model, move them to lora_config and watch adapter load rate.
Key takeaway: Ray Serve LLM maps one LLMConfig to a deployment whose replicas each reserve tensor_parallel_size x pipeline_parallel_size GPUs as a placement group. Size a replica from memory: weights plus the KV cache your concurrency needs. Split across GPUs only until the KV budget is large enough, keep tensor parallel workers on one node with STRICT_PACK, and add replicas after that. Most production problems are placement groups that cannot be satisfied or cold starts, not model code.