Almost every large language model you will train, fine-tune or serve is driven from Python. Python itself is far too slow to multiply billion-element matrices, so what you actually use is a stack: thin Python APIs on top, compiled C++ and CUDA code underneath, and a GPU driver at the bottom. When everything lines up you write ten lines and the GPU runs flat out. When one layer disagrees with another you get cryptic import errors, silent fallbacks to slow kernels or out-of-memory crashes at step 400.

This article walks the stack from the bottom up. It explains what happens when Python calls a matrix multiply, which version numbers must agree and how to check them, what each popular library contributes and where its GPU work happens, how to do the memory arithmetic for a real model before you rent hardware, and how the training and serving halves of the stack differ. It ends with a checklist you can run against your own environment.

Advertisement

The stack in one picture

Read the diagram from the top. Your code calls a training library or a serving engine. Both sit on model libraries that know model architectures and file formats. All of them call PyTorch, which turns tensor operations into GPU work. PyTorch does not contain every kernel itself: matrix multiplies go to NVIDIA's cuBLAS, collective communication goes to NCCL, attention may go to FlashAttention, 4-bit matmuls go to bitsandbytes, and fused element-wise code may be generated on the fly by Triton. Everything below PyTorch talks to the GPU through the CUDA runtime and the kernel-mode driver.

Two properties follow. First, each layer is a separate package with its own release cycle, so compatibility is a matrix, not a single version. Second, almost all the time is spent in the bottom three layers. The Python you write matters for correctness and orchestration; performance is decided by which kernels end up running and how full the GPU stays.

The Python LLM stack: every layer is Python on top and compiled GPU code underneathYour codetraining script, eval harness, API serverTraining librariesaccelerate, peft, trl, DeepSpeedServing enginesvLLM, SGLang, TGIModel librariestransformers, tokenizers, safetensors, datasetsPyTorchPython API, autograd, dispatcher, ATen, caching allocatorVendor librariescuBLAS, cuDNN, NCCLCustom kernelsflash-attn, bitsandbytesGenerated kernelsTriton, InductorCUDA runtime + kernel-mode driverGPU: SMs, tensor cores, HBM, NVLink
The layers of a typical Python LLM environment. Arrows show calls flowing down; data stays on the GPU throughout.

What happens when Python calls a matmul

Take y = x @ w with both tensors on the GPU. The Python interpreter calls into PyTorch's C++ core. The dispatcher looks at the tensors' device, dtype and whether autograd is recording, and picks an implementation: for a bf16 CUDA matmul that is a cuBLAS or cuBLASLt routine. PyTorch allocates the output from its caching allocator, which hands out memory it already reserved from the driver rather than calling the driver each time, and enqueues the kernel on the current CUDA stream. Then it returns to Python immediately. The kernel runs later, asynchronously.

That asynchrony is why GPU code looks fast in naive timing and why some lines look slow for no reason. A call such as loss.item(), print(tensor) or tensor.cpu() forces Python to wait for every queued kernel to finish. It also explains the main overhead in small-batch inference: each kernel launch costs a few microseconds of CPU time, and a transformer decode step launches hundreds of them. When the GPU finishes a tiny kernel faster than Python can launch the next one, the GPU idles. CUDA graphs, which record a sequence of launches and replay it with one call, and torch.compile, which fuses many small operations into fewer generated kernels, both exist to remove that overhead.

import torch, time

x = torch.randn(8192, 8192, device="cuda", dtype=torch.bfloat16)
w = torch.randn(8192, 8192, device="cuda", dtype=torch.bfloat16)

t0 = time.perf_counter(); y = x @ w
print("returned after", time.perf_counter() - t0)   # microseconds: only enqueued

torch.cuda.synchronize()                              # wait for the kernel
print("finished after", time.perf_counter() - t0)    # the real cost

# Correct timing uses CUDA events recorded on the stream.
start, end = torch.cuda.Event(enable_timing=True), torch.cuda.Event(enable_timing=True)
start.record(); y = x @ w; end.record(); torch.cuda.synchronize()
ms = start.elapsed_time(end)
print(f"{2 * 8192**3 / (ms / 1e3) / 1e12:.0f} TFLOP/s achieved")
Advertisement

The version contract

Four version numbers must agree, and most broken environments break one of these rules.

  1. Driver. The kernel-mode driver installed on the host (or the node image) supports CUDA up to some version. nvidia-smi prints that maximum in its header. It is a ceiling, not the version you are using.
  2. CUDA runtime in the wheel. PyTorch wheels are built against one CUDA version and, on Linux, pull the matching runtime, cuBLAS, cuDNN and NCCL as pip packages. You choose the build with the index URL, such as https://download.pytorch.org/whl/cu126 or .../cu128; pytorch.org lists which CUDA builds exist for each PyTorch release. The wheel's CUDA version should not be newer than the driver supports; do not rely on minor-version compatibility to bridge the gap.
  3. Compiled extensions. Packages with their own CUDA code, such as flash-attn, bitsandbytes, xformers and vLLM, are compiled against a specific PyTorch ABI and CUDA version. A wheel built for a different torch fails at import with undefined-symbol errors, or worse, imports and then falls back to a slower path.
  4. GPU architecture. Kernels are compiled for specific compute capabilities. A library whose wheel omits your GPU's architecture either errors with "no kernel image is available" or JIT-compiles slowly at first use. Some kernels, such as FlashAttention variants, require a minimum architecture.

The practical consequence is that you never install a CUDA toolkit to make PyTorch work. The wheel brings its own runtime; you only need a new enough driver. You need a full toolkit only when you compile extensions from source.

# One environment per role; let the index URL pick the CUDA build.
python -m venv .venv-train && . .venv-train/bin/activate
pip install torch --index-url https://download.pytorch.org/whl/cu128
pip install transformers accelerate peft datasets safetensors bitsandbytes

# Verify every layer of the contract in one script.
python - <<'EOF'
import torch
print("torch", torch.__version__, "built for CUDA", torch.version.cuda)
print("cuDNN", torch.backends.cudnn.version(), "NCCL", torch.cuda.nccl.version())
print("GPU", torch.cuda.get_device_name(0), "capability", torch.cuda.get_device_capability(0))
print("bf16 supported:", torch.cuda.is_bf16_supported())
EOF

What each library does on the GPU

LibraryRoleWhere its GPU work happens
transformersModel definitions, loading checkpoints, generation loopCalls PyTorch ops; picks an attention implementation (eager, SDPA, FlashAttention 2)
tokenizersFast Rust tokenizerNone; runs on CPU and can starve the GPU if the data loader is single-threaded
safetensorsCheckpoint formatNone directly; memory-maps weights so loading avoids pickle and extra copies
accelerateDevice placement, mixed precision, launching multi-GPU jobsWraps PyTorch DDP and FSDP; communication goes to NCCL
peftLoRA and other adaptersAdds small trainable matrices next to frozen weights; ordinary PyTorch matmuls
bitsandbytes8-bit and 4-bit weights, 8-bit optimizersIts own CUDA kernels for quantized matmul and dequantization
flash-attnFused attentionCustom kernels that keep attention tiles in on-chip SRAM instead of writing the full score matrix to HBM
vLLM, SGLangHigh-throughput servingPaged KV cache, continuous batching, custom attention and sampling kernels, CUDA graphs

The two places where the GPU actually runs your model are therefore PyTorch's own kernels and the custom kernel packages. Everything else is orchestration. That is useful when debugging: if throughput is poor, the question is almost always which kernel ran, not which Python library you called.

Worked example: memory arithmetic for an 8B model

Before renting a GPU, do the arithmetic. Take Llama 3.1 8B: about 8 billion parameters, 32 layers, 8 key-value heads and a head dimension of 128.

ItemFormulaResult
Weights in bf168e9 params x 2 bytesabout 16 GB
KV cache per token2 (K and V) x 32 layers x 8 heads x 128 dims x 2 bytes131,072 bytes = 128 KiB
KV cache, one 8K-token sequence8,192 x 128 KiB1 GiB
Weights in 4-bit (QLoRA)8e9 x 0.5 bytes plus scalesroughly 5 GB
Full fine-tune with Adam, mixed precisionabout 16 bytes per param before activationsabout 128 GB, so multiple GPUs

Now apply it. On a 24 GB GPU, a serving engine allowed to use 90% of memory has 21.6 GB. Subtract 16 GB of weights and a gigabyte or two for activations and CUDA graphs, and roughly 4 GB remain for KV cache: about 32,000 tokens, or four concurrent 8K-token conversations. That is why serving engines expose a maximum context length and a memory-utilization fraction, and why a 4-bit or 8-bit model can serve many more users on the same card. For training, the last row shows that a full fine-tune of even an 8B model does not fit on one card, which is the reason LoRA and QLoRA exist.

The fine-tuning half of the stack

QLoRA combines three layers. bitsandbytes stores the frozen base weights in 4-bit NF4 and dequantizes tiles on the fly into a bf16 compute dtype. peft adds LoRA adapters, small low-rank matrices that are the only trainable parameters. transformers ties it together and runs the training loop. The code below is the minimal shape; real runs add a dataset, a collator and evaluation.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training

model_id = "meta-llama/Llama-3.1-8B"          # any causal LM you have access to
bnb = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,       # matmuls run in bf16 after dequantizing
)
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    quantization_config=bnb,
    device_map={"": 0},
    attn_implementation="sdpa",                  # "flash_attention_2" if flash-attn is installed
)
model = prepare_model_for_kbit_training(model)   # enables gradient checkpointing, casts norms
lora = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
                  target_modules=["q_proj", "k_proj", "v_proj", "o_proj"])
model = get_peft_model(model, lora)
model.print_trainable_parameters()               # a small fraction of a percent of the model

batch = tok(["The capital of France is"], return_tensors="pt").to(0)
out = model(**batch, labels=batch["input_ids"])
out.loss.backward()                              # gradients flow only into the LoRA matrices
print(torch.cuda.max_memory_allocated() / 2**30, "GiB peak")

Two habits make this reliable. Print torch.cuda.max_memory_allocated() after the first step so you know your real headroom before scaling sequence length. And check which attention implementation actually loaded, because transformers falls back to a slower one when flash-attn is missing or incompatible, without failing.

The serving half of the stack

Serving is a different workload. Training processes large fixed batches; serving receives requests of unpredictable length at unpredictable times and spends most of its time in memory-bound decode steps. Serving engines therefore replace the transformers generation loop with their own scheduler: continuous batching, which adds and removes requests from the running batch at every step, a paged KV cache that allocates cache in fixed blocks instead of one contiguous buffer per request, and captured CUDA graphs for decode. The engines differ in kernels and features; the serving stacks comparison covers vLLM, TGI, TensorRT-LLM and SGLang side by side.

# Offline batch inference with vLLM's Python API.
from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct",
          gpu_memory_utilization=0.90,   # fraction of GPU memory vLLM may claim
          max_model_len=8192)            # caps per-request KV cache
params = SamplingParams(temperature=0.2, max_tokens=256)
for out in llm.generate(["Explain paged attention in two sentences."], params):
    print(out.outputs[0].text)

# The same engine as an OpenAI-compatible server, split across two GPUs:
#   vllm serve meta-llama/Llama-3.1-8B-Instruct --tensor-parallel-size 2 --max-model-len 8192

Serving engines pin a specific PyTorch version because their custom kernels are compiled against it. Installing vLLM into your training environment will usually upgrade or downgrade torch underneath bitsandbytes and flash-attn. Give each engine its own virtual environment or container image.

Failure modes

  • Silent CPU tensors. A tensor created without device= lands on the CPU, and the operation either fails with a device-mismatch error or, in data-loading code, runs slowly on the CPU. Check tensor.device at boundaries.
  • Kernel fallback. A missing or mismatched flash-attn, or a GPU too old for it, causes a fallback to slower attention. The job works but runs much slower. Log the attention implementation at start-up.
  • Driver too old. torch.cuda.is_available() returns False on a machine with a GPU. Compare the wheel's torch.version.cuda with the maximum CUDA version shown by nvidia-smi.
  • Fragmentation out of memory. The error reports plenty of reserved but unallocated memory. Variable sequence lengths fragment the caching allocator; bucket sequence lengths, or try PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True.
  • Hidden synchronization. Logging .item() every step serializes CPU and GPU. Log every N steps.
  • Starved GPU. Tokenization or decompression on one CPU core keeps utilization low. Use multiple data-loader workers, or pre-tokenize the dataset.
  • NCCL hangs. Multi-GPU jobs freeze with no error when one rank crashes or the network is misconfigured. Set NCCL_DEBUG=INFO and a collective timeout; see NCCL collectives.

Operating the stack: trade-offs that matter

Pin, then upgrade deliberately. Lock every package in a lock file and rebuild images from it. Upgrading torch is a project, because every compiled extension must move with it. Upgrade the whole set together and rerun a fixed benchmark and a fixed evaluation before rolling out.

Containers versus virtual environments. Containers capture system libraries and are the norm for serving and cluster training. Plain virtual environments are fine for development as long as the driver on the host is new enough. Either way, the driver belongs to the host, never to the image.

Eager versus compiled. Eager PyTorch is easiest to debug. Compilation and CUDA graphs buy throughput but add warm-up time and new failure modes. Start eager, measure, then compile the hot path.

Precision. bf16 is the default for training on GPUs that support it. 8-bit and 4-bit weights trade a small quality loss for capacity; measure that loss on your own evaluation set, not a leaderboard. Knowing where the bytes go, as in the GPU memory hierarchy, tells you which of these levers will help.

What to do next

  1. Run the verification script in every environment you own and record torch, CUDA build, cuDNN, NCCL and compute capability next to the driver version.
  2. Split training and serving into separate environments or images, each from its own lock file.
  3. Do the memory arithmetic for your target model: weights, KV cache per token, and the concurrency you need.
  4. Run one QLoRA step and one vLLM generation, printing peak memory and the attention implementation used.
  5. Time a forward pass correctly with CUDA events and compare it with the GPU's peak throughput.
  6. Profile one training step and one decode step to see which kernels dominate before optimizing anything.
Key takeaway: The Python LLM stack is a set of layers: your code, training libraries and serving engines, model libraries, PyTorch, kernel libraries, and the CUDA runtime and driver. Python only enqueues work; the GPU time goes to cuBLAS, NCCL, FlashAttention, bitsandbytes and generated kernels. Keep the version contract in order: a driver new enough for the wheel's CUDA build, and compiled extensions built for your exact torch. Do the memory arithmetic before buying hardware, keep training and serving in separate pinned environments, and always check which kernels actually ran.