GPU Technology

GPU Technology

Deep technical articles on this topic.

579Articles
579Topics covered
Articles in this category

All 467 articles, sorted alphabetically

ARTICLE · 001

100+ MW AI Clusters, in depth: what the power budget buys, how often a 50,000-GPU job breaks, checkpoint intervals and goodput, power swings and fabric depth

A 100 MW AI site from the training job's side: per-GPU all-in power and the GPU count it buys, failure rates scaled from Llama 3 data, Young/Daly…

Read article →
ARTICLE · 002

GPU Allocation Queues, in depth: quota, provider capacity and cluster admission as one pipeline, why large gangs wait longest, and how to shape requests

GPU allocation as a queueing problem: the quota, provider capacity and cluster admission queues, an Erlang C model of how gang size and load drive wai…

Read article →
ARTICLE · 003

Amazon Inferentia2 + Inferentia3, in depth: the Inferentia line in 2026, why there is no Inferentia3, what still belongs on Inf2, leaving Inf1 and the LLM decision

The Inferentia product line as it stands in October 2026: Inf1 and Inf2 compared, the absence of an Inferentia3, which workloads still belong on Inf2,…

Read article →
ARTICLE · 004

Amazon Trainium 2 + UltraServers, in depth: the 64-chip NeuronLink domain, where 1.28 TB/s goes, mapping parallelism onto 4x4x4, failure blast radius and Trn3

How a Trn2 UltraServer works as a system: the 4x4 torus plus cross-instance rings, reconciling the bandwidth figures, collective cost at each level, p…

Read article →
ARTICLE · 005

AMD MI325X + MI350, in depth: the CDNA 3 to CDNA 4 step for kernels, FP8 encodings, MXFP4, memory plans and migration

What changes from AMD's MI325X to the MI350X and MI355X for software: 256 wider CUs and 160 KB LDS, machine balance, OCP FP8 versus FNUZ, MX form…

Read article →
ARTICLE · 006

AMD ROCm Software Stack, in depth: the layers from amdgpu and KFD to HIP and libraries, code objects, containers and triage

A layer map of the AMD ROCm stack for operators: amdgpu and KFD device nodes, the ROCr HSA runtime, the HIP runtime and compiler, code objects and gfx…

Read article →
ARTICLE · 007

Apple M-Series Neural Engine, in depth: what is known, ANE-friendly layers, checking op placement with MLComputePlan, fp16 limits and W8A8

How software reaches the Apple Neural Engine: the Core ML pipeline and graph partitioning, the (B, C, 1, S) transformer layout with 1x1 convolutions, …

Read article →
ARTICLE · 008

ARM Neoverse for AI, in depth: core families, SVE2, BF16 and I8MM, runtime dispatch and bandwidth-bound CPU inference

How ML software uses Arm Neoverse cores in Graviton, Grace, Axion and Cobalt: which core ships where, what SVE2, BFMMLA and SMMLA compute, HWCAP featu…

Read article →
ARTICLE · 009

AWS P5 and P4 GPU Instances, in depth: six instance types, per-GPU network math, EFA setup, NCCL and choosing between them

AWS P4 (A100) and P5 (H100, H200) instances for training and serving: specs checked against AWS documentation, the NVSwitch and EFA fabrics, per-GPU n…

Read article →
ARTICLE · 010

AWS SageMaker HyperPod, in depth: instance groups, lifecycle scripts, health agents, deep health checks, auto-resume and what failures still cost

How SageMaker HyperPod clusters work for distributed training: the instance group model, a CreateCluster definition, lifecycle scripts, the health age…

Read article →
ARTICLE · 011

AWS UltraCluster, in depth: getting co-located capacity, Capacity Blocks, topology-aware rank ordering, EFA collectives and failure handling at thousands of GPUs

EC2 UltraClusters explained for the engineer running the job: which instances live in them and their EFA bandwidth, Capacity Blocks and placement, the…

Read article →
ARTICLE · 012

Azure ML Compute, in depth: GPU compute targets, cluster settings, quota, multi-node PyTorch, InfiniBand and low-priority capacity

How Azure Machine Learning compute runs GPU training: choosing between clusters, serverless and Kubernetes, the cluster settings that decide cost, how…

Read article →
ARTICLE · 013

Azure ND-Series GPU VMs, in depth: the current sizes, one InfiniBand adapter per GPU, images and NCCL topology, sizing by memory and node health

What Azure's ND GPU VM sizes contain (A100 v4, H100 v5, H200 v5, MI300X v5, GB200 v6), how NVSwitch and per-GPU InfiniBand adapters shape paralle…

Read article →
ARTICLE · 014

Bandwidth per GPU (Bisection), in depth: injection vs bisection vs NVLink, oversubscription, which collectives care, a model and how to measure it

What per-GPU bisection bandwidth means, how to compute it from a leaf-spine or rail-optimised fabric, why all-to-all is bisection-bound and ring all-r…

Read article →
ARTICLE · 015

Shared Memory Bank Conflicts, in depth: the bank function, counting wavefronts, padding versus XOR swizzle, and proving the fix in Nsight Compute

How shared memory banks work and how to remove conflicts: the bank function and broadcast rule, the gcd stride rule, 64-bit and 128-bit phases, a Pyth…

Read article →
ARTICLE · 016

Bare Metal vs Cloud GPU Cost Analysis, in depth: utilization break-even, commitment layering and a model you can run

How to decide between on-demand cloud, committed cloud, rented bare metal and owned GPUs: what bare metal changes technically, the commit-level rule f…

Read article →
ARTICLE · 017

Batching Impact on Inference Cost, in depth: achieved versus configured batch, fleet consolidation, the price of a latency target, padding and cost attribution

Why the batch an LLM server actually runs is set by traffic per replica, not the configured maximum: Little's law for achieved batch, a worked fl…

Read article →
ARTICLE · 018

Datacenter Busbars and Power Distribution, in depth: the in-rack DC busbar, power shelves, voltage drop and heat, 800 VDC, and what training jobs do to the rack

How AI racks deliver power to GPUs: the conversion chain from grid to die, why 48 to 54 V DC busbars replaced cords, a worked busbar sizing example wi…

Read article →
ARTICLE · 019

Cerebras WSE-3, in depth: the wafer-scale processor, the PE mesh, weight streaming for training, multi-wafer inference and how software reaches it

How the Cerebras WSE-3 works and how software uses it: 900,000 PEs with local SRAM, why 21 PB/s changes the decode bottleneck, weight streaming with M…

Read article →
ARTICLE · 020

Chilled Water Systems for AI DCs, in depth: chillers and lift, warm and cold loops, sizing flows and N+1, staging, ride-through storage and failure modes

How the chilled water plant behind a liquid-cooled AI hall works: chiller lift and COP, ASHRAE W-classes, two-temperature loops, flow and tonnage arit…

Read article →
ARTICLE · 021

Coalesced Memory Access

How a warp's 32 lane addresses turn into global-memory transactions: sector and cache-line granularity, sector efficiency as requested bytes over…

Read article →
ARTICLE · 022

Coherent Optics for Long-Distance, in depth: field modulation, the DSP chain, 400ZR and ZR+, OSNR link budgets and training across sites

How coherent optics carry data centre interconnects: dual-polarisation QAM, the receive DSP chain, the FEC cliff, the 400ZR family, an OSNR and disper…

Read article →
ARTICLE · 023

Cold Plate Manufacturing, in depth: fin forming, lid joining, flatness, cleanliness, factory tests, traceability and measuring plate quality from GPU telemetry

How GPU cold plates are made and how manufacturing shows up in telemetry: materials, skived, machined, bonded and printed fins, brazing and friction s…

Read article →
ARTICLE · 024

Colossal-AI Framework, in depth: the Booster and its plugins, Gemini chunk memory, HybridParallelPlugin with Shardformer, and sizing a 13B run

Colossal-AI as a practitioner uses it: the Booster API and the six plugins, a minimal training program, Gemini's chunk manager and placement knob…

Read article →
ARTICLE · 025

Compute + Storage Co-Design, in depth: the four traffic classes of a GPU cluster, a worked 1,024-GPU storage budget, tier placement, job shaping and interference failures

How to size and shape storage for GPU training and serving: training reads, checkpoint bursts, restart reads and model loads; a worked 1,024-GPU budge…

Read article →
ARTICLE · 026

CUDA Cooperative Groups: Thread Block Tiles, Grid Sync and Cooperative Launch

A practical guide to CUDA Cooperative Groups: thread block tiles and warp collectives without masks, cg::reduce, coalesced and labeled partitions for …

Read article →
ARTICLE · 027

Copper vs Optical for Intra-DC, in depth: cable families, why reach halves per lane-rate step, a cable planner and what failing links look like to software

Copper versus optics inside an AI datacenter, hop by hop: DAC, ACC, AEC, AOC and transceivers, the loss-budget physics behind copper reach, power and …

Read article →
ARTICLE · 028

CoreWeave

How the GPU neocloud model works, with CoreWeave as the worked example: non-blocking InfiniBand fabric versus oversubscribed general-purpose cloud net…

Read article →
ARTICLE · 029

Cost per Token Trained, in depth: marginal versus fully loaded cost, the denominator problem, live tracking and amortisation

How to compute the cost of each training token: the 6N derivation with MFU and goodput, a 7B on 2T tokens worked example, overhead multipliers for exp…

Read article →
ARTICLE · 030

Crusoe Energy, in depth: from flare gas to gigawatt campuses, what power-first siting means for latency and data, and running jobs on Crusoe Cloud

Crusoe from first principles: why it agreed to sell its flare-gas business, how power-first siting shapes latency, data gravity, capacity timing and c…

Read article →
ARTICLE · 031

cuBLAS, in depth: column-major GEMMs from row-major frameworks, handles and workspace, cuBLASLt heuristics and epilogues, FP8 rules and reproducibility

cuBLAS as an API from first principles: the row-major trick, handles, streams and workspace, compute types and emulation, cuBLASLt descriptors, heuris…

Read article →
ARTICLE · 032

CUDA Cores and Streaming Multiprocessors: Inside the GPU, from Warps to Register Files

What a CUDA core really is, what sits inside an H100 streaming multiprocessor, how blocks and warps map onto SMs, the register file as the largest on-…

Read article →
ARTICLE · 033

CUDA Streams and Concurrency, in depth: default-stream semantics, stream-ordered APIs and libraries, a chunked copy-compute pipeline, host functions and async errors

CUDA streams at the runtime level: legacy versus per-thread default streams, which APIs and libraries are stream-ordered, writing stream-correct libra…

Read article →
ARTICLE · 034

cuDNN, in depth: the graph API, engines and execution plans, a fused convolution with bias and ReLU, PyTorch switches and failure modes

How cuDNN turns deep-learning operations into GPU kernels: legacy versus graph API, the cuDNN 9 sub-libraries, heuristics, engine configs and plans, a…

Read article →
ARTICLE · 035

Data Curation Cost for Foundation Models, in depth: cost per retained token, filter ordering, the ablation bill and a break-even against training compute

What curating pretraining data really costs: cost per retained token through the funnel, a runnable cost model, a 2-billion-page worked example where …

Read article →
ARTICLE · 036

Data Parallel Training, in depth: the DDP reducer and its buckets, a complete torchrun script, accumulation, scaling measurements and the traps that hang jobs

Data parallel training with PyTorch DDP as an engine: the reducer, gradient buckets and overlap, a complete torchrun script, gradient accumulation wit…

Read article →
ARTICLE · 037

DGX H100 System Architecture, in depth: the three fabrics, GPU-to-NIC and NUMA affinity, the local NVMe cache, a worked all-reduce estimate, validation and failure modes

How a DGX H100 looks to software: eight H100 GPUs on four NVSwitch chips, one 400 Gb/s ConnectX-7 rail per GPU, two NUMA domains, a RAID 0 NVMe cache …

Read article →
ARTICLE · 038

Distillation for Cost Savings, in depth: GPU-seconds per request, the one-time ledger, fallback and upkeep, and the break-even volume

When does distilling a large model into a small one actually save money? Measuring GPU-seconds per request at the SLO, pricing data, training, evaluat…

Read article →
ARTICLE · 039

Distributed Checkpointing, in depth: global-offset metadata, the save and load planners, resharding by chunk intersection, deduplication and format conversion

How PyTorch Distributed Checkpoint makes checkpoints independent of the parallel layout that wrote them: the .metadata index and .distcp files, the pl…

Read article →
ARTICLE · 040

Dragonfly Network Topology, in depth: groups and global links, the radix arithmetic, minimal versus Valiant and UGAL routing, and placing training jobs

The dragonfly interconnect from first principles: p, a and h, router radix and maximum size, minimal local-global-local routes, the group-shift patter…

Read article →
ARTICLE · 041

CUDA Dynamic Parallelism, in depth: CDP2 tail launch and fire-and-forget streams, memory visibility, launch-pool limits and migrating from CDP1

How CUDA Dynamic Parallelism works since CUDA 12 (CDP2): parent and child grids, device streams including tail launch and fire-and-forget, memory visi…

Read article →
ARTICLE · 042

AWS EFA, in depth: interface types, the SRD transport, the libfabric stack, network rules, counters for triage and EFA on EKS

AWS Elastic Fabric Adapter from the software side: ENA, EFA and EFA-only interfaces, why SRD is reliable but unordered and multipath, the NCCL to libf…

Read article →
ARTICLE · 043

Fat-Tree Network Topology, in depth: radix math, bisection, oversubscription, routing collisions and job placement for GPU clusters

How fat-tree (folded Clos) networks for GPU clusters are sized from switch radix, what bisection bandwidth and oversubscription mean for all-reduce an…

Read article →
ARTICLE · 044

FLOPS Budget for LLM Training, in depth: counting 6N plus attention, MFU versus HFU, goodput and a worked GPU-hour ledger

Plan the compute for an LLM training run: derive FLOPs per token from 6N plus the attention term, use dense rather than sparsity peaks, convert with m…

Read article →
ARTICLE · 045

FluidStack, in depth: what a dedicated neocloud GPU cluster is, how its fabric shapes training, acceptance burn-in, topology-aware Slurm, checkpoint cadence and the contract questions that matter

Fluidstack and the dedicated-cluster neocloud model, explained for training teams: node and fabric layout, a reproducible acceptance test with nvidia-…

Read article →
ARTICLE · 046

GB200 Grace Blackwell, in depth: the compute tray as a Linux host, Arm porting, NUMA affinity and tray acceptance testing

GB200 from the host's point of view: what a compute tray exposes to Linux, porting a training stack to aarch64, 4 KB versus 64 KB pages, NUMA and…

Read article →
ARTICLE · 047

GKE with TPU, in depth: slice node pools, labels and chip requests, Indexed Jobs, Multislice with JobSet, capacity options and failure modes

How TPUs work on Google Kubernetes Engine: single-host and multi-host slice node pools and their atomic behaviour, the accelerator and topology labels…

Read article →
ARTICLE · 048

Google Cloud TPU Pods, in depth: pods, slices and cubes, torus wraparound, a collective cost model and mapping JAX meshes onto ICI and DCN

Cloud TPU pods as a network: pod, slice, cube, host and DCN vocabulary, v5p, v6e and Ironwood pod figures, when torus wraparound exists, twisted topol…

Read article →
ARTICLE · 049

Google TPU v4 + v5 (v5e + v5p), in depth: megacore, the optically switched torus, SparseCores, and moving v4 jobs to v5p and v5e

TPU v4 as the reference design for v5e and v5p: megacore and TensorCore-counted names, 4x4x4 cubes joined by optical circuit switches, twisted tori, S…

Read article →
ARTICLE · 050

Ada Lovelace / Blackwell architecture

Deep-dive on NVIDIA Hopper→Blackwell arc: SM redesign, FP8/FP4, memory bandwidth, NVLink generations, GB200 Grace-Blackwell.

Read article →
ARTICLE · 051

Admission Control for LLM Serving

How admission control protects p99 latency in LLM serving: why unbounded queues fail, sizing a queue-depth cap from the SLO with Little's Law, lo…

Read article →
ARTICLE · 052

AI Gateway Overview

How an AI gateway works as a policy plane in front of LLM backends: virtual API keys and identity, per-key and per-team budget enforcement, rate limit…

Read article →
ARTICLE · 053

Alignment Data, in depth: SFT and preference formats, chat templates, loss masks, packing, decontamination and GPU cost

Engineering alignment data for training: the three record shapes, chat templates and loss masking, packing with cu_seqlens, cached reference log-probs…

Read article →
ARTICLE · 054

AMD Instinct GPUs

AMD Instinct and ROCm as the practical alternative to CUDA: how HIP maps onto the CUDA programming model, what hipify translates and what it silently …

Read article →
ARTICLE · 055

Anthropic Workbench + Claude, in depth: what replaced the Console Workbench, and how to rebuild its prompt-evaluation loop in code

Workbench (legacy) in the Claude Console is retired and replaced by a stateless playground. What the playground does, what was lost, and how to rebuil…

Read article →
ARTICLE · 056

Apigee AI Gateway, in depth: token quotas, prompt spike limits, semantic caching and Model Armor in front of GPU-backed models

How to turn an Apigee proxy into an AI gateway that protects GPU serving capacity: LLMTokenQuota enforce and count pairs, PromptTokenLimit spike contr…

Read article →
ARTICLE · 057

Apple Silicon, in depth: unified memory, the GPU and its Neural Accelerators, the Neural Engine, and how MLX, PyTorch and llama.cpp use them for local inference and fine-tuning

How Apple Silicon works as a machine-learning computer: unified memory and why it changes what fits, the CPU, GPU, M5 Neural Accelerators and Neural E…

Read article →
ARTICLE · 058

GPU Architecture Explained: Streaming Multiprocessors and Warps

How GPU architecture works: streaming multiprocessors (SMs), warps of 32 threads and SIMT execution, the memory hierarchy, occupancy and tensor cores.

Read article →
ARTICLE · 059

ASR (Speech-to-Text) Serving, in depth: request pipelines, long-form chunking, batching, RTFx capacity planning and failure guards

How to serve speech-to-text on GPUs: batch versus streaming workloads, the decode-VAD-chunk-batch-stitch pipeline, where Whisper spends GPU time, a VA…

Read article →
ARTICLE · 060

GPU asynchronous copy and software pipelining

Deep-dive on the asynchronous global-to-shared copy pipeline that fast GPU kernels are built on: cp.async streams tiles directly into a multi-stage sh…

Read article →
ARTICLE · 061

Audio Encoders, in depth: log-mel and learned front ends, Whisper and wav2vec frame math, GPU cost, batching and streaming

How audio encoders such as Whisper, wav2vec 2.0 and HuBERT turn waveforms into frames, how to estimate their GPU cost from frame counts, and how fixed…

Read article →
ARTICLE · 062

Axolotl, in depth: config-driven LLM fine-tuning, from YAML to packed batches, QLoRA memory, multi-GPU sharding and the mistakes that waste a run

How Axolotl turns one YAML file into a fine-tuning run: what each config block does at run time, chat_template datasets and label masking, the prepare…

Read article →
ARTICLE · 063

Azure AI Foundry, in depth: resources, projects, deployment types and planning GPU capacity

How Azure AI Foundry (now Microsoft Foundry) works: the resource and project model, RBAC, how Standard, Provisioned, Batch and managed compute deploym…

Read article →
ARTICLE · 064

NVIDIA B100, in depth: Blackwell at Hopper's 700 W, what a power-capped GPU means for training and inference, and how to plan around it

The B100 was announced as the 700 W Blackwell GPU for existing HGX H100 power and cooling. What its launch figures mean, why a power cap hurts compute…

Read article →
ARTICLE · 065

NVIDIA B200, in depth: dual-die Blackwell, FP4 and microscaling, tensor memory, and what changes for training and inference

A practitioner's guide to the NVIDIA B200: the dual-die package, HBM3e capacity and bandwidth, fifth-generation tensor cores with tensor memory, …

Read article →
ARTICLE · 066

GPU Batching Strategies Overview

A decision guide to GPU batching for inference: static, dynamic, continuous and chunked prefill compared on the workload signals that actually select …

Read article →
ARTICLE · 067

Arena-Hard, in depth: BenchBuilder prompts, the pairwise judging protocol, scoring code, serving a candidate on your GPUs and gating releases

How Arena-Hard-Auto works and how to run it: v0.1 and v2.0 prompt sets, judges and baselines, the five-verdict two-game protocol, weighted bootstrap s…

Read article →
ARTICLE · 068

Custom LLM Benchmarks, in depth: item sets from your own traffic, programmatic scorers, paired statistics and the GPU noise floor

Build a custom LLM benchmark that decides between models, quantisations and serving configs: items from production traffic, stratification, programmat…

Read article →
ARTICLE · 069

LLM Eval Frameworks, in depth: how lm-evaluation-harness turns benchmarks into GPU work, from request types and batching to reproducible scores

How LLM eval frameworks work on the GPU: the request-centric pipeline, loglikelihood versus generate requests, boundary tokenisation, acc versus acc_n…

Read article →
ARTICLE · 070

HELM Benchmark for Serving, in depth: wiring the harness to your inference server, reading its outputs and gating quantization and engine changes with paired statistics

How to use Stanford CRFM HELM against your own vLLM-style endpoint: scenarios, adapters and run entries, model_deployments.yaml with VLLMClient, the p…

Read article →
ARTICLE · 071

LiveBench, in depth: contamination-limited releases, objective scorers, running it against a model on your own GPUs and comparing serving configurations

How LiveBench works and how to run it on your own GPUs: its categories and tasks, dated releases and private questions, the answer-judge-show pipeline…

Read article →
ARTICLE · 072

LLMPerf Benchmark, in depth: how the load test measures, the inter-token metric that misleads, recomputing TPOT and goodput, and a worked concurrency sweep

How Ray's archived LLMPerf load generator measures TTFT, inter-token latency and throughput, read from its source; why its inter-token figure inc…

Read article →
ARTICLE · 073

LMSYS Chatbot Arena, in depth: anonymous battles, Bradley-Terry scores, bootstrap ranks, style control and running a private arena

How Chatbot Arena (LMArena) turns anonymous pairwise votes into scores: the battle protocol, Bradley-Terry instead of online Elo, bootstrap intervals …

Read article →
ARTICLE · 074

MLPerf Benchmark, in depth: how Training and Inference results are produced, LoadGen scenarios, LLM latency limits, divisions and how to read a result

How MLPerf measures GPUs and other accelerators: Training time-to-quality with olympic scoring and reference convergence points, Inference LoadGen sce…

Read article →
ARTICLE · 075

SWE-bench for Serving Model Choice, in depth: resolve rate, cost and latency on your own endpoint

Use SWE-bench to choose a serving configuration, not a leaderboard winner: the serving knobs that move resolve rate, a pinned-scaffold rig, harness co…

Read article →
ARTICLE · 076

vLLM Benchmark Suite, in depth: bench latency, throughput and serve, arrival processes, concurrency caps, goodput and a reproducible sweep

How to use vLLM's own benchmarks: offline latency and throughput versus online bench serve, TTFT, TPOT, ITL and goodput, request rate, burstiness…

Read article →
ARTICLE · 077

BentoML for LLM Serving, in depth: services, vLLM integration, composition, adaptive batching, concurrency autoscaling and cold starts

A practical guide to serving LLMs with BentoML: what the framework does above the inference engine, a verified vLLM service using __command__, service…

Read article →
ARTICLE · 078

Bricktree / Traefik AI, in depth: Traefik Hub's AI Gateway middlewares as a capacity control for GPU inference

Traefik Hub AI Gateway in front of self-hosted GPU inference: enabling it, middleware order, Chat Completion, token rate limits and quotas, semantic c…

Read article →
ARTICLE · 079

GPU Capacity Planning, in depth: demand ledgers, stranded GPUs, spare pools and the buy trigger for a shared fleet

Capacity planning for a running, shared GPU fleet: a GPU-hour demand ledger with per-class utilisation targets, a placement simulation showing strande…

Read article →
ARTICLE · 080

GPU Carbon Footprint, in depth: measuring energy per job, operational and embodied emissions, and the software levers that cut them

How to measure and reduce the carbon footprint of GPU training and inference: what is counted, reading the NVML energy counter, node overhead and PUE,…

Read article →
ARTICLE · 081

GPU Checkpointing Deep Dive, in depth: bytes per parameter, the save path, async sharded saves and commits that survive crashes

How large-model training checkpoints work end to end: what state costs per parameter, the GPU-to-storage save path and which hop blocks training, a wo…

Read article →
ARTICLE · 082

Chunked Prefill in Serving

How chunked prefill splits a long prompt into decode-sized slices so it stops blocking other requests' decode: the head-of-line stall it removes,…

Read article →
ARTICLE · 083

CLIP Training on GPU, in depth: the contrastive loss, why batch size is everything, distributed local loss, memory arithmetic, precision and the data pipeline

How contrastive image-text pre-training actually runs on GPUs: the symmetric InfoNCE loss, why the negative set makes batch size a quality knob, the q…

Read article →
ARTICLE · 084

GPU Cluster Bandwidth, in depth: the HBM-to-NIC ladder, what each parallelism strategy sends, a per-step communication budget, and how to measure it

The bandwidth ladder of a GPU training cluster with checked H100 and B200 figures, algbw versus busbw, bytes per step for DDP, FSDP and tensor paralle…

Read article →
ARTICLE · 085

GPU Collective Operations, in depth: all-reduce, all-gather, reduce-scatter and all-to-all mapped to DDP, FSDP, tensor and expert parallelism

GPU collectives from first principles: what each operation does, which parallelism strategy issues it, an alpha-beta cost model, algorithm versus bus …

Read article →
ARTICLE · 086

Collective Communication Overlap

How data-parallel training hides gradient all-reduce behind backward compute: gradient bucketing and why buckets fire during backward instead of after…

Read article →
ARTICLE · 087

GPU-Enabled Containers

How to run GPU workloads in Docker containers via nvidia-container-toolkit.

Read article →
ARTICLE · 088

Continuous batching

Deep-dive on continuous (in-flight) batching for LLM inference: the scheduling technique that forms a fresh batch every decode iteration, letting fini…

Read article →
ARTICLE · 089

GPU Data Center Cooling: Air, Direct-to-Chip Liquid, Immersion

How GPU server cooling is chosen by kW per rack: air, rear-door heat exchangers, direct-to-chip liquid cooling, immersion, and the CDU coolant loop be…

Read article →
ARTICLE · 090

Core ML for LLM, in depth: stateful KV cache, flexible shapes, int4 weights and the on-device generation loop

How to run a decoder language model through Core ML on Apple devices: why LLMs fit Core ML awkwardly, wrapping the KV cache as model state, converting…

Read article →
ARTICLE · 091

GPU Cost Optimization

An ordered lever list for cutting GPU spend: utilization first, then right-sizing, batching, cache hit rate, quantization, request routing, context tr…

Read article →
ARTICLE · 092

GPU CUDA Programming

The CUDA programming model: how kernels launch threads organized into blocks and grids, and the __global__/__device__ function distinction.

Read article →
ARTICLE · 093

CUDA streams and graphs -- overlap and launch-overhead elimination

Deep-dive on CUDA streams and graphs: streams as ordered work queues enabling concurrency and copy-compute overlap, events for cross-stream synchroniz…

Read article →
ARTICLE · 094

CUTLASS, in depth: the GEMM hierarchy, CuTe layouts, Hopper kernel schedules and the Python DSL

How NVIDIA CUTLASS builds matrix multiplications: tiling from first principles, the device, kernel, collective and atom layers, CuTe layouts, a comple…

Read article →
ARTICLE · 095

GPU Data Pipeline, in depth: measuring data stalls, DataLoader workers, pinned memory, GPU decode and sharded storage

How to keep GPUs fed during training: the input pipeline as a five-stage throughput budget, measuring data wait against GPU time, PyTorch DataLoader w…

Read article →
ARTICLE · 096

GPU Datacenter Deployment

Deploying GPU capacity as a build project: turning a megawatt envelope into a GPU count, floor loading and freight paths, the cable plant, commissioni…

Read article →
ARTICLE · 097

GPU Datacenter Power Requirements

How power is delivered to a GPU hall: MW-per-rack math, why synchronized training looks like a square wave to the grid, step loads and breaker coordin…

Read article →
ARTICLE · 098

GPU DataLoader Bottleneck, in depth: inside the PyTorch DataLoader, proving the loader is the limit, and fixing it

How the PyTorch DataLoader really works (workers, index queues, shared memory, pin thread, reorder buffer), three measurements that prove an input bot…

Read article →
ARTICLE · 099

NVIDIA DCGM, in depth: the host engine, profiling metrics, health watches, diagnostics as a node gate and dcgm-exporter

NVIDIA Data Center GPU Manager as a fleet system: the host engine and its watch cache, groups and fields, what the profiling metrics say about a train…

Read article →
ARTICLE · 100

Decode Compute Math, in depth: a per-operator FLOP and byte ledger for one LLM decode step

The arithmetic of one LLM decode step, operator by operator: FLOPs and HBM bytes for QKV, attention, MLP and LM head, why attention intensity equals t…

Read article →
ARTICLE · 101

DeepInfra, in depth: the four ways to run a model, the per-model concurrency limit as a throughput ceiling, custom deployments, the Batch API and break-even math

DeepInfra for engineers: OpenAI-compatible and native APIs, the 200-concurrent-requests-per-model limit and Little's law, a client limiter with b…

Read article →
ARTICLE · 102

DeepSpeed, in depth: the engine, the launcher, custom ops, pipeline and MoE engines, and when to pick it over FSDP

How the DeepSpeed library is put together and how to run it: what deepspeed.initialize wraps, the batch-size triangle, backward and step at accumulati…

Read article →
ARTICLE · 103

GPU Depreciation Timeline, in depth: book life versus economic life, what useful-life disclosures mean, the training-to-inference cascade and when to replace a fleet

GPU depreciation as two clocks: straight-line book value against front-loaded economic value, hyperscaler useful-life changes, the five drivers of val…

Read article →
ARTICLE · 104

Diffusion Model Serving, in depth: the cost model, shape buckets, step-level batching, prompt caching, compilation, LoRA handling and a split decode stage

How to serve text-to-image diffusion models on GPUs: a cost model built from steps, resolution and guidance, a staged request path, resolution buckets…

Read article →
ARTICLE · 105

Direct Liquid Cooling, in depth: the heat path from die to coolant, inside a cold plate, flow and pressure drop, the residual air load and what a training job sees

How direct-to-chip liquid cooling works for GPU servers: the thermal resistance stack from junction to coolant, cold plate construction, flow per plat…

Read article →
ARTICLE · 106

Disaggregated LLM Serving, in depth: the request path through a real proxy, vLLM and SGLang launch, routing, failure handling, and splits beyond prefill and decode

An operator's guide to disaggregated LLM serving: the kv_transfer_params handshake through a proxy, launching vLLM with NixlConnector and SGLang …

Read article →
ARTICLE · 107

DPO Training on GPU, in depth: four forward passes, the reference-model bill, logits memory and a run that fits

The systems side of Direct Preference Optimization: what one DPO step computes on the GPU, where memory goes for an 8B model, the four ways to host th…

Read article →
ARTICLE · 108

Dropless MoE, in depth: sort and permute dispatch, grouped and block-sparse GEMM, variable all-to-all, worst-case memory and stragglers

How dropless mixture-of-experts works on GPUs: why capacity factors drop and pad, the sort, permute and unpermute data flow, a PyTorch reference, grou…

Read article →
ARTICLE · 109

PyTorch Dynamo, in depth: PEP 523 frame hooks, symbolic bytecode interpretation, sources and guards, dynamic shape symbols, graph breaks as continuations and writing your own backend

How TorchDynamo, the graph-capture front end of torch.compile, really works: the PEP 523 frame evaluation hook, InstructionTranslator and VariableTrac…

Read article →
ARTICLE · 110

Edge Inference for LLM, in depth: bandwidth-bound decode, memory and KV budgets, runtimes, thermals and a hybrid local-cloud router

How to run language models on edge GPUs and SoCs: why decode is bandwidth-bound, sizing weights and KV cache, ceilings for Jetson Orin and Apple M4 pa…

Read article →
ARTICLE · 111

Edge NPUs, in depth: how on-device neural accelerators execute models, and how to quantize, compile and deploy for them

How edge neural processing units work and how software uses them: MAC arrays, scratchpad SRAM and DMA tiling, why integer quantization is the entry fe…

Read article →
ARTICLE · 112

Embedding Model Serving, in depth: token-budget batching, pooling and prefixes, TEI configuration, capacity math and versioned embedding contracts

How to serve embedding models on GPUs: what the encoder forward pass costs, CLS, mean and last-token pooling, instruction prefixes, token-budget batch…

Read article →
ARTICLE · 113

GPU ECC and Memory Reliability, in depth: SECDED from first principles, containment, row remapping budgets and a drain-reset-replace policy

How GPU error-correcting codes work and how to operate them: a runnable SECDED encoder and decoder, where ECC and parity sit on the die and in HBM, vo…

Read article →
ARTICLE · 114

GPU Experiment Tracking, in depth: stack fingerprints, honest step timing, MFU and goodput, cost ledgers and fair run comparisons

The efficiency and cost half of GPU experiment tracking: per-node stack fingerprints as comparison keys, step timing on an asynchronous device, MFU wi…

Read article →
ARTICLE · 115

Expert Parallelism

Expert parallelism on the training side: how EP, tensor and data parallelism compose into one device mesh, which parameters shard by expert and which …

Read article →
ARTICLE · 116

DoRA Fine-Tuning, in depth: magnitude and direction, a from-scratch layer, PEFT internals and GPU cost

DoRA from first principles: decomposing weights into a trainable magnitude and a LoRA-updated direction, a from-scratch PyTorch layer with merging, ho…

Read article →
ARTICLE · 117

DPO Fine-Tuning, in depth: what beta controls, sweeping beta with learning rate on a GPU budget, measuring drift and choosing a checkpoint with real confidence intervals

Running a DPO fine-tune as an experiment: beta as the KL price with a worked toy, the gradient weight that couples beta and learning rate, a LoRA swee…

Read article →
ARTICLE · 118

Full Fine-Tuning in Depth: Memory Budget, Data Preparation, Training Recipe, and When It Beats LoRA, with Pseudocode

How to fully fine-tune a pretrained language model on GPUs: the 16 bytes per parameter model-state budget and which ZeRO or FSDP stage fits 13B and 70…

Read article →
ARTICLE · 119

GRPO Fine-Tuning, in depth: a run plan from task fit and difficulty filtering to rewards, LoRA or vLLM server layouts, sizing and checkpoint choice

A practical GRPO fine-tuning runbook: when GRPO fits, filtering prompts by pass rate so groups carry signal, testable reward functions, TRL configs fo…

Read article →
ARTICLE · 120

LoRA Fine-Tuning, in depth: where GPU memory and time go when you train adapters on an 8B model

What a LoRA training step does on the GPU: why frozen weights still cost an input-gradient GEMM, 4N against 6N FLOPs per token, a memory budget for an…

Read article →
ARTICLE · 121

Fine-Tuning Ops on GPU, in depth: a runbook for memory planning, launching, checkpointing, monitoring and shipping adapters

An operations runbook for fine-tuning LLMs on GPUs: reproducible run specs, a worked LoRA and QLoRA memory plan for an 8B model, the fp32 logits trap,…

Read article →
ARTICLE · 122

ORPO Fine-Tuning, in depth: building preference pairs, templates and masking, memory planning, a lambda sweep and promotion gates

A practitioner's recipe for ORPO fine-tuning: when one-stage alignment fits, building preference pairs from your own model, chat templates and tr…

Read article →
ARTICLE · 123

PPO Fine-Tuning for LLMs, in depth: a run plan from method choice and reward-model audits to memory sizing, reward-hacking guards and checkpoint selection

A practical PPO fine-tuning plan for LLMs: when PPO beats DPO or GRPO, the 2026 tooling change after TRL removed PPOTrainer, auditing the reward model…

Read article →
ARTICLE · 124

SFT on GPUs, in depth: length distributions, padding versus packing, padding-free attention, logit memory and token-weighted loss

The GPU side of supervised fine-tuning: why SFT batches are mostly padding, length grouping versus BFD packing, varlen attention with cu_seqlens and r…

Read article →
ARTICLE · 125

GPU Fine-Tuning Costs

Estimate fine-tuning spend before you launch: GPU-hour arithmetic, the bytes-per-parameter memory floor, LoRA and QLoRA as cost decisions, MFU, spot e…

Read article →
ARTICLE · 126

Fireworks AI, in depth: deployments, precision and KV memory, speculative decoding with Predicted Outputs, structured output and a measurement harness

How to run open-weight models on Fireworks AI well: model strings, firectl deployments with shapes, accelerators, FP8 and replica limits, the GPU memo…

Read article →
ARTICLE · 127

FlashAttention architecture

The GPU-side view of FlashAttention: the memory-traffic arithmetic that puts attention on the wrong side of the roofline, why a generic fusion compile…

Read article →
ARTICLE · 128

FlashInfer, in depth: paged-KV attention kernels, the plan and run split, load-balanced scheduling and how serving engines use it

How FlashInfer accelerates LLM serving on NVIDIA GPUs: why serving attention differs from training attention, the block-sparse page-table format with …

Read article →
ARTICLE · 129

FLUX on the GPU, in depth: tokens per image, the double-stream transformer, FLOP and memory budgets, offload, quantization and serving

How FLUX.1 image generation uses a GPU: the pipeline stages, how resolution becomes latent tokens, double-stream and single-stream transformer blocks,…

Read article →
ARTICLE · 130

PyTorch FSDP, in depth: sharding parameters, gradients and optimizer state with fully_shard

PyTorch Fully Sharded Data Parallel explained from first principles: the per-block all-gather and reduce-scatter lifecycle, memory arithmetic for a 7B…

Read article →
ARTICLE · 131

NVIDIA GB200, in depth: the Grace Blackwell superchip, the NVL72 NVLink domain and how software uses them

How the NVIDIA GB200 is built and how training and inference software uses it: the Grace CPU and two Blackwell GPUs on NVLink-C2C, HBM3E versus LPDDR5…

Read article →
ARTICLE · 132

NVIDIA GH200, in depth: the Grace Hopper superchip, NVLink-C2C coherence and how software uses a CPU memory tier

What is on the GH200 module, how ATS and NVLink-C2C give the GPU coherent access to CPU memory, first touch and migration, a malloc-based CUDA example…

Read article →
ARTICLE · 133

GPU Sharing Strategies Overview

A decision matrix for GPU sharing: MPS, time-slicing, MIG and vGPU scored on isolation strength, achievable utilization, blast radius, failure contain…

Read article →
ARTICLE · 134

GPUDirect

A map of the GPUDirect family: which path a transfer actually takes -- P2P between GPUs, RDMA to the NIC, Storage from NVMe, or Async -- what each one…

Read article →
ARTICLE · 135

GPUDirect Storage architecture

Deep-dive on GPUDirect Storage: the cuFile API and kernel driver, peer-to-peer PCIe DMA from NVMe and NVMe-oF into GPU HBM, the CPU bounce buffer it e…

Read article →
ARTICLE · 136

GRPO, in depth: the group-relative objective from scratch, where the GPU time and memory go, and the Dr. GRPO and DAPO fixes

Group Relative Policy Optimization for LLMs, implemented and budgeted: the objective with group-normalised advantages, a PyTorch loss with three aggre…

Read article →
ARTICLE · 137

NVIDIA H100 Architecture: Hopper GPU Explained

NVIDIA H100 Hopper architecture explained: the GH100 die and SMs, 80 GB HBM3, L2, NVLink 4, SXM5 vs PCIe, the FP8 Transformer Engine, MIG, and the A10…

Read article →
ARTICLE · 138

NVIDIA H100 Architecture: Inside the Hopper SM, Tensor Cores

NVIDIA H100 (Hopper) GPU architecture at the SM level: four warp schedulers, register file limits, L1 and shared memory, tensor cores, TMA async copy,…

Read article →
ARTICLE · 139

NVIDIA H200, in depth: 141 GB of HBM3e, why memory sets LLM decode speed, KV-cache sizing, training gains and deployment pitfalls

What the NVIDIA H200 changes and what it keeps from the H100: 141 GB HBM3e at 4.8 TB/s with the same Hopper compute, roofline reasoning for prefill an…

Read article →
ARTICLE · 140

GPU Hardware Faults, in depth: how ECC errors, Xids, NVLink failures and silent corruption reach a training job, and how the job survives them

How GPU hardware faults show up to software: the failure arithmetic of large jobs, a taxonomy from correctable ECC to fallen-off-the-bus with NVIDIA&#…

Read article →
ARTICLE · 141

GPU HBM architecture

How GPU HBM actually works: 3D stacking and TSVs, channels and banks, row buffers and the activate/precharge cycle, refresh, ECC, and why achieved ban…

Read article →
ARTICLE · 142

HCCL, in depth: how Huawei Ascend collectives bootstrap, choose algorithms, use buffers and fail, with a PyTorch benchmark

Huawei HCCL explained for engineers training on Ascend NPUs: host NIC versus device RoCE links, rank tables and root info, the hierarchical algorithms…

Read article →
ARTICLE · 143

Hugging Face Ecosystem, in depth: one checkpoint from Hub to GPU and back, through safetensors, device maps, attention kernels, adapters and serving

Follow a model through the Hugging Face stack on GPUs: pinned downloads, the safetensors format, meta-device loading and device_map placement, 4-bit l…

Read article →
ARTICLE · 144

GPU Hyperparameter Tuning, in depth: systems knobs vs optimisation knobs, batch and learning-rate coupling, precision traps and budget-fair search

Tune training hyperparameters the GPU-aware way: separate micro-batch, accumulation and checkpointing from the learning rate, benchmark throughput and…

Read article →
ARTICLE · 145

Image Classification Serving, in depth: GPU decode, TensorRT engines, Triton ensembles, dynamic batching and fleet sizing

How to serve image classifiers on GPUs: where the time goes, nvJPEG and DALI preprocessing, ONNX to TensorRT FP16 and INT8 engines, a Triton ensemble …

Read article →
ARTICLE · 146

Immersion Cooling for GPU, in depth: single- and two-phase physics, sizing fluid flow, what changes inside the server, GPU telemetry and scheduler integration

Immersion cooling for GPU clusters from first principles: single-phase versus two-phase, a worked flow-sizing example, server changes for fans, heatsi…

Read article →
ARTICLE · 147

GPU Incremental Training, in depth: continuing from checkpoints, LR re-warming, replay, sharded state and forgetting

How to add data to a trained model without retraining from scratch on GPUs: resume versus continual pre-training versus incremental fine-tuning, what …

Read article →
ARTICLE · 148

PyTorch Inductor, in depth: lowering, the loop-level IR, fusion scheduling, Triton code generation, autotuning and AOT packaging

How PyTorch Inductor turns an ATen graph into GPU kernels: decompositions, the Pointwise and Reduction IR, vertical and horizontal fusion, generated T…

Read article →
ARTICLE · 149

GPU Inference Latency, in depth: TTFT, TPOT and end-to-end time from first principles, with a runnable latency model

Where the milliseconds of an LLM request go on a GPU: TTFT, TPOT and end-to-end latency defined, prefill as a compute floor, decode as a memory-bandwi…

Read article →
ARTICLE · 150

LLM Inference Optimization Overview

A triage guide to LLM inference optimization on GPUs: the levers in the order you should try them, from measuring whether you are prefill- or decode-b…

Read article →
ARTICLE · 151

InfiniBand + NVLink architecture

How the scale-out fabric works for GPU training: RDMA and kernel bypass, GPUDirect RDMA into HBM, InfiniBand HCAs, switches and the subnet manager, fa…

Read article →
ARTICLE · 152

GPU Infrastructure Planning, in depth: from workload demand to GPU count, fabric, storage bandwidth and a failure budget

How to plan a GPU cluster from the workload down: training compute from 6ND and MFU, inference capacity from measured throughput and headroom, memory …

Read article →
ARTICLE · 153

GPU Infrastructure Cost, in depth: building the fully loaded cost of a useful GPU-hour from capital, power, facility and people

How to compute what a GPU-hour really costs on owned or colocated infrastructure: capital recovery with cost of capital, fabric and storage, power tim…

Read article →
ARTICLE · 154

Intel Gaudi + GPU Max, in depth: two different architectures, how PyTorch drives each, porting from CUDA and running existing fleets

A practical guide to Intel's two data-center AI accelerator lines: Gaudi's graph-compiled matrix engines, tensor processor cores and integra…

Read article →
ARTICLE · 155

Iterative DPO, in depth: on-policy rounds, pair building, the moving reference and scheduling GPUs between generation and training

Iterative DPO as a GPU workload: why one offline round goes stale, the generate-score-pair-train round loop, a worked budget where generation costs as…

Read article →
ARTICLE · 156

ITL, in depth: inter-token latency as a distribution, the stalls behind its tail, the p99 cliff and ITL SLOs

Inter-token latency for LLM serving: ITL versus TPOT, the decode-step floor, prefill interference, preemption, speculative bursts and proxy buffering,…

Read article →
ARTICLE · 157

GPU kernel fusion -- fewer kernels, less memory traffic

Deep-dive on GPU kernel fusion: the memory-bound HBM-traffic problem, fusing operations to keep intermediates on-chip (registers/SRAM), elementwise/ep…

Read article →
ARTICLE · 158

Kong AI Gateway, in depth: ai-proxy plugins, balancing across vLLM GPU pools, token budgets, guards, caching and failure modes

Running Kong Gateway as an AI gateway in front of self-hosted vLLM GPU pools and hosted LLM APIs: the request path through the ai-* plugin chain, ai-p…

Read article →
ARTICLE · 159

KServe for LLMs, in depth: InferenceService vs LLMInferenceService, model loading, KEDA autoscaling on engine metrics, multi-node serving and safe rollouts

How KServe serves LLMs on Kubernetes: the Hugging Face runtime with vLLM, OpenAI-style routes, LLMInferenceService for multi-node and disaggregated se…

Read article →
ARTICLE · 160

KTO Training, in depth: Kahneman-Tversky Optimization from thumbs-up data, the in-batch KL reference point, GPU cost, imbalance weights and reading the metrics

A practical guide to KTO (Kahneman-Tversky Optimization) for aligning LLMs with binary desirable/undesirable feedback: the loss and its prospect-theor…

Read article →
ARTICLE · 161

Kubeflow, in depth: Pipelines, Trainer v2 TrainJobs, Katib and Kueue on a shared GPU cluster

How Kubeflow turns Python into GPU pods: KFP components, artifacts and caching, Trainer v2 TrainJobs and runtimes, Kueue admission that prevents parti…

Read article →
ARTICLE · 162

Kueue, in depth: job-level admission, quota, cohort borrowing, preemption and readiness for GPU training on Kubernetes

How Kueue queues GPU jobs on Kubernetes: ResourceFlavor, ClusterQueue, LocalQueue and Workload, v1beta2 YAML, queueing strategies, cohort borrowing an…

Read article →
ARTICLE · 163

KV Cache Disk Offload, in depth: break-even arithmetic, chained chunk keys, NVMe layout, LMCache configuration and failure modes

When reading LLM KV cache from NVMe beats recomputing it: bytes per token, a break-even calculator, safe cache keys, O_DIRECT layout, a reference disk…

Read article →
ARTICLE · 164

KV Cache Sizing for Deployments, in depth: bytes per token, the per-GPU budget, tensor parallelism, concurrency and worked examples

How to size the KV cache when deploying an LLM: the bytes-per-token formula for MHA, GQA, MLA and sliding-window layers, how the per-GPU budget is lef…

Read article →
ARTICLE · 165

Liquid Cooling for GPU Datacenters, in depth: sizing the loop, the seconds between a pump fault and a throttled GPU, CDU telemetry and wiring cooling alarms into the scheduler

An operator's guide to liquid-cooled GPU clusters: the secondary loop as a system, heat-balance and flow sizing with worked numbers, capture rati…

Read article →
ARTICLE · 166

LiteLLM, in depth: one OpenAI-compatible gateway over vLLM GPU pools and hosted models, with routing, fallbacks, virtual-key budgets and supply-chain hygiene

LiteLLM as GPU-serving infrastructure: SDK versus proxy, a config for two vLLM pools and a hosted fallback, routing strategies, retries, cooldowns and…

Read article →
ARTICLE · 167

LLM A/B Testing, in depth: experimenting with quantisation, engines and GPU changes without fooling yourself

How to A/B test LLM serving changes on GPUs: user-level randomisation, per-arm replica pools that avoid batching interference, guardrail and quality m…

Read article →
ARTICLE · 168

LLM Active Learning, in depth: acquisition signals from logprobs, LoRA ensembles, GPU selection and the cost of rescoring the pool

Active learning for LLM tasks as a GPU workload: label margin and truncated top-k entropy, multi-LoRA disagreement, a vLLM scoring sketch, two-stage u…

Read article →
ARTICLE · 169

LLM Annotation Tools, in depth: task design, GPU pre-annotation, active learning, agreement and exporting SFT and DPO data

How human-in-the-loop annotation tools fit the GPU training pipeline: task shapes for LLM data, the six-part architecture, Label Studio, Argilla and P…

Read article →
ARTICLE · 170

LLM Bill of Materials, in depth: the five layers of a served model, runtime capture on a GPU node, CycloneDX output, diffs and fleet queries

What a bill of materials for a served LLM must contain, from system prompt and chat template to kernels, CUDA libraries, driver and GPU; capturing it …

Read article →
ARTICLE · 171

LLM Blue-Green Deployment, in depth: environment parity, an atomic switch, cold caches after cutover and the price of the rollback window

How to run blue-green releases for GPU-served LLMs: a hashed release manifest, parity checks for tokenizer, chat template and generation defaults, loa…

Read article →
ARTICLE · 172

LLM Bottleneck Analysis, in depth: naming the bound with MFU and MBU, reading the right counters, and confirming with perturbation tests

A procedure for diagnosing slow LLM inference and training: the five bounds (compute, memory bandwidth, communication, host launch, capacity), MFU and…

Read article →
ARTICLE · 173

LLM Canary Deployment

How to run a canary rollout for a model or inference-engine change: why the quality signal lags, sticky per-conversation traffic assignment, which met…

Read article →
ARTICLE · 174

LLM Capacity Analysis, in depth: goodput at SLO from load-test sweeps, naming the binding resource, model ceilings and an honest replica count

How to analyse the capacity of an LLM serving deployment from evidence: open-loop sweeps, goodput at SLO, finding the knee, vLLM saturation metrics, f…

Read article →
ARTICLE · 175

LLM Capacity Forecasting, in depth: forecasting tokens not requests, quantile peaks, backtesting and converting demand into replicas

How to forecast LLM serving capacity: measuring input and output tokens, decomposing growth and weekly shape, overlays for launches and mix shifts, a …

Read article →
ARTICLE · 176

LLM Change Management, in depth: release manifests, risk classes, gates and rollback for GPU-served models

Change management for GPU-served LLMs: the layers that change outputs and capacity, a content-addressed release manifest, automatic risk classificatio…

Read article →
ARTICLE · 177

LLM Chargeback, in depth: cost pools, billable units, budget rates, idle capacity policy, an append-only ledger and the monthly close

How to build internal chargeback for shared LLM GPU fleets: cost pools and overhead uplift, reserved GPU-hours versus weighted token units, budget rat…

Read article →
ARTICLE · 178

LLM CI/CD Pipeline, in depth: the serving release unit, GPU test tiers, numerical parity and performance gates

How to build CI/CD for LLM serving on GPUs: a release manifest pinning image digest, engine, CUDA and host driver, weights and serving config; CPU and…

Read article →
ARTICLE · 179

LLM Committed Use Discounts, in depth: commitment instruments, hourly matching, the quantile sizing rule, ladders and GPU-specific risks

How to buy and manage GPU commitment discounts for LLM fleets: resource-based and spend-based instruments on Google Cloud and AWS, how discounts are m…

Read article →
ARTICLE · 180

LLM Content Moderation, in depth: a GPU classifier cascade for user content, one-token logprob scoring, threshold selection, capacity math and policy backfills

How to moderate user-generated content with LLMs at platform scale: a three-stage cascade of embedding heads, an LLM safety classifier scored from one…

Read article →
ARTICLE · 181

LLM Cost Analysis, in depth: from a GPU-hour to cost per million tokens, per request and per break-even

How to turn a GPU-hour price into cost per million input and output tokens for LLM inference: bandwidth-bound decode, compute-bound prefill, KV cache …

Read article →
ARTICLE · 182

LLM Cost Attribution, in depth: splitting shared GPU steps among batched requests, KV memory-time, prefix cache credit and reconciliation

How to attribute GPU inference cost to individual requests under continuous batching: a calibrated step-time model, compute and KV memory shares, a co…

Read article →
ARTICLE · 183

LLM Data Curation Pipelines, in depth: running dedup, embedding and quality classifiers on GPUs at billion-document scale

How large-scale LLM data curation runs as a GPU job: stage placement, MinHash and LSH fuzzy deduplication with buckets-to-edges and connected componen…

Read article →
ARTICLE · 184

LLM Data Flywheel, in depth: production signals, preference pairs, GPU budgets per turn and the gates that keep the loop honest

How to run an LLM data flywheel as an engineered loop: which user signals to capture and how they are biased, an event schema with consent, building p…

Read article →
ARTICLE · 185

LLM Deployment Pattern Overview

The five LLM deployment topologies compared: dedicated single-tenant, pooled multi-tenant, serverless scale-to-zero, on-prem and hybrid burst. What is…

Read article →
ARTICLE · 186

LLM Distillation Data, in depth: the teacher scoring pass as a GPU workload

How to produce token-level distillation data on GPUs: why teacher scoring is a prefill-only, compute-bound job, the logits memory wall and chunked on-…

Read article →
ARTICLE · 187

LLM Disaster Recovery Drill, in depth: per-asset RTO and RPO, the weights and GPU bottleneck, restore validation and a scripted drill

How to run disaster recovery drills for LLM serving: scenarios from regional loss to account compromise, per-asset RTO and RPO, weight transfer arithm…

Read article →
ARTICLE · 188

LLM Error Budgets, in depth: budget arithmetic, a written policy and cause attribution for GPU serving fleets

How to run error budgets for GPU-backed LLM serving: request versus token weighting, availability, latency and quality budgets, a three-state policy, …

Read article →
ARTICLE · 189

LLM FinOps

Cost attribution for a shared GPU fleet: the GPU-hour to cost-per-1k-tokens unit model and why utilization sits in the denominator, the request-time m…

Read article →
ARTICLE · 190

Flame Graphs for LLMs, in depth: CPU, wall and GPU-weighted stacks, folding Kineto traces, and finding hidden syncs in decode

Build flame graphs that answer LLM questions: py-spy for host and wait time, PyTorch export_stacks and a correlation-based trace fold for GPU time, a …

Read article →
ARTICLE · 191

LLM Gameday, in depth: rehearsing GPU serving and training failures with hypotheses, safe fault injection and a scorecard

How to run gamedays for GPU-backed LLM systems: a failure catalogue, steady-state hypotheses as code, a replica-loss scenario with capacity and cold-s…

Read article →
ARTICLE · 192

GitOps for LLM Deployments, in depth: pinning the release tuple, moving weights, slow-start health checks, drift and rollback

Run LLM model servers with Argo CD or Flux: a deploy repo layout, pinning engine digest and model revision, prefetching weights before rollout, startu…

Read article →
ARTICLE · 193

LLM Guardrails

LLM guardrails viewed as a serving component rather than a policy document: where to place the input and output classifiers, what each placement costs…

Read article →
ARTICLE · 194

LLM System Health Scoring, in depth: turning GPU, runtime and service signals into scores that route traffic, place jobs and trigger repair

How to design a health score for LLM serving replicas and training nodes: deriving it from the decisions it drives, signal categories, normalisation w…

Read article →
ARTICLE · 195

Helm for LLM Deployments, in depth: model profiles and schemas, GPU-aware templates, probes sized to load time, weights outside the chart and safe upgrades

How to build a Helm chart for LLM serving: what a release stores, model profiles with a values schema and fail guards, a vLLM Deployment template with…

Read article →
ARTICLE · 196

LLM Human-in-the-Loop, in depth: confidence signals from the serving stack, review routing, reviewer capacity and turning decisions into training data

How to run human review for LLM serving: logprob, verifier and self-consistency signals and their GPU cost, calibrated routing with an audit sample, r…

Read article →
ARTICLE · 197

HITL Review Patterns, in depth: pre-delivery gates, judge pre-screens, stratified audits, reviewer agreement and shadow mode on a GPU serving stack

A catalogue of human review patterns for LLM output and what each costs the serving stack: buffered gates and regeneration with prefix caching, batche…

Read article →
ARTICLE · 198

LLM Incident Communication Channels, in depth: channel topology, telemetry snapshots and model-fallback notices for GPU serving incidents

How to run communication for LLM serving incidents on GPU fleets: a channel topology including upstream providers and API consumers, a declare bot and…

Read article →
ARTICLE · 199

LLM KPI Dashboards, in depth: a four-layer metric tree from cost per token down to SM activity, with PromQL, recording rules and the traps that make panels lie

How to build an LLM serving KPI dashboard that answers questions instead of decorating a wall: a four-layer metric tree from business KPIs to service …

Read article →
ARTICLE · 200

LLM Individual Contributor KPIs, in depth: MFU with the attention term, goodput, cost per token at the SLO, and guard metrics that stop gaming

The few metrics an engineer on an LLM training, inference or eval team should own: MFU and HFU derived from first principles, goodput, cost per millio…

Read article →
ARTICLE · 201

LLM Leadership KPIs, in depth: a six-number GPU fleet scorecard, ratio-of-sums rollups, commitment coverage and decision thresholds

The KPIs a platform leader needs to run an LLM GPU fleet: cost per qualified task, busy fraction of paid GPU-hours, commitment coverage, SLO attainmen…

Read article →
ARTICLE · 202

LLM North Star Metric, in depth: qualified successful tasks, the factor tree that links model, serving and GPU capacity, and the counter-metrics that stop gaming

How to choose and govern one North Star metric for a GPU-served LLM product: why tokens and utilisation fail, a precise definition of qualified succes…

Read article →
ARTICLE · 203

Kustomize for LLM Deployments, in depth: GPU components, patching server args safely, generators and pinned images, and render-time checks

Kustomize for GPU inference fleets: base, components and overlays, why strategic merge replaces vLLM args, JSON 6902 appends, label selector immutabil…

Read article →
ARTICLE · 204

LLM Data Labeling, in depth: prefill-bound batch inference, prefix caching, constrained labels with logprobs, calibration and cascades

Labeling data with LLMs as a GPU workload: why it is prefill-bound, prefix caching, constrained single-token labels with vLLM structured outputs and l…

Read article →
ARTICLE · 205

LLM Load Testing

How to load test a token-streaming LLM service correctly: why closed-loop generators can never overload a server, Little's Law and coordinated om…

Read article →
ARTICLE · 206

LLM Serving Math 101, in depth: five numbers, a Little's law calculator and a full Llama 3.1 70B sizing example

The back-of-envelope math of LLM serving from first principles: weight bytes, KV bytes per token, prefill FLOPs, decode bytes per step and Little&#x27…

Read article →
ARTICLE · 207

LLM Memory Profiling, in depth: predict the GPU budget, measure each phase, and attribute every gigabyte before the OOM

A practical method for LLM GPU memory profiling: allocated vs reserved vs device memory, a five-term budget for training and inference, per-phase peak…

Read article →
ARTICLE · 208

LLM Metrics Deep Dive, in depth: TTFT, ITL and queue histograms, KV cache saturation, DCGM and cardinality for GPU serving

How to measure LLM serving in production: vLLM and OpenTelemetry GenAI metric names, TTFT versus ITL versus TPOT, correct histogram_quantile aggregati…

Read article →
ARTICLE · 209

NVIDIA Nsight for LLMs, in depth: NVTX for prefill and decode, capture windows, CUDA graphs, kernel families, decode GEMMs and tensor-parallel all-reduce

A working method for profiling LLM training and serving with Nsight Systems and Nsight Compute: annotating steps and layers with NVTX, capturing stead…

Read article →
ARTICLE · 210

LLM On-Call Playbook, in depth: the first fifteen minutes, a five-layer triage tree, a mitigation ladder and sustainable GPU serving shifts

A responder's playbook for LLM serving on GPUs: the first fifteen minutes, a five-layer triage tree from gateway to model output, a Prometheus tr…

Read article →
ARTICLE · 211

LLM Output Validation, in depth: constrained decoding on the GPU, token masks, retry economics and semantic validators

How LLM output validation works in a GPU serving stack: grammar-constrained decoding and token bitmasks, why mask computation overlaps the forward pas…

Read article →
ARTICLE · 212

LLM Ownership Matrix, in depth: a RACI over the GPU serving stack, ownership as code with one Accountable per row, and wiring it into reviews, alerts and incidents

How to build an ownership matrix for a self-hosted LLM serving stack: RACI rules, a matrix from weights to quotas, a YAML ownership file with a CI val…

Read article →
ARTICLE · 213

LLM On-Call Rotation, in depth: page-load arithmetic, rotation shapes, escalation and handoffs for GPU serving teams

Designing on-call for GPU LLM serving: what makes it different, sizing with page-load arithmetic, rotation shapes compared, a rotation generator and b…

Read article →
ARTICLE · 214

LLM Penetration Testing, in depth: assessing the self-hosted GPU inference stack below the prompt

A defensive penetration-testing plan for self-hosted LLM infrastructure: scoping and rules of engagement, mapping the serving stack, auditing exposed …

Read article →
ARTICLE · 215

LLM Performance Profiling, in depth: decode and prefill budgets, MBU and MFU, torch.profiler, Nsight timelines and a worked example

A top-down method for profiling LLM inference: compute bandwidth and FLOP floors for decode and prefill, measure TTFT and inter-token latency correctl…

Read article →
ARTICLE · 216

LLM Pricing Models, in depth: input and output meters, cache writes and reads, reasoning tokens, batch and flex tiers, length tiers and committed capacity

How LLM API prices are actually computed: why output costs more than input, cache write surcharges and read discounts with break-even hit rates, hidde…

Read article →
ARTICLE · 217

LLM Health Probes, in depth: startup, readiness and liveness for GPU inference servers without restart storms

How to design Kubernetes startup, readiness and liveness probes for LLM inference servers: sizing startup budgets from measured load times, keeping lo…

Read article →
ARTICLE · 218

LLM Prompt Injection Defense, in depth: serving-side controls with windowed scanners, spotlighting cost, constrained tool calls, quarantined pools and cache isolation

Prompt injection defences that live in the GPU serving stack: provenance tags, windowed Prompt Guard 2 scanning for its 512-token limit with capacity …

Read article →
ARTICLE · 219

LLM Python Frameworks, in depth: how LangChain, LlamaIndex and DSPy drive a self-hosted GPU server, and how to stop them wasting it

What Python LLM application frameworks make the GPU do: connecting LangChain, LlamaIndex and DSPy to an OpenAI-compatible vLLM or SGLang server, mappi…

Read article →
ARTICLE · 220

PyTorch Profiler for LLMs, in depth: schedules, key_averages, traces, memory snapshots and multi-GPU profiling

A working manual for torch.profiler on LLM training and inference: what it records, the step schedule, reading key_averages by self time and input sha…

Read article →
ARTICLE · 221

LLM Ranking and Comparison UI, in depth: blind side-by-side serving on GPUs, vote logging, Bradley-Terry ranks with intervals, style control and anti-gaming

How to build a blind side-by-side LLM comparison system: the four guarantees, concurrent generation with identical parameters and a parity buffer, GPU…

Read article →
ARTICLE · 222

LLM Performance Regression Analysis, in depth: noise floors, bootstrap comparisons, bisecting and kernel profile diffs

How to detect and localise LLM serving performance regressions: mapping TTFT and TPOT to prefill and decode, the usual causes from drivers to chat tem…

Read article →
ARTICLE · 223

LLM Reserved Capacity, in depth: capacity guarantees versus discounts, sizing from demand, filling the pool and surviving the end date

How reserved GPU capacity works for LLM training and serving: capacity reservations versus billing commitments, Capacity Blocks and calendar-mode rese…

Read article →
ARTICLE · 224

LLM Rollback Strategies, in depth: where rollback time goes on GPUs, the serving tuple, automatic triggers, adapter rollback and the cost of a warm pool

How to roll back an LLM serving release on GPUs: weight transfer, engine startup and cold prefix caches, rolling back the full serving tuple, an autom…

Read article →
ARTICLE · 225

LLM Routing Strategies, in depth: model-tier routing, cascades and KV-cache-aware replica routing

The two LLM routing problems and how to solve each: rules, classifiers, preference-trained routers such as RouteLLM and cascades for choosing a model;…

Read article →
ARTICLE · 226

LLM Runbook Index, in depth: a symptom-keyed catalogue for GPU fleets, alert wiring, CI lint and the entries that matter

How to build a runbook index for GPU LLM serving and training: runbooks as data in git, an entry schema, symptom-keyed catalogue with Xid 79, 48, 63, …

Read article →
ARTICLE · 227

GPU Serving Architecture for LLMs: The Full Stack

An architecture map of a GPU LLM serving stack: ingress and admission, the scheduler, the model executor and its tensor/pipeline layout, the KV cache …

Read article →
ARTICLE · 228

LLM Serving Stacks

How to choose an LLM serving engine: the axes that actually differentiate vLLM, TGI, TensorRT-LLM and SGLang - scheduling model, KV-cache management, …

Read article →
ARTICLE · 229

LLM Shadow Deployment, in depth: mirroring live traffic to a candidate model without side effects, wasted GPUs or misleading diffs

How to shadow-deploy an LLM: gateway mirroring, inert tool calls, comparing nondeterministic outputs, cache and batching confounders on GPUs, a worked…

Read article →
ARTICLE · 230

LLM SLO Burn Rate Alerts, in depth: TTFT and inter-token SLIs, multiwindow rules for GPU serving, and what really burns the budget

How to page on error-budget burn for LLM inference: defining good and bad events for streaming responses, TTFT and inter-token latency SLIs from vLLM …

Read article →
ARTICLE · 231

LLM Spend Alerts, in depth: gateway metering, multiwindow burn rates in dollars, idle GPU spend and an enforcement ladder

Designing spend alerts for LLM APIs and self-hosted GPU fleets: why billing data is too late, meter events with a versioned price table, threshold, bu…

Read article →
ARTICLE · 232

Python LLM Stack Overview, in depth: from pip install to GPU kernels, the version contract, memory arithmetic, fine-tuning and serving layers

A first-principles map of the Python LLM stack: how a PyTorch call becomes a GPU kernel, the driver and CUDA wheel version contract, what transformers…

Read article →
ARTICLE · 233

LLM Status Page, in depth: components customers can act on, mapping GPU and serving signals to status, quality degradation, and a feed that survives the outage

Designing a status page for GPU-backed LLM serving: components by surface, model and region, thresholds for errors, TTFT, capacity rejections and qual…

Read article →
ARTICLE · 234

Synthetic Data Generation for LLMs on GPUs, in depth: cost per accepted sample, parallel layout, offline vLLM batches, judging, dedup and resumable jobs

Running LLM synthetic data generation as a GPU workload: why decode and acceptance rate drive cost, a GPU-hours budget model, tensor-parallel versus r…

Read article →
ARTICLE · 235

LLM Synthetic Monitoring, in depth: outside-in probes for GPU serving, cache-proof prompts, correctness checks and quorum alerts

How to build synthetic monitoring for an LLM serving fleet: client-side TTFT and decode timing, probe prompt design that defeats prefix caching, a Pyt…

Read article →
ARTICLE · 236

LLM TCO, in depth: a three-year total cost of ownership model for API, cloud and owned GPU serving

How to build a three-year total cost of ownership model for an LLM workload: demand growth, measured GPU throughput, capacity for peak, API versus on-…

Read article →
ARTICLE · 237

LLM Trace Analysis, in depth: reading PyTorch profiler traces to find idle GPUs, exposed communication and launch overhead

How to capture and analyse PyTorch profiler (Kineto) traces of LLM training and inference: the trace format, a capture schedule that measures steady s…

Read article →
ARTICLE · 238

LLM Uptime Calculation, in depth: good-event definitions, three measurement methods, series and k-of-n composition, and GPU failure maths

How to calculate availability for an LLM endpoint: defining a good request with TTFT and output checks, time vs request vs probe availability, series,…

Read article →
ARTICLE · 239

LMDeploy, in depth: TurboMind, blocked KV cache sizing, AWQ and KV quantization, and serving it in production

How LMDeploy serves LLMs: TurboMind and PyTorch engines, persistent batching, the cache_max_entry_count share of free memory, worked KV sizing for an …

Read article →
ARTICLE · 240

Lookahead Decoding

Draft-model-free speculative decoding on the GPU: Jacobi-style parallel decoding, lookahead branches and the n-gram pool, prompt-lookup drafting that …

Read article →
ARTICLE · 241

LoRAX, in depth: serving hundreds of LoRA adapters on one base model with SGMV, exchange scheduling and tiered weight caching

How LoRAX serves many LoRA fine-tunes from one GPU deployment: router and per-adapter queues, heterogeneous continuous batching with SGMV kernels, ada…

Read article →
ARTICLE · 242

Megatron-LM, in depth: how tensor, pipeline, data and context parallelism compose, sizing a 70B run on 512 GPUs, and the flags that matter

Megatron-LM and Megatron-Core explained for practitioners: the parallel dimensions and their communication, the tp-cp-ep-dp-pp rank layout, a 512-GPU …

Read article →
ARTICLE · 243

GPU Memory Hierarchy

The GPU memory hierarchy as one system: register file, shared memory and L1, L2, and HBM. How capacity rises while bandwidth falls at each level, why …

Read article →
ARTICLE · 244

CUDA Memory Pool, in depth: the stream-ordered allocator, release thresholds, cross-stream reuse, CUDA graphs and PyTorch's caching allocator

How GPU memory pools work: why cudaMalloc and cudaFree are slow, the CUDA stream-ordered allocator with cudaMallocAsync and explicit pools, the releas…

Read article →
ARTICLE · 245

AMD MI300X, in depth: eight XCDs, private 4 MB L2 slices, a 256 MB Infinity Cache and how kernels should use them

MI300X from the kernel's point of view: the XCD and I/O die hierarchy, round-robin workgroup dispatch and a program-id remap for L2 reuse, wave64…

Read article →
ARTICLE · 246

AMD MI325X, in depth: 256 GB of HBM3E on MI300X compute, KV-cache planning, FNUZ FP8, partitioning and porting traps

What the AMD Instinct MI325X changes for software: the same CDNA 3 compute as MI300X with 256 GB HBM3E at 6 TB/s and 1,000 W, a worked KV-cache and de…

Read article →
ARTICLE · 247

NVIDIA MIG architecture

How Multi-Instance GPU partitions a datacenter GPU: SM slices, L2 slices and the memory-controller path that makes isolation a hardware property rathe…

Read article →
ARTICLE · 248

Mixed Precision Training, in depth: autocast, GradScaler, BF16 versus FP16 on GPUs, FSDP policies and debugging NaNs

How mixed precision training actually runs on NVIDIA GPUs with PyTorch: what autocast casts and caches, when you need GradScaler, choosing FP16, BF16 …

Read article →
ARTICLE · 249

ML Framework Comparison, in depth: PyTorch, JAX and TensorFlow by execution model, compilation and scaling

PyTorch, JAX and TensorFlow compared the way that matters on GPUs: eager dispatch versus traced functions, torch.compile versus jax.jit versus tf.func…

Read article →
ARTICLE · 250

GPU ML Ops, in depth: running a GPU fleet as a control loop, from node lifecycle and failure triage to checkpoint intervals and goodput

How to operate a GPU fleet for training and serving: the node lifecycle state machine, burn-in gates, classifying faults into drain or ignore, the ari…

Read article →
ARTICLE · 251

ML Reproducibility on GPUs, in depth: floating-point order, PyTorch determinism controls, seeded data pipelines, distributed reductions and a bitwise replay test

Why identical GPU training runs diverge and how to stop it: non-associative float sums, atomics, cuDNN autotuning, TF32, seeded DataLoaders, RNG state…

Read article →
ARTICLE · 252

MLflow for GPU Training Ops, in depth: server topology, rank-0 logging, metric volume, system metrics, checkpoints and the model registry

How to run MLflow underneath multi-GPU and multi-node training without slowing it down or losing data: tracking server topology, which rank logs, why …

Read article →
ARTICLE · 253

Multimodal LLM Training, in depth: the encoder-connector-decoder pipeline, staged freezing, packing and the GPU load-balancing problem

How vision-language models are trained on GPUs: vision encoder, connector and LLM decoder, how many tokens an image becomes at fixed and dynamic resol…

Read article →
ARTICLE · 254

MLX, in depth: unified-memory arrays, lazy evaluation, composable transforms, quantisation, mlx-lm fine-tuning and custom Metal kernels

How Apple MLX works and how to use it well: arrays without devices, lazy evaluation and mx.eval, grad, vmap and compile, a training loop, affine quant…

Read article →
ARTICLE · 255

MoE All-to-All Communication

The MoE dispatch and combine all-to-all as a communication primitive: why every expert-parallel layer pays for two of them, what each one moves, why t…

Read article →
ARTICLE · 256

MoE Expert Parallelism Deployment, in depth: sizing, multi-node launch, network prerequisites, failure domains and rollout

How to deploy a large mixture-of-experts model with wide expert parallelism across nodes: sizing experts and KV memory per GPU, the divisibility rule …

Read article →
ARTICLE · 257

Grouped GEMM for MoE, in depth: jagged expert batches in one launch, persistent tile scheduling, alignment, the backward pass and choosing a kernel

How grouped GEMM runs every MoE expert's matmul in a single kernel: offsets and problem descriptors, persistent tile scheduling, alignment and ti…

Read article →
ARTICLE · 258

MoE Kernels, in depth: the six kernels of an expert layer, block alignment, fused expert GEMMs, the decode memory wall and profiling

How a mixture-of-experts layer actually runs on a GPU: fused router top-k, block-aligned sorting with padding, gathers folded into a grouped expert GE…

Read article →
ARTICLE · 259

MoE Load Balancing, in depth: measuring imbalance, capacity factors, auxiliary losses, DeepSeek-V3's bias controller and expert replication

How mixture-of-experts load balancing works in practice: why routing collapses, how to measure imbalance per expert and per GPU, capacity factors and …

Read article →
ARTICLE · 260

MoE Routing Math, in depth: scoring recipes, where the gradient flows, capacity and drop arithmetic, balance losses and the dispatch index math on the GPU

The arithmetic of mixture-of-experts routing as GPUs execute it: softmax and sigmoid scoring in Switch, Mixtral and DeepSeek-V3, gradients through gat…

Read article →
ARTICLE · 261

Mixture-of-Experts Serving Architecture in Depth

How mixture-of-experts models actually behave on serving GPUs: expert placement and expert parallelism, the dispatch and combine all-to-all that domin…

Read article →
ARTICLE · 262

MSCCL, in depth: custom GPU collective algorithms with MSCCLang, the XML IR and the selection rules that decide whether they run

How MSCCL runs custom collective schedules on top of NCCL: MSCCLang programs, the compiler and XML IR, synthesis, the runtime selection predicate behi…

Read article →
ARTICLE · 263

Multi-Instance GPU Deployment, in depth: profiles, mig-parted geometries, Kubernetes strategies, bin-packing, monitoring and safe reconfiguration

How to deploy NVIDIA MIG in production: GPU and compute instances, reading the driver's profile tables, manual setup with nvidia-smi, declarative…

Read article →
ARTICLE · 264

Multi-LoRA Serving, in depth: what the GPU does when one batch carries many adapters, gathered shrink and expand kernels, adapter memory arithmetic and the decode cost model

A GPU-level guide to serving many LoRA adapters on one base model: the unmerged LoRA computation, why grouping by adapter fails, gathered shrink and e…

Read article →
ARTICLE · 265

Multi-Provider LLM Strategy, in depth: task-named APIs, adapters, safe failover, capacity arithmetic and self-hosted GPUs as a provider

How to run LLM workloads across several providers and your own GPUs: goals and their designs, a task-named internal API, adapters and circuit breakers…

Read article →
ARTICLE · 266

Multi-Stream Execution, in depth: stream semantics, events, copy-compute overlap, PyTorch side streams, allocator hazards and proving overlap in a trace

A practical guide to CUDA multi-stream execution: what streams guarantee, ordering with events, a double-buffered CUDA pipeline, a PyTorch prefetcher …

Read article →
ARTICLE · 267

Multi-Turn KV Cache, in depth: keeping a conversation's attention state alive across turns, think time, routing and prompt rendering

How multi-turn KV cache reuse works in LLM serving: why each chat turn re-sends the whole history, the prefill arithmetic of reuse versus recompute, K…

Read article →
ARTICLE · 268

Multimodal Model Serving, in depth: the request path, media token accounting, CPU preprocessing, encoder scheduling and caching, KV budgets and encoder disaggregation

How to serve vision and audio language models in production: the request path, computing image and video token costs, safe CPU preprocessing, encoder …

Read article →
ARTICLE · 269

NCCL Collectives Explained: All-Reduce, Ring vs Tree, Topology

How NCCL runs multi-GPU collectives: all-reduce and other ops, ring vs tree algorithms, NVLink, NVSwitch and InfiniBand topology, GPUDirect RDMA, tuni…

Read article →
ARTICLE · 270

NVIDIA Nemotron, in depth: lineage, hybrid Mamba-Transformer MoE layers, KV memory math, FP8 and NVFP4, pruning and serving

A technical guide to NVIDIA Nemotron models: the lineage from Nemotron-4 340B to Nemotron 3, the hybrid Mamba-Transformer MoE layer pattern, KV cache …

Read article →
ARTICLE · 271

nvidia-smi

How nvidia-smi provides quick GPU status, and the key views for debugging.

Read article →
ARTICLE · 272

NVLink and NVSwitch Explained: The Multi-GPU Scale-Up Fabric

How NVLink and NVSwitch connect GPUs: memory-semantic links, the NVSwitch crossbar, peer-to-peer access, NVLink vs PCIe bandwidth, topology and collec…

Read article →
ARTICLE · 273

NVLink Switch, in depth: rack-scale NVLink domains, in-switch reduction and the software that makes them work

How the NVLink Switch turns NVLink into a rack-scale fabric: the GB200 NVL72 wiring of 72 GPUs to 18 switch chips, NVLink SHARP and NCCL's NVLS a…

Read article →
ARTICLE · 274

Object Detection Serving, in depth: GPU decode, letterboxing, TensorRT engines with NMS, Triton dynamic batching and capacity planning

How to serve object detectors on GPUs: what YOLO-style heads output, confidence filtering and NMS, letterbox coordinate mapping, GPU decode with nvJPE…

Read article →
ARTICLE · 275

GPU Occupancy: Warps per SM, Registers, Shared Memory

What GPU occupancy measures: resident warps per SM against the maximum, the register, shared memory and block-slot limits, a worked example, when low …

Read article →
ARTICLE · 276

Ollama on the GPU, in depth: VRAM budgets, layer offload, multi-GPU placement, KV cache precision and measuring decode speed

How Ollama uses a GPU: the VRAM budget of weights, KV cache and compute buffers, why a partial CPU offload collapses decode speed, how models are plac…

Read article →
ARTICLE · 277

Online DPO, in depth: on-policy pairs every step, reward models and LLM judges in the loop, vLLM colocation, weight sync and the GPU budget

Online DPO (online AI feedback) on GPUs: how it differs from offline and iterative DPO, the loss with a worked number, a minimal PyTorch step, reward …

Read article →
ARTICLE · 278

ONNX Runtime, in depth: execution providers, graph partitioning, IOBinding, CUDA graphs, TensorRT and quantization

How ONNX Runtime executes a model: graph optimisation levels, how execution providers claim nodes, memcpy at device boundaries, CUDA EP options, IOBin…

Read article →
ARTICLE · 279

OpenAI Realtime API, in depth: transports, events, turn detection, barge-in and tool calls for voice agents

How the OpenAI Realtime API works: WebRTC, WebSocket and SIP transports, ephemeral client secrets, the session and event model, server and semantic VA…

Read article →
ARTICLE · 280

OpenRouter, in depth: model and provider routing, the provider object, fallbacks, streaming errors and production operation

How OpenRouter routes a request to a model and a GPU provider, the documented provider fields for data policy, precision and price, fallback behaviour…

Read article →
ARTICLE · 281

OpenVINO for LLMs, in depth: export and weight compression, a memory budget for an 8B model, and running on Intel CPUs, GPUs and NPUs

How OpenVINO runs language models: Optimum Intel export, stateful IR, int4 compression flags, an 8B memory and bandwidth budget, LLMPipeline code, CPU…

Read article →
ARTICLE · 282

Hugging Face Optimum, in depth: the package split, what the ONNX exporter does, O1-O4 graph optimization, int8 quantization and when not to use it

A practical guide to Hugging Face Optimum: which package targets which hardware, how the ONNX exporter infers tasks and validates graphs, ORTModel and…

Read article →
ARTICLE · 283

GPU Orchestration Overview, in depth: inventory, quota, gang and topology-aware placement, failure handling, and Kubernetes versus Slurm

How GPU clusters decide which jobs run where: device plugins and Kubernetes DRA, Kueue and Slurm gang admission, topology-aware placement with tested …

Read article →
ARTICLE · 284

ORPO, in depth: the odds-ratio objective, one training stage without a reference model, and what it costs on the GPU

Odds Ratio Preference Optimization from first principles: how ORPO folds preference alignment into supervised fine-tuning, the loss and its gradient, …

Read article →
ARTICLE · 285

P2P KV Cache Transfer

How the KV cache actually moves from a prefill GPU to a decode GPU: bytes-per-token arithmetic under GQA, layer-wise streaming that overlaps transfer …

Read article →
ARTICLE · 286

GPU P2P Memory Access, in depth: enabling peers, copies versus direct loads, cross-device ordering, CUDA IPC and proving the fast path

How GPU peer-to-peer memory access works in software: UVA and peer mappings, cudaDeviceEnablePeerAccess per direction, cudaMemcpyPeerAsync versus dire…

Read article →
ARTICLE · 287

Page-Locked (Pinned) Host Memory, in depth: why DMA needs it, the CUDA allocation and registration APIs, PyTorch pin_memory and non_blocking, and pinning without starving the node

How pinned host memory works and how to use it: why DMA cannot read pageable pages, cudaHostAlloc and cudaHostRegister flags, mapped and write-combine…

Read article →
ARTICLE · 288

Paged KV cache architecture

Paged KV cache as GPU memory management: why contiguous per-sequence allocation fragments HBM, how fixed-size blocks and a block-table indirection fix…

Read article →
ARTICLE · 289

PCIe host interconnect architecture

How the PCIe host interconnect really behaves for GPUs: lanes and generation scaling, the host-to-device transfer path, why pinned memory matters, cop…

Read article →
ARTICLE · 290

Prefill/Decode Disaggregation Architecture in Depth

Why LLM prefill and decode have opposite hardware profiles, how they interfere when they share a GPU, what splitting them into separate pools actually…

Read article →
ARTICLE · 291

Pipeline Parallelism: GPipe vs 1F1B and the Pipeline Bubble

How pipeline parallelism trains big models across GPUs: layer stages, microbatches, GPipe vs 1F1B schedules, the bubble fraction, activation stashing,…

Read article →
ARTICLE · 292

GPU Datacenter Placement Strategy, in depth: matching training and inference workloads to sites by power, latency, data and risk

How to decide where GPU capacity should live and which workloads go where: workload classes and their constraints, power availability as the gating fa…

Read article →
ARTICLE · 293

GPU Pod Network, in depth: scale-up and scale-out domains, rail-optimised fabrics, mapping parallelism and pod bring-up

How a GPU pod network is designed and validated: the four networks in a pod, NVLink scale-up versus 400 Gb/s scale-out bandwidth, rail-optimised fat-t…

Read article →
ARTICLE · 294

Portkey AI Gateway, in depth: config trees, fallback and load-balance semantics, retries, guardrail status codes and routing to self-hosted GPU pools

How the open-source Portkey AI Gateway evaluates its routing config: strategy modes, inheritance, the exact fallback stop condition, weighted load bal…

Read article →
ARTICLE · 295

PPO for LLMs, in depth: implementing the training step, from tensor masks and per-token rewards to the diagnostics that tell you a run is going wrong

An implementer's guide to PPO for language models on GPUs: tensor shapes and response masks, per-token KL rewards with the score on the last toke…

Read article →
ARTICLE · 296

GPU Preemption Handling

Taking work back from a GPU that is already running it: request-level preemption inside an inference server, victim selection and the swap-versus-reco…

Read article →
ARTICLE · 297

Prefill vs Decode Split, in depth: measuring each phase, a fitted decode cost model, and choosing a split from your own traffic

How to decide between colocated, chunked and disaggregated LLM serving from evidence: an open-loop streaming harness for TTFT and inter-token latency,…

Read article →
ARTICLE · 298

Prefix Caching in Depth: What the GPU Skips, What It Still Pays For, and How Cached KV Blocks Live and Die in HBM

Prefix caching from the GPU's point of view: prefill FLOPs and what a cache hit removes, the full-block rule, reference-counted KV blocks and the…

Read article →
ARTICLE · 299

GPU Profiling

How the Nsight tools actually work: what each Nsight Systems timeline row records, how correlation links a CUDA API call to its kernel, reading gaps, …

Read article →
ARTICLE · 300

LLM Provider Failover, in depth: error classification, health-scored routing, deadline budgets, model equivalence, mid-stream failures and a warm backup

How to build failover across LLM providers that works under a real outage: classifying errors into retry, fail over or fail fast, per-target circuit b…

Read article →
ARTICLE · 301

PUE, in depth: meter boundaries, partial-load behaviour of GPU halls, pPUE and WUE, the reporting rules, and charging facility energy to a training job

Power usage effectiveness for GPU datacenters, measured properly: where the meters sit, interval versus annual PUE, why lightly loaded AI halls score …

Read article →
ARTICLE · 302

PyTorch Lightning, in depth: the module and Trainer split, hook order, FSDP strategies, memory arithmetic, checkpoints and Fabric

A practical guide to PyTorch Lightning 2.x for GPU training: what the LightningModule and Trainer each own, a full fine-tuning module, the hook order,…

Read article →
ARTICLE · 303

QPS Math for LLM Serving, in depth: from measured slots and service time to queueing tails, pooled routing, goodput and fleet size

Turn measured TTFT and TPOT into requests per second: slots and service time, Little's law, mixed workloads, Erlang C tail waits in code, pooled …

Read article →
ARTICLE · 304

Ray for GPU Workloads, in depth: tasks, actors, GPU scheduling, placement groups, Ray Data and Ray Train

How Ray runs GPU work from first principles: logical GPU resources and CUDA_VISIBLE_DEVICES, fractional GPUs, placement groups for gang scheduling, th…

Read article →
ARTICLE · 305

Ray Serve for LLM, in depth: from LLMConfig to GPU placement groups, replica sizing, multi-LoRA and failure modes

How Ray Serve LLM turns one LLMConfig into replicas on GPUs: placement group bundles and strategies for tensor and pipeline parallelism, a worked KV c…

Read article →
ARTICLE · 306

GPU Regulatory Landscape, in depth: how export controls measure chips, how AI laws measure training compute, and the compliance engineering that keeps a GPU team shippable

An engineer's guide to AI hardware regulation at the end of September 2026: 3A090 total processing performance and performance density, the Janua…

Read article →
ARTICLE · 307

Replicate, in depth: predictions, Cog packaging, deployments, signed webhooks, cold starts and cost control

How Replicate works as production infrastructure: models, versions and the prediction lifecycle, packaging with Cog, the three billing models and cold…

Read article →
ARTICLE · 308

Response Caching for LLM APIs, in depth: exact-match keys, determinism, streaming, invalidation and stampede control

How to build an application-side response cache in front of an LLM API: what must go into an exact-match cache key, request normalisation, the determi…

Read article →
ARTICLE · 309

GPU Pricing: Retail vs Cloud vs Enterprise, in depth: what each channel's price buys, licence and capacity constraints, hidden costs, and normalising to dollars per useful GPU-hour

How to compare GPU prices across retail cards, cloud instances and enterprise servers: what each price includes, the GeForce datacenter licence clause…

Read article →
ARTICLE · 310

Reward Model Training on GPU, in depth: the paired training step, last-token pooling, sizing, evaluation and the bugs that silently ruin a reward model

How to implement and run reward model training on GPUs: architecture, the pairwise loss with worked numbers, a PyTorch training step, the padding and …

Read article →
ARTICLE · 311

Ring Attention

Ring attention explained from the GPU side: sharding the sequence across devices, rotating KV blocks around a ring so every query sees every key, merg…

Read article →
ARTICLE · 312

RLHF Pipeline on GPU

RLHF and post-training as a GPU systems problem: holding a policy, reference, reward and critic model in one memory budget, the generation-versus-trai…

Read article →
ARTICLE · 313

RoCE (RDMA over Ethernet), in depth: verbs, queue-pair states, go-back-N retransmission, timeouts and debugging a hung all-reduce

How the RoCE reliable transport behaves under GPU training traffic: verbs objects and GPU memory registration, the RC queue-pair state machine with re…

Read article →
ARTICLE · 314

RoCE v2, in depth: running GPU training traffic over routed Ethernet with GID selection, PFC, ECN, DCQCN and NCCL tuning

How RDMA over Converged Ethernet v2 carries GPU collectives: the UDP 4791 encapsulation and ECMP entropy, GID selection, lossless versus lossy designs…

Read article →
ARTICLE · 315

Run:ai, in depth: quotas, over-quota borrowing, preemption, gang scheduling and fractional GPUs on Kubernetes

How NVIDIA Run:ai and the open-source KAI Scheduler share GPUs between teams: projects and queues, fair share, priorities, podgroups, fractions, a wor…

Read article →
ARTICLE · 316

GPU scheduling architecture

How work actually gets onto a GPU and shares it: the hardware work distributor and warp scheduler, stream priorities, why concurrent kernel execution …

Read article →
ARTICLE · 317

GPU Scheduling With Slurm

How Slurm schedules GPUs on an HPC cluster: GRES configuration, request flags, cgroup enforcement of CUDA_VISIBLE_DEVICES, affinity, partitions and Qo…

Read article →
ARTICLE · 318

Segmentation Model Serving, in depth: output arithmetic, GPU post-processing, mask encoding, promptable encoder caches and tiling

How to serve semantic, instance and promptable segmentation models on GPUs: output size arithmetic, GPU upsample and argmax, RLE and PNG mask encoding…

Read article →
ARTICLE · 319

GPU Selection in Depth: A Decision Method Built on Memory, Bandwidth, Compute and Cost per Unit of Work

How to choose a GPU for training or inference from first principles: memory-capacity math for weights, KV cache and optimizer state, why decode speed …

Read article →
ARTICLE · 320

Active-Passive LLM Serving, in depth: standby tiers, cold-start budgets, fenced promotion and in-flight streams

Active-passive LLM serving on GPUs: hot, warm, pilot-light and cold standby tiers, vLLM sleep mode, a stage-by-stage cold-start budget, generating pro…

Read article →
ARTICLE · 321

Anycast for LLM Serving, in depth: BGP catchments, long-lived token streams, the PoP-terminated front door, health-driven announcements and gentle traffic engineering

How to put an LLM API behind BGP anycast: catchments and ECMP, why long token streams break, anycast at PoPs with unicast GPU regions, resumable strea…

Read article →
ARTICLE · 322

GPU Serving Capacity Estimation, in depth: memory budgets, decode bandwidth, prefill cost and the latency-bounded request rate

How to estimate the request rate one GPU replica can serve for an LLM: KV pool and concurrency, decode step time from HBM bandwidth, prefill cost from…

Read article →
ARTICLE · 323

DNS Failover for LLM Serving, in depth: health-checked records, the real failover timeline, GPU-aware health probes and warm standby

How DNS failover works for LLM inference: failover record pairs, the detection, TTL and client-cache timeline, a cached synthetic-generation health pr…

Read article →
ARTICLE · 324

LLM Serving DR, in depth: DR tiers costed in GPUs, the release manifest and parity check, where standby GPUs come from, a degraded-mode ladder and the declare-disaster rule

Disaster recovery for LLM inference fleets: events to plan for, backup, pilot light, warm standby and active-active priced in GPUs, digest-based repli…

Read article →
ARTICLE · 325

Edge PoP LLM Serving, in depth: what the edge saves, small models and caches near users, capacity fragmentation and model rollout

How to serve LLM traffic from edge points of presence: TTFT versus decode latency, which work belongs at the edge, a vendor-neutral edge router with t…

Read article →
ARTICLE · 326

Geographic Routing for LLM Serving, in depth: residency filters, TTFT-based region scoring, cache affinity and spillover with hysteresis

How to choose a serving region per LLM request: why nearest-region routing fails under GPU queueing, the TTFT budget, residency and model filters, loa…

Read article →
ARTICLE · 327

Multi-Region LLM Serving

Why and when to run LLM inference in more than one region: latency, data residency, GPU capacity access and blast radius; the cost of idle standby acc…

Read article →
ARTICLE · 328

LLM Serving Reliability, in depth: failure domains, truthful health probes, preemption storms, retry budgets, graceful drains and stuck GPUs

How to keep LLM serving reliable: a failure-domain map, liveness versus readiness versus deep probes for vLLM, KV pressure and preemption alerts, per-…

Read article →
ARTICLE · 329

LLM Serving SLOs, in depth: choosing TTFT, TPOT, ITL and E2E targets, joint attainment versus percentiles, goodput and the load sweep

How to define LLM serving SLOs that match what users feel: per-class SLIs and targets, why marginal p90s overstate attainment (worked example), goodpu…

Read article →
ARTICLE · 330

SGLang, in depth: RadixAttention, cache-aware scheduling, structured decoding and running it in production

How SGLang serves LLM programs: the process architecture, the RadixAttention prefix tree and eviction, scheduling policies, the frontend language, con…

Read article →
ARTICLE · 331

GPU shared memory architecture

Deep-dive on GPU shared memory: why the memory hierarchy makes on-SM shared memory the linchpin of performance, the 32-bank structure and how warp-wid…

Read article →
ARTICLE · 332

Shared Prefix KV Caching, in depth: reading the prefix once per batch with cascade attention, the bandwidth arithmetic and when it pays

Shared-prefix attention for LLM decode: why prefix caching stores the prefix once but reads it per sequence, the log-sum-exp split behind Hydragen and…

Read article →
ARTICLE · 333

SLO-Aware Scheduling

Deadline-aware ordering of already-admitted LLM requests: deriving a per-request deadline from a TTFT or TPOT budget, earliest-deadline-first and why …

Read article →
ARTICLE · 334

S-LoRA, in depth: serving thousands of LoRA adapters with unified paging, adapter prefetch, early abort and a LoRA-aware tensor-parallel layout

How S-LoRA serves thousands of LoRA adapters on one base model: why unmerged serving wins, unified paging of adapters and KV cache with worked memory …

Read article →
ARTICLE · 335

Speculative Decoding Architecture in Depth

The GPU-execution view of speculative decoding: why decode is memory-bandwidth bound, how verifying k tokens raises arithmetic intensity and turns a G…

Read article →
ARTICLE · 336

GPU Spot Instances, in depth: how clouds sell and reclaim spare GPUs, the real price per useful hour, and making training and inference survive it

How AWS, Google Cloud and Azure sell spot GPU capacity and reclaim it, the notice signals and metadata endpoints, why GPU spot differs from CPU spot, …

Read article →
ARTICLE · 337

Stable Diffusion Training + Inference, in depth

Latent diffusion on GPUs: VAE, text encoders, UNet vs MMDiT, epsilon, v and rectified-flow objectives, a training step with min-SNR, memory budgets, g…

Read article →
ARTICLE · 338

Static Batching

Static batching on the GPU: why fixed-shape batches lose badly for online serving but still win for offline and batch jobs. The slowest-sequence probl…

Read article →
ARTICLE · 339

CUDA Stream Capture, in depth: recording stream work into graphs, fork and join, capture modes, invalidation, memory and graph updates

How CUDA stream capture turns stream code into a graph: what gets frozen, cross-stream fork and join, what Global, ThreadLocal and Relaxed modes reall…

Read article →
ARTICLE · 340

StreamingLLM

StreamingLLM keeps a transformer decoding forever in constant KV memory by pinning a handful of attention-sink tokens and rolling a recent window over…

Read article →
ARTICLE · 341

Streaming LLM Serving, in depth: SSE token delivery, TTFT and inter-token latency, cancellation, backpressure and proxies

How token streaming really works between the GPU and the user: TTFT versus inter-token latency, the SSE wire format, incremental detokenisation with s…

Read article →
ARTICLE · 342

GPU Supply Chain, in depth: from wafer to rack, where the bottlenecks sit, and how to engineer training around the capacity you actually get

How data-center GPUs are made and delivered: logic die, HBM, advanced packaging, substrates, modules, servers, networking and power; why yield compoun…

Read article →
ARTICLE · 343

Switch Transformer Architecture

The architectural lineage of sparse mixture-of-experts models: what GShard's top-2 routing established, why Switch argued top-1 is enough and wha…

Read article →
ARTICLE · 344

GPU Tail Latency Management, in depth: where p99 comes from in LLM serving, measuring it honestly and the levers that move it

Why tail latency dominates LLM and agent workloads, the queueing math, nine causes of slow requests on GPU serving replicas with their signatures, an …

Read article →
ARTICLE · 345

What Is a Tensor Core? GPU Tensor Core Architecture Explained

What GPU tensor cores are: warp-level matrix multiply-accumulate, WMMA and MMA, narrow inputs with wide accumulators, fp16, bf16, fp8, 2:4 sparsity.

Read article →
ARTICLE · 346

GPU Tensor Cores: How GEMM Tiling Keeps Them Fed

Why tensor cores make most GPU kernels memory-bound and how GEMMs keep them fed: arithmetic intensity, the warp tile hierarchy, wave quantization and …

Read article →
ARTICLE · 347

Tensor Parallelism in Depth: Splitting Transformer Layers Across GPUs, the Communication It Costs, and How to Run It

How tensor parallelism splits individual transformer layers across GPUs: the column and row matrix splits, Megatron's f and g operators, attentio…

Read article →
ARTICLE · 348

TensorRT-LLM, in depth: the PyTorch-based runtime, in-flight batching, paged KV cache, quantization, parallelism and how to run it

A practical guide to NVIDIA TensorRT-LLM as it works today: why the engine-build workflow is gone, how the scheduler and paged KV cache use GPU memory…

Read article →
ARTICLE · 349

TensorFlow Lite for LLM, in depth

How LLMs run on TensorFlow Lite, now LiteRT: flatbuffers and delegates, prefill and decode signatures with an explicit KV cache, conversion with liter…

Read article →
ARTICLE · 350

GPU Thermal Management, in depth: heat as power, the clock-management loop, reading throttle reasons, and why one hot GPU slows a whole training job

How GPU temperature turns into lost training throughput: dynamic and leakage power, the thermal resistance chain and its time constants, how firmware …

Read article →
ARTICLE · 351

GPU Throughput Math for LLM, in depth: three ceilings, decode batch curves, Little's law and cost per million tokens

A first-order throughput model for LLM serving and training: compute, bandwidth and capacity ceilings, why decode stays bandwidth-bound, a runnable es…

Read article →
ARTICLE · 352

GPU Time Slicing

GPU time slicing on Kubernetes from a cluster-operations point of view: how the device plugin advertises one physical GPU as N schedulable replicas, w…

Read article →
ARTICLE · 353

GPU Tensor Memory Accelerator (TMA): Architecture Deep-Dive

How Hopper's Tensor Memory Accelerator moves tensor tiles asynchronously via descriptors and mbarriers, freeing compute warps and enabling deep p…

Read article →
ARTICLE · 354

Together AI, in depth: serverless, dedicated endpoints, batch and fine-tuning, with integration code and a break-even model

An engineer's guide to Together AI: how serverless, dedicated endpoints, batch and fine-tuning map onto GPU economics, OpenAI-compatible client c…

Read article →
ARTICLE · 355

GPU Topology Awareness, in depth: reading the node graph, how NCCL detects it, binding processes to GPUs, CPUs and memory, and making schedulers respect it

How to make GPU jobs topology-aware: reading nvidia-smi topo -m and NVML, how NCCL detects and dumps topology and the variables that steer it, binding…

Read article →
ARTICLE · 356

torch.compile, in depth: Dynamo graph capture, guards and graph breaks, AOTAutograd, Inductor code generation, modes, dynamic shapes and compile-time caching

A first-principles guide to torch.compile: how TorchDynamo turns Python bytecode into FX graphs with guards, why graph breaks and recompilations happe…

Read article →
ARTICLE · 357

Google TPU

TPU as an architectural contrast to the GPU: what a systolic array does well and badly, XLA ahead-of-time compilation against static shapes versus CUD…

Read article →
ARTICLE · 358

Google TPU v5, in depth

Running training on Google TPU v5e and v5p in depth: accelerator naming and provisioning, one process per host, compile-once train steps, input pipeli…

Read article →
ARTICLE · 359

Google TPU v6 Trillium, in depth: the 256x256 MXU, shapes that fill it, collectives on a 16x16 torus, and measuring a JAX step

Working on Google TPU v6e (Trillium): documented chip and pod specifications, how the 256x256 MXU tiles matrix multiplications, a JAX benchmark and pr…

Read article →
ARTICLE · 360

LLM Training Checkpointing, in depth: choosing the interval, pricing failures, exact-resume state, restarting on a different GPU count and tiered retention

Checkpoint policy for large training runs: the Young and Daly interval, a waste model priced with Llama 3's interruption rate, every piece of sta…

Read article →
ARTICLE · 361

LLM Training Cluster Design, in depth: sizing GPUs, fabric, storage, power and restarts from the 6ND budget

Design an LLM training cluster from the job backwards: the 6ND compute budget and MFU, a worked 70B on 15T tokens example on 4,096 H100s, parallel lay…

Read article →
ARTICLE · 362

LLM Training Data Pipeline, in depth: deterministic token indexes, mixture blending, rank partitioning, exact resumption and replayable batches

How the online data pipeline for LLM pretraining works: memory-mapped token shards, packed sample indexes, deterministic mixture blending, data-parall…

Read article →
ARTICLE · 363

LLM Training Debugging, in depth: expected curves, a symptom triage table, the shrink-the-problem ladder, data and NaN localisation, gradient flow and distributed-only bugs

How to diagnose LLM training bugs: what a healthy curve predicts (ln V start loss), a symptom-to-suspect table, overfitting one batch, decoding real b…

Read article →
ARTICLE · 364

LLM Training Experiment Tracking, in depth: metrics, token axes, run lineage and spike alerts

How to track LLM training runs: a metric taxonomy with cadences, tokens-seen as the x-axis, rank-0 async logging without host syncs, run lineage acros…

Read article →
ARTICLE · 365

LLM Training Health Checks, in depth: preflight gates, in-loop numerics, loss-spike rollback, NCCL hang detection and rank-agreed responses

How to keep a distributed LLM training run healthy: DCGM and nccl-tests preflight gates, a matmul canary, global nonfinite checks, a robust loss-spike…

Read article →
ARTICLE · 366

Hyperparameter Tuning in Practice: Grid, Random, Bayesian, ASHA, and PBT, with Pseudocode

A practical guide to hyperparameter search on GPUs: what is worth tuning for neural network and LLM training, why random search beats grid, Bayesian o…

Read article →
ARTICLE · 367

LLM Training Orchestration, in depth: launch and rendezvous, preflight checks, hang detection, a restart supervisor with spares, and goodput

How to orchestrate a long LLM training run: Slurm plus torchrun launch, node preflight checks, detecting crashes, hangs and stragglers, a restart supe…

Read article →
ARTICLE · 368

LLM Training Reproducibility, in depth: run manifests, exact resume, data order, parallel layouts, seed variance and bisecting divergent runs

Run-level reproducibility for multi-GPU LLM training: levels of guarantee, a run manifest, checkpoints complete enough for exact resume, sample order …

Read article →
ARTICLE · 369

Anatomy of One GPU Training Step: Forward, Backward, Optimizer Update, and Where the Memory Goes, with Pseudocode

A timeline of one training step on a GPU: the data copy, the forward pass that saves activations, the backward pass that frees them and produces gradi…

Read article →
ARTICLE · 370

Triton Inference Server for LLMs, in depth: decoupled streaming, vLLM and TensorRT-LLM backends, KV cache sizing and multi-GPU modes

How to serve large language models on NVIDIA Triton Inference Server: why LLM backends own batching, decoupled streaming and the generate endpoints, a…

Read article →
ARTICLE · 371

Triton, in depth: the block programming model, the compiler pipeline, autotuning and debugging GPU kernels in Python

A first-principles guide to writing GPU kernels in Triton: programs and blocks instead of threads, masks and constexprs, a worked fused softmax kernel…

Read article →
ARTICLE · 372

TRL Deep Dive, in depth: trainers, data formats, memory and scaling

A practical deep dive into Hugging Face TRL: the config and trainer architecture, the v1.0 stable and experimental tiers, dataset formats, GPU memory …

Read article →
ARTICLE · 373

TTFT, in depth: a full time-to-first-token ledger with prefill FLOPs, the attention term, prefix caching, tensor parallelism and queueing

Time to first token from first principles: non-embedding prefill FLOPs, the causal attention term, the HBM floor, prefix-cache arithmetic, tensor-para…

Read article →
ARTICLE · 374

TTS (Text-to-Speech) Serving, in depth: time to first audio, real-time factor, chunked decoding, batching and barge-in

Running text-to-speech as a GPU service: TTFA and per-stream real-time factor as SLOs, streaming text segmentation, codec-token generation with contin…

Read article →
ARTICLE · 375

UCX, in depth: Unified Communication X layers, GPU data paths and debugging transport selection

How UCX moves data for MPI, UCC, NIXL and RAPIDS: UCP, UCT and UCS layers, workers, endpoints and progress, eager versus rendezvous, GPU memory paths …

Read article →
ARTICLE · 376

Unsloth, in depth: where fine-tuning memory goes, Triton kernels, offloaded checkpointing, an end-to-end QLoRA run and export paths

How Unsloth speeds up LLM fine-tuning: a worked memory budget for QLoRA on an 8B model, the Triton kernels, hand-derived LoRA backward and offloaded g…

Read article →
ARTICLE · 377

Vertex AI Gemini, in depth: capacity lanes, dynamic shared quota, Provisioned Throughput sizing and a per-request router

Gemini on Google Cloud (now the Gemini Enterprise Agent Platform) as a capacity system: Standard, Priority and Flex PayGo, batch inference and Provisi…

Read article →
ARTICLE · 378

NVIDIA vGPU

NVIDIA vGPU explained as a hypervisor-mediated sharing model: mediated passthrough and what the guest driver actually talks to, time-sliced engine sha…

Read article →
ARTICLE · 379

Video Diffusion Models, in depth: 3D VAEs, spacetime tokens, diffusion transformers and what quadratic attention over video does to GPU training and inference

How modern video diffusion models work and what they cost on GPUs: causal 3D VAE compression, patchified spacetime tokens, diffusion transformers with…

Read article →
ARTICLE · 380

Video Generation Serving, in depth: async jobs, shape buckets, sequence-parallel GPU groups and split decode stages

How to serve text-to-video diffusion models: an asynchronous idempotent job API, a token-count cost model by shape bucket, scheduling onto sequence-pa…

Read article →
ARTICLE · 381

Video LLM, in depth: frames, token budgets, encoder cost and KV memory on the GPU

How video language models turn a clip into tokens and what that costs on a GPU: decode, frame sampling, 3D patches and token merging, time-aware posit…

Read article →
ARTICLE · 382

Vision Encoders, in depth: patch math, token budgets, FLOPs by resolution, native-resolution packing and serving ViT, CLIP, SigLIP and DINOv2 on the GPU

Vision encoders as a GPU workload: patchify as a matmul, token count and FLOP formulas, ViT-L/14 worked from 224 to 672 pixels, tiling versus native r…

Read article →
ARTICLE · 383

Vision Transformer Serving, in depth: the quadratic crossover, fused attention, resolution buckets with CUDA graphs, token merging and quantization

How to serve a standalone ViT for classification, embeddings and dense prediction: a per-resolution cost model, where attention starts to dominate, SD…

Read article →
ARTICLE · 384

vLLM on GPU, in Depth: How the Engine Spends GPU Memory, Schedules Every Step, and What to Tune When It Misbehaves

Run vLLM on GPUs with intent: the process layout, how the startup memory profile becomes a KV-cache token budget (worked for Llama-3.1-8B and 70B), ho…

Read article →
ARTICLE · 385

Vision-Language Model Architectures, in depth: projectors, learned queries, cross-attention, early fusion and what each costs on a GPU

How vision-language models connect an image encoder to a language model: LLaVA projectors, Q-Former and Perceiver queries, Flamingo and Llama 3.2 cros…

Read article →
ARTICLE · 386

GPU vs CPU for Deep Learning: Latency Machines, Throughput Machines, and Why Training Runs on GPUs

Why deep-learning training runs on GPUs, worked from the workload: training-step FLOP counts, one roofline for a CPU and an H100, a time-to-train esti…

Read article →
ARTICLE · 387

GPU warp specialization architecture

Deep-dive on GPU warp specialization: splitting a thread block's warps into producer warps that drive cp.async/TMA copies and consumer warps that…

Read article →
ARTICLE · 388

WebAssembly for LLM Inference, in depth: SIMD kernels, threads, memory limits and when to move to WebGPU

How LLM inference runs in WebAssembly: why decode is memory-bandwidth bound, 128-bit SIMD int8 dot-product kernels, threads with SharedArrayBuffer and…

Read article →
ARTICLE · 389

WebGPU for LLM Inference, in depth: buffers and limits, f16 and subgroups, a quantised matvec kernel and the decode budget

How WebGPU runs language models in the browser: adapters, devices, buffers and pipelines, default limits that force weight sharding, shader-f16 and su…

Read article →
ARTICLE · 390

GPU Workload Types, in depth: classifying training, inference, rendering and HPC jobs by the resource they saturate

A practical taxonomy of GPU workloads by bottleneck: arithmetic intensity and the ridge point computed in code, why training and prefill are tensor-bo…

Read article →
ARTICLE · 391

ZeRO Optimizer, in depth: what each stage adds to a training step, a 7B memory budget, and running it with DeepSpeed

How to operate the ZeRO optimizer in practice: the collectives stages 1, 2 and 3 add to every step, a worked memory budget for a 7B model on eight GPU…

Read article →
ARTICLE · 392

ZeRO sharding architecture

Deep-dive on ZeRO: the 16-bytes-per-parameter memory math, stages 1/2/3, reduce-scatter and all-gather volume, prefetch overlap, hybrid sharding, ZeRO…

Read article →
ARTICLE · 393

Gradient Clipping, in depth: norm versus value, the AMP and accumulation order, global norms under sharding, Adam's moments and threshold choice

How gradient norm clipping really works: the clamped coefficient, where it sits relative to loss scaling and accumulation, computing the global norm u…

Read article →
ARTICLE · 394

Gradient Checkpointing (Activation Recomputation), in depth: activation memory arithmetic, full versus selective recompute, PyTorch and Megatron switches, and reading MFU afterwards

Activation recomputation from first principles: why activations dominate at long context, the sbh(34 + 5as/h) memory formula, a 7B worked example, sqr…

Read article →
ARTICLE · 395

Graphcore IPU, in depth: 1,472 tiles, bulk synchronous execution, Poplar and PopTorch, and the memory arithmetic that decides what fits

How the Graphcore GC200 IPU works and how software uses it: tiles with private SRAM, BSP compute-sync-exchange, a Poplar vertex example, PopTorch pipe…

Read article →
ARTICLE · 396

Grid Connection Challenges for AI, in depth: how a training site gets power, why connections take years, what utilities now require of large loads, and the software controls that make a cluster connectable

How AI data centers connect to the grid: load versus generator interconnection, the large-load study process, what training workloads do to the grid, …

Read article →
ARTICLE · 397

Groq LPU, in depth: SRAM-resident weights, the tensor streaming processor, compiler-scheduled determinism and sizing a 70B deployment

How Groq's LPU works and when to use it: why decode is bandwidth-bound, the tensor streaming processor layout, static compiler scheduling, multi-…

Read article →
ARTICLE · 398

H100 Availability History (2023-2025), in depth: from scarcity to commodity, quota versus capacity, and a procurement model for the next GPU generation

How H100 availability changed from 2023 to 2025 and what it means for engineers: a dated timeline, the three phases, why quota is not capacity, Capaci…

Read article →
ARTICLE · 399

H100-Hours per Foundation Model, in depth: reading published disclosures, back-solving utilisation, goodput and estimating your own run

What published GPU-hour figures for Llama 3.1, Llama 2 and DeepSeek-V3 actually count, how to back-solve delivered throughput and MFU from them, why g…

Read article →
ARTICLE · 400

H100 System Network Breakdown, in depth: the hop-by-hop bandwidth budget, which parallelism rides which link, and a ladder for finding the slow one

The network of an H100 system broken into hops: HBM, NVLink and NVSwitch, PCIe Gen5, ConnectX-7 rails, leaf, spine and storage. Per-direction bandwidt…

Read article →
ARTICLE · 401

H100 Pricing History, in depth: reading tiered price data, normalising quotes, cost per useful hour and per token, and pricing a contract

H100 rental prices as data: a tiered index from 2023 to 2025, why hyperscaler, neocloud and marketplace prices diverge, unit traps, a script that conv…

Read article →
ARTICLE · 402

H200 vs H100 Economics, in depth: cost per token, the break-even price ratio and when extra HBM pays

When an H200 is worth its premium over an H100: the break-even rule in cost per token, why freed KV-cache capacity can raise throughput faster than ba…

Read article →
ARTICLE · 403

Immersion Cooling, in depth: choosing a fluid, qualifying servers and optics, floor loading, fluid health monitoring and servicing submerged GPU servers

Running an immersion cooling programme for GPU servers: fluid families and PFAS exposure, material compatibility soak tests, optics and warranty, floo…

Read article →
ARTICLE · 404

Inference Cost per Query, in depth: pricing the whole request graph, the cost distribution and cost per successful answer

How to price one user query end to end: token, GPU-time and fixed-fee spans, a runnable cost model, worked retrieval and agent examples, why caching c…

Read article →
ARTICLE · 405

InfiniBand NDR (400G) + XDR (800G), in depth: lanes, radix, fabric size, host PCIe limits and collective-time math

What changes between InfiniBand NDR and XDR and why training jobs care: per-lane rates and port split, Quantum-2 versus Quantum-X800 radix, fat-tree s…

Read article →
ARTICLE · 406

Intel Gaudi 2 + Gaudi 3, in depth: the generations compared, the in-box Ethernet mesh, scale-out arithmetic and porting between them

Gaudi 2 against Gaudi 3 for the people who run software on them: published figures, the all-to-all RoCE mesh and why it favours tensor parallelism of …

Read article →
ARTICLE · 407

Kubernetes GPU Operator, in depth: operands and their order, ClusterPolicy, per-node control, sharing hooks and driver upgrades without lost jobs

How the NVIDIA GPU Operator turns a Kubernetes node into a GPU node: driver, toolkit, validator, device plugin, feature discovery and DCGM exporter in…

Read article →
ARTICLE · 408

KV Cache Reuse and Prefix Caching, in depth: hash chains versus radix trees, prompt layouts that earn hits, cache-aware routing across a fleet, and isolation

How KV cache reuse works beyond a single engine: hash-chain and radix-tree indexes, prompt layouts that keep prefixes stable, why round robin destroys…

Read article →
ARTICLE · 409

Lambda Labs, in depth: instances, the Cloud API lifecycle, region-locked filesystems, 1-Click Clusters and the billing traps

How to operate on the Lambda GPU cloud: on-demand instances, 1-Click Clusters and private cloud, a controller that finds capacity and launches within …

Read article →
ARTICLE · 410

Latitude.sh

Bare metal GPU cloud with reserved and on-demand capacity. Predictable pricing, physical servers, minimal managed services. Allocation shape, contract…

Read article →
ARTICLE · 411

Meta AI Research SuperCluster (RSC), in depth: DGX A100 nodes on a non-blocking InfiniBand Clos, the AIRStore data path, isolation and reliability at scale

Meta's AI Research SuperCluster as a design case study: 6,080 A100 GPUs at launch on a non-blocking two-level InfiniBand Clos, tiered flash and A…

Read article →
ARTICLE · 412

Meta MTIA, in depth: the processing-element grid and memory hierarchy, why recommendation models shaped it, the PyTorch-to-Triton path, and what MTIA 300 and 400 change

Meta's MTIA accelerator from a software engineer's view: generation names, MTIA 200's PE grid, SRAM and LPDDR5, roofline arithmetic for…

Read article →
ARTICLE · 413

Azure AI Supercomputer for OpenAI, in depth: five generations from 2020 to Fairwater, NVLink domains, fat trees, checkpoints, failure rates and multi-site training

The Microsoft systems built to train OpenAI's models, from the 2020 10,000-GPU cluster through Eagle to Fairwater's GB200 NVL72 halls and AI…

Read article →
ARTICLE · 414

MIG, in depth: what a CUDA process sees inside a slice, sizing weights and KV cache to an instance, MPS on MIG, training limits and measuring isolation

MIG from the application's side: GPU and compute instances, CUDA enumeration before and after R570, IPC and P2P rules, sizing an 8B model and its…

Read article →
ARTICLE · 415

MIG for Multi-Tenant GPU Sharing, in depth: tenant classes, layout planning, quotas, re-layout drains and chargeback

Running MIG as a shared service: mapping tenants to H100 profiles, GPU Operator single and mixed strategies, mig-parted layouts, ResourceQuota per pro…

Read article →
ARTICLE · 416

Modal Labs, in depth: serverless GPUs in Python, cold-start anatomy and memory snapshots, autoscaling knobs, batch fan-out, preemption-safe training and multi-node clusters

How to build on Modal: apps, images and GPU functions, GPU type strings and fallbacks, what a cold start costs and how min_containers, scaledown_windo…

Read article →
ARTICLE · 417

Model Lifetime Utility, in depth: amortising training GPU-hours over a serving life, demand decay, the replica floor and when to retire a model

A GPU-hour ledger for the life of a model version: fixed versus recurring hours, ramp, plateau and decay of demand, the replica floor, a runnable simu…

Read article →
ARTICLE · 418

Modular Datacenters for AI, in depth: interface contracts, modules as failure domains, factory and site testing, and wiring modules into the scheduler

How prefabricated power, cooling and IT modules work for AI datacenters: interface contracts with transients, modules as failure and maintenance domai…

Read article →
ARTICLE · 419

Small Modular Reactors (SMR) for AI DCs, in depth: the designs, connection models, sizing around refuelling, training load swings and realistic timelines

An engineering view of feeding an AI data center from small modular reactors: what counts as an SMR, the designs with real licensing progress, behind-…

Read article →
ARTICLE · 420

MoE Model Inference Cost, in depth: capacity, bandwidth and compute bills, experts touched per batch, and a dollars-per-million-tokens model

A first-principles cost model for serving mixture-of-experts LLMs: why capacity scales with total parameters, decode bandwidth with experts touched an…

Read article →
ARTICLE · 421

Multi-GPU Peer-to-Peer Access, in depth: the server as a graph, NCCL transport choice, one-shot all-reduce over peer pointers and placing groups on P2P cliques

Peer access across four to eight GPUs: modelling a server as a connectivity graph and finding cliques, how NCCL chooses P2P, SHM or network transport …

Read article →
ARTICLE · 422

Multi-Year GPU Capacity Commitments, in depth: a locked rate against falling prices, sizing under demand risk, contract clauses and the fleet lifecycle

How to evaluate a two-to-five-year GPU commitment: a present-value model against a falling market price for the same work, break-even decline, term co…

Read article →
ARTICLE · 423

All-Reduce, in depth: the ring and tree algorithms, NCCL protocols, busbw versus algbw, the APIs, a worked cost estimate and how training jobs fail

A first-principles guide to all-reduce for GPU training: the alpha-beta cost model, ring reduce-scatter plus all-gather and why it is bandwidth-optima…

Read article →
ARTICLE · 424

NCCL, in depth: communicators, topology search, transports, buffer registration, fault recovery and a debugging runbook

NCCL as a runtime you operate: communicator init and split, topology detection and transports, channels and the proxy thread, buffer registration and …

Read article →
ARTICLE · 425

NCCL Topology-Aware Communication, in depth: the path-type ladder, topology files for VMs, NIC selection and diagnosing what NCCL detected

How NCCL turns a node's topology into decisions: the NVL, PIX, PXB, PHB and SYS path types and the variables that gate on them, NVB and PXN, topo…

Read article →
ARTICLE · 426

Nebius AI Cloud, in depth: regions and InfiniBand fabrics, building a GPU cluster, P-Key isolation, Soperator's shared root and acceptance testing

A practitioner's runbook for Nebius AI Cloud: the region, fabric, GPU cluster and VM model, the current fabric and platform table, CLI cluster cr…

Read article →
ARTICLE · 427

NCCL + NIXL, in depth: lock-step collectives versus one-sided transfers, the NIXL agent lifecycle, and moving KV cache between instances

When to use NCCL and when to use NIXL: communicators versus agents, memory sections, backend plug-ins and metadata, the NIXL Python lifecycle, prepare…

Read article →
ARTICLE · 428

NVIDIA Nsight Systems, in depth: capture windows for training jobs, nsys stats reports, the SQLite export, scripted idle-gap analysis and recipes

Using nsys as an analysis pipeline for GPU training: NVTX and profiler start-stop capture windows, the stats reports, the SQLite export schema, a merg…

Read article →
ARTICLE · 429

NVIDIA Rubin, in depth: HBM4 and NVLink 6 through a software lens, the NVL72 domain, NVFP4 training and how to get ready

What NVIDIA Rubin changes for people who train and serve models: verified specifications, roofline and decode arithmetic, the 72-GPU NVLink 6 domain a…

Read article →
ARTICLE · 430

NVIDIA Spectrum-X, in depth: why ECMP fails AI traffic, per-packet adaptive routing, out-of-order placement, NIC-side congestion control and operating it

How NVIDIA Spectrum-X changes RoCE fabrics for AI training: an ECMP collision worked example, per-packet adaptive routing, SuperNIC reordering and dir…

Read article →
ARTICLE · 431

GB200 NVL72 Rack, in depth: the rack as a physical and failure unit, link arithmetic, power and cooling, acceptance testing, spare trays and checkpoint cadence

The GB200 NVL72 rack from an operator's point of view: 18 compute and 9 NVLink switch trays, why 1,296 links make 18 planes, facility power and l…

Read article →
ARTICLE · 432

GPU Occupancy Calculator, in depth: the exact arithmetic, a CI-ready calculator, the CUDA occupancy API and reading the result

How a GPU occupancy calculator works: kernel and device inputs from ptxas and the compute capability table, register and shared memory allocation gran…

Read article →
ARTICLE · 433

OCI GPU SuperCluster, in depth: bare-metal GPU shapes, the RoCEv2 cluster network, compute clusters, bandwidth arithmetic and bring-up

How Oracle's GPU SuperCluster works for training jobs: bare-metal H100, H200, B200, B300 and GB200 shapes, per-GPU RDMA bandwidth, the RoCEv2 clu…

Read article →
ARTICLE · 434

Optical Transceivers for AI Networking, in depth: the signal path, reach classes, linear and co-packaged optics, FEC budgets and fleet telemetry

How optical transceivers work in GPU training fabrics: inside a pluggable module, VR, SR, DR, FR and LR reach classes, DSP versus linear and co-packag…

Read article →
ARTICLE · 435

Optimizer State Offloading to CPU/NVMe

Optimizer state offloading as an engineering decision: the per-parameter byte cost of Adam, why optimizer state is the first tier to move off the GPU,…

Read article →
ARTICLE · 436

PCIe Gen4 + Gen5 for GPUs, in depth: what the generation changes, why links fall back, how to prove your link speed, and which workloads feel the doubling

PCIe Gen4 and Gen5 for GPU systems: per-direction bandwidth derived from the line rate, which GPUs use which generation, signal integrity and retimers…

Read article →
ARTICLE · 437

Power Backup for AI Datacenters, in depth: the ride-through timeline, why GPU loads stress generators, a checkpoint budget and wiring UPS events into the scheduler

How backup power protects AI training: PSU holdup, rack BBUs, UPS, transfer switches and generators in time order, why synchronized GPU loads stress g…

Read article →
ARTICLE · 438

Power Density in AI Datacenters, in depth: from GPU watts to rack kilowatts, the airflow and coolant physics, why AI racks are built dense, and density-aware scheduling

Why AI racks run from 40 kW to over 120 kW: how per-GPU power compounds, the airflow and coolant flow arithmetic, electrical and floor-loading limits,…

Read article →
ARTICLE · 439

PUE for AI Datacenters, in depth: a loss model for liquid-cooled GPU halls, facility water temperature, heat reuse and the levers that move the number

A design-side guide to PUE for AI datacenters: building an hour-by-hour loss model of a liquid-cooled GPU hall, electrical conversion losses, liquid c…

Read article →
ARTICLE · 440

Qualcomm Cloud AI 100, in depth: AI cores, SRAM versus LPDDR, compiling to a QPC and sizing LLM inference by bandwidth

The Qualcomm Cloud AI 100 as a software target: SKU specifications, the tensor, vector and scalar units of an AI core, the scratchpad memory model, th…

Read article →
ARTICLE · 441

Rail-Aligned Topology for AI Clusters, in depth: why collectives follow rails, how NCCL and PXN use them, and how placement keeps them

Rail-aligned (rail-optimized) GPU fabrics from the software side: the NIC-i-to-rail-i rule, hierarchical all-reduce on rails, NCCL_CROSS_NIC and PXN f…

Read article →
ARTICLE · 442

Ray on GPU and Distributed Training, in depth: FSDP inside Ray Train, NCCL setup, data shards, sharded checkpoints and elastic recovery

How Ray Train runs multi-node GPU training: controller and worker group, ScalingConfig and TorchConfig, FSDP with Distributed Checkpoint, Ray Data ing…

Read article →
ARTICLE · 443

Rear-Door Heat Exchangers, in depth: passive versus active doors, sizing the energy balance, dew point, telemetry and GPU throttling

Rear-door heat exchangers for GPU racks from first principles: passive and active doors, air and water energy balance, effectiveness, a worked 40 kW s…

Read article →
ARTICLE · 444

GPU Register Pressure

How the CUDA compiler allocates registers and what drives per-thread demand up: live ranges, loop unrolling, inlining and large local arrays. Why spil…

Read article →
ARTICLE · 445

Reserved GPU Capacity vs On-Demand, in depth: the run-level decision, gang acquisition, sizing a reservation window from a duration distribution and placing each workload

Reserved versus on-demand GPUs for a finite training run: why on-demand fails for gangs of nodes, a Monte Carlo duration model, choosing a reservation…

Read article →
ARTICLE · 446

Ring All-Reduce, in depth: the chunk schedule step by step, a working implementation, pipelining, multi-ring and hierarchical variants

How the ring all-reduce algorithm works: reduce-scatter and all-gather schedules, a traced four-GPU example, a verified simulator and a torch.distribu…

Read article →
ARTICLE · 447

SambaNova RDU

SambaNova RDU reconfigurable dataflow architecture for inference: how spatial mapping of computation graphs onto distributed compute and memory units …

Read article →
ARTICLE · 448

Second-Hand GPU Market

Enterprise churn feeds resale.

Read article →
ARTICLE · 449

SHARP

How SHARP moves reduction arithmetic into the switch ASIC: aggregation-tree construction and the finite switch state that backs it, which collectives …

Read article →
ARTICLE · 450

SIMT Execution: How Thousands of GPU Threads Compute in Parallel, with Worked Examples

How the SIMT model maps scalar-looking CUDA threads onto 32-wide warps: active masks, branch divergence and its cost in worked numbers, predication, i…

Read article →
ARTICLE · 451

DGX SuperPOD Topology, in depth: scalable units, four fabrics, the switch arithmetic, the UFM node and how Slurm and NCCL use the layout

The DGX SuperPOD as a system: the 32-node scalable unit and why pods have 127 nodes, the compute, storage and two management fabrics, non-blocking lea…

Read article →
ARTICLE · 452

TensorRT, in depth: the builder, tactic selection, strong typing and ModelOpt precision in TensorRT 11, optimization profiles, the runtime API and engine portability

How TensorRT compiles and runs inference: graph fusion, tactic timing and the timing cache, strongly typed precision with ModelOpt AutoCast and Q/DQ q…

Read article →
ARTICLE · 453

Foundation Model Project Cost Breakdown, in depth: the full ledger beyond the final run, compute from 6ND and MFU, a runnable cost model, a worked 30B example, sensitivity and tracking actuals

What a foundation-model project really costs: reading published figures for scope, a ten-line ledger covering research, failures, post-training, evalu…

Read article →
ARTICLE · 454

Training Cost Trajectory 2024-2027, in depth: what the 2.4x-a-year trend measures, a cost identity, calibration and a forecast you can run

Frontier training cost from 2024 to 2027: final-run versus programme cost, Epoch AI's 2.4x-per-year estimate and its uncertainty, a FLOPs times p…

Read article →
ARTICLE · 455

Training on Thousands of GPUs, in depth: composing TP, PP and DP, mapping the mesh to the network, and a full step budget for 70B on 1,024 GPUs

How to lay out an LLM training job across thousands of GPUs: the parallelism axes, building a device mesh, mapping groups to NVLink and the scale-out …

Read article →
ARTICLE · 456

Tree All-Reduce, in depth: NCCL's double binary tree built bit by bit, a simulator, the pipelined cost model and when rings win

Tree all-reduce from the inside: NCCL's rank bit arithmetic for the binary tree, the mirror and shift rules for the second tree, a Python port an…

Read article →
ARTICLE · 457

OpenAI Triton Language, in depth: blocks, constexpr and specialization, masks, tl.dot, and a fused matmul written line by line

Triton as a language: program instances and blocks, shape and dtype rules, constexpr and integer specialization, masked loads, tl.dot and pipelined lo…

Read article →
ARTICLE · 458

Ultra Ethernet Consortium, in depth: the UET transport layer by layer, packet spraying, trimming, NSCC and RCCC, LLR and CBFC, profiles, and what training traffic gains

How the Ultra Ethernet specification works and what it changes for GPU training collectives: SES, PDS, CMS and TSS sublayers, job-based addressing, co…

Read article →
ARTICLE · 459

Vast.ai, in depth: the GPU marketplace model, choosing offers, interruptible bidding, and training jobs that survive a kill

How Vast.ai works for ML: offers, hosts and the three billing meters, on-demand versus reserved versus interruptible bidding, the search query languag…

Read article →
ARTICLE · 460

Voltage Park, in depth: a foundation-owned H100 cloud, its bare-metal on-demand and reserved offers, the Lightning AI merger, and how to validate and run training on it

Voltage Park explained with dated facts: Navigation Fund ownership, about 24,000 H100s at launch, HGX nodes on Quantum-2 InfiniBand and VAST storage, …

Read article →
ARTICLE · 461

Warp Scheduling on GPU

How a GPU SM warp scheduler picks which warp issues each cycle: eligible versus merely resident warps, the scoreboard, and the stall-reason taxonomy (…

Read article →
ARTICLE · 462

Waste Heat Reuse from Datacenters, in depth: heat grade, heat pumps, ERF and why training load shapes heat supply

Reusing GPU datacenter heat: temperature grade and liquid cooling, the heat path to a district network, heat-pump COP in kelvin, the Energy Reuse Fact…

Read article →
ARTICLE · 463

Water Usage in AI Datacenters, in depth: withdrawal versus consumption, evaporation physics, cycles of concentration, source water and a per-job water model

How AI datacenters use water, built from first principles: withdrawal, consumption and discharge, about 1.5 litres evaporated per kWh of heat, blowdow…

Read article →
ARTICLE · 464

xAI Colossus, in depth: the 64-GPU rack and 512-GPU array, Spectrum-X Ethernet for collectives, power transients and failure math at 100,000 GPUs

xAI's Memphis training cluster read as engineering: the reported server, rack and array building blocks, rail-optimised Spectrum-X Ethernet and w…

Read article →
ARTICLE · 465

XLA, in depth: tracing to StableHLO, HLO passes, fusion, layout and buffer assignment, SPMD partitioning and avoiding recompilation

How the XLA compiler turns a JAX function into device code: jaxpr, StableHLO, HLO fusion, layout and buffer assignment, Shardy partitioning, GPU code …

Read article →
ARTICLE · 466

Zero-Copy Memory (Pinned Host), in depth: what a kernel load over PCIe costs, the break-even against copying, sparse gathers and host-visible results

How CUDA zero-copy memory works: mapped pinned host memory with cudaHostAllocMapped and cudaHostRegisterMapped, the cost of a kernel load over PCIe, c…

Read article →
ARTICLE · 467

DeepSpeed ZeRO Stages, in depth: choosing stage 0 to 3 by memory and traffic, gradient accumulation, the stage ladder and moving checkpoints between stages

How to choose a DeepSpeed ZeRO stage: per-stage sharding and traffic per optimizer step, why gradient accumulation makes stage 2 and 3 move more bytes…

Read article →