All 467 articles, sorted alphabetically
100+ MW AI Clusters, in depth: what the power budget buys, how often a 50,000-GPU job breaks, checkpoint intervals and goodput, power swings and fabric depth
A 100 MW AI site from the training job's side: per-GPU all-in power and the GPU count it buys, failure rates scaled from Llama 3 data, Young/Daly…
Read article →GPU Allocation Queues, in depth: quota, provider capacity and cluster admission as one pipeline, why large gangs wait longest, and how to shape requests
GPU allocation as a queueing problem: the quota, provider capacity and cluster admission queues, an Erlang C model of how gang size and load drive wai…
Read article →Amazon Inferentia2 + Inferentia3, in depth: the Inferentia line in 2026, why there is no Inferentia3, what still belongs on Inf2, leaving Inf1 and the LLM decision
The Inferentia product line as it stands in October 2026: Inf1 and Inf2 compared, the absence of an Inferentia3, which workloads still belong on Inf2,…
Read article →Amazon Trainium 2 + UltraServers, in depth: the 64-chip NeuronLink domain, where 1.28 TB/s goes, mapping parallelism onto 4x4x4, failure blast radius and Trn3
How a Trn2 UltraServer works as a system: the 4x4 torus plus cross-instance rings, reconciling the bandwidth figures, collective cost at each level, p…
Read article →AMD MI325X + MI350, in depth: the CDNA 3 to CDNA 4 step for kernels, FP8 encodings, MXFP4, memory plans and migration
What changes from AMD's MI325X to the MI350X and MI355X for software: 256 wider CUs and 160 KB LDS, machine balance, OCP FP8 versus FNUZ, MX form…
Read article →AMD ROCm Software Stack, in depth: the layers from amdgpu and KFD to HIP and libraries, code objects, containers and triage
A layer map of the AMD ROCm stack for operators: amdgpu and KFD device nodes, the ROCr HSA runtime, the HIP runtime and compiler, code objects and gfx…
Read article →Apple M-Series Neural Engine, in depth: what is known, ANE-friendly layers, checking op placement with MLComputePlan, fp16 limits and W8A8
How software reaches the Apple Neural Engine: the Core ML pipeline and graph partitioning, the (B, C, 1, S) transformer layout with 1x1 convolutions, …
Read article →ARM Neoverse for AI, in depth: core families, SVE2, BF16 and I8MM, runtime dispatch and bandwidth-bound CPU inference
How ML software uses Arm Neoverse cores in Graviton, Grace, Axion and Cobalt: which core ships where, what SVE2, BFMMLA and SMMLA compute, HWCAP featu…
Read article →AWS P5 and P4 GPU Instances, in depth: six instance types, per-GPU network math, EFA setup, NCCL and choosing between them
AWS P4 (A100) and P5 (H100, H200) instances for training and serving: specs checked against AWS documentation, the NVSwitch and EFA fabrics, per-GPU n…
Read article →AWS SageMaker HyperPod, in depth: instance groups, lifecycle scripts, health agents, deep health checks, auto-resume and what failures still cost
How SageMaker HyperPod clusters work for distributed training: the instance group model, a CreateCluster definition, lifecycle scripts, the health age…
Read article →AWS UltraCluster, in depth: getting co-located capacity, Capacity Blocks, topology-aware rank ordering, EFA collectives and failure handling at thousands of GPUs
EC2 UltraClusters explained for the engineer running the job: which instances live in them and their EFA bandwidth, Capacity Blocks and placement, the…
Read article →Azure ML Compute, in depth: GPU compute targets, cluster settings, quota, multi-node PyTorch, InfiniBand and low-priority capacity
How Azure Machine Learning compute runs GPU training: choosing between clusters, serverless and Kubernetes, the cluster settings that decide cost, how…
Read article →Azure ND-Series GPU VMs, in depth: the current sizes, one InfiniBand adapter per GPU, images and NCCL topology, sizing by memory and node health
What Azure's ND GPU VM sizes contain (A100 v4, H100 v5, H200 v5, MI300X v5, GB200 v6), how NVSwitch and per-GPU InfiniBand adapters shape paralle…
Read article →Bandwidth per GPU (Bisection), in depth: injection vs bisection vs NVLink, oversubscription, which collectives care, a model and how to measure it
What per-GPU bisection bandwidth means, how to compute it from a leaf-spine or rail-optimised fabric, why all-to-all is bisection-bound and ring all-r…
Read article →Shared Memory Bank Conflicts, in depth: the bank function, counting wavefronts, padding versus XOR swizzle, and proving the fix in Nsight Compute
How shared memory banks work and how to remove conflicts: the bank function and broadcast rule, the gcd stride rule, 64-bit and 128-bit phases, a Pyth…
Read article →Bare Metal vs Cloud GPU Cost Analysis, in depth: utilization break-even, commitment layering and a model you can run
How to decide between on-demand cloud, committed cloud, rented bare metal and owned GPUs: what bare metal changes technically, the commit-level rule f…
Read article →Batching Impact on Inference Cost, in depth: achieved versus configured batch, fleet consolidation, the price of a latency target, padding and cost attribution
Why the batch an LLM server actually runs is set by traffic per replica, not the configured maximum: Little's law for achieved batch, a worked fl…
Read article →Datacenter Busbars and Power Distribution, in depth: the in-rack DC busbar, power shelves, voltage drop and heat, 800 VDC, and what training jobs do to the rack
How AI racks deliver power to GPUs: the conversion chain from grid to die, why 48 to 54 V DC busbars replaced cords, a worked busbar sizing example wi…
Read article →Cerebras WSE-3, in depth: the wafer-scale processor, the PE mesh, weight streaming for training, multi-wafer inference and how software reaches it
How the Cerebras WSE-3 works and how software uses it: 900,000 PEs with local SRAM, why 21 PB/s changes the decode bottleneck, weight streaming with M…
Read article →Chilled Water Systems for AI DCs, in depth: chillers and lift, warm and cold loops, sizing flows and N+1, staging, ride-through storage and failure modes
How the chilled water plant behind a liquid-cooled AI hall works: chiller lift and COP, ASHRAE W-classes, two-temperature loops, flow and tonnage arit…
Read article →Coalesced Memory Access
How a warp's 32 lane addresses turn into global-memory transactions: sector and cache-line granularity, sector efficiency as requested bytes over…
Read article →Coherent Optics for Long-Distance, in depth: field modulation, the DSP chain, 400ZR and ZR+, OSNR link budgets and training across sites
How coherent optics carry data centre interconnects: dual-polarisation QAM, the receive DSP chain, the FEC cliff, the 400ZR family, an OSNR and disper…
Read article →Cold Plate Manufacturing, in depth: fin forming, lid joining, flatness, cleanliness, factory tests, traceability and measuring plate quality from GPU telemetry
How GPU cold plates are made and how manufacturing shows up in telemetry: materials, skived, machined, bonded and printed fins, brazing and friction s…
Read article →Colossal-AI Framework, in depth: the Booster and its plugins, Gemini chunk memory, HybridParallelPlugin with Shardformer, and sizing a 13B run
Colossal-AI as a practitioner uses it: the Booster API and the six plugins, a minimal training program, Gemini's chunk manager and placement knob…
Read article →Compute + Storage Co-Design, in depth: the four traffic classes of a GPU cluster, a worked 1,024-GPU storage budget, tier placement, job shaping and interference failures
How to size and shape storage for GPU training and serving: training reads, checkpoint bursts, restart reads and model loads; a worked 1,024-GPU budge…
Read article →CUDA Cooperative Groups: Thread Block Tiles, Grid Sync and Cooperative Launch
A practical guide to CUDA Cooperative Groups: thread block tiles and warp collectives without masks, cg::reduce, coalesced and labeled partitions for …
Read article →Copper vs Optical for Intra-DC, in depth: cable families, why reach halves per lane-rate step, a cable planner and what failing links look like to software
Copper versus optics inside an AI datacenter, hop by hop: DAC, ACC, AEC, AOC and transceivers, the loss-budget physics behind copper reach, power and …
Read article →CoreWeave
How the GPU neocloud model works, with CoreWeave as the worked example: non-blocking InfiniBand fabric versus oversubscribed general-purpose cloud net…
Read article →Cost per Token Trained, in depth: marginal versus fully loaded cost, the denominator problem, live tracking and amortisation
How to compute the cost of each training token: the 6N derivation with MFU and goodput, a 7B on 2T tokens worked example, overhead multipliers for exp…
Read article →Crusoe Energy, in depth: from flare gas to gigawatt campuses, what power-first siting means for latency and data, and running jobs on Crusoe Cloud
Crusoe from first principles: why it agreed to sell its flare-gas business, how power-first siting shapes latency, data gravity, capacity timing and c…
Read article →cuBLAS, in depth: column-major GEMMs from row-major frameworks, handles and workspace, cuBLASLt heuristics and epilogues, FP8 rules and reproducibility
cuBLAS as an API from first principles: the row-major trick, handles, streams and workspace, compute types and emulation, cuBLASLt descriptors, heuris…
Read article →CUDA Cores and Streaming Multiprocessors: Inside the GPU, from Warps to Register Files
What a CUDA core really is, what sits inside an H100 streaming multiprocessor, how blocks and warps map onto SMs, the register file as the largest on-…
Read article →CUDA Streams and Concurrency, in depth: default-stream semantics, stream-ordered APIs and libraries, a chunked copy-compute pipeline, host functions and async errors
CUDA streams at the runtime level: legacy versus per-thread default streams, which APIs and libraries are stream-ordered, writing stream-correct libra…
Read article →cuDNN, in depth: the graph API, engines and execution plans, a fused convolution with bias and ReLU, PyTorch switches and failure modes
How cuDNN turns deep-learning operations into GPU kernels: legacy versus graph API, the cuDNN 9 sub-libraries, heuristics, engine configs and plans, a…
Read article →Data Curation Cost for Foundation Models, in depth: cost per retained token, filter ordering, the ablation bill and a break-even against training compute
What curating pretraining data really costs: cost per retained token through the funnel, a runnable cost model, a 2-billion-page worked example where …
Read article →Data Parallel Training, in depth: the DDP reducer and its buckets, a complete torchrun script, accumulation, scaling measurements and the traps that hang jobs
Data parallel training with PyTorch DDP as an engine: the reducer, gradient buckets and overlap, a complete torchrun script, gradient accumulation wit…
Read article →DGX H100 System Architecture, in depth: the three fabrics, GPU-to-NIC and NUMA affinity, the local NVMe cache, a worked all-reduce estimate, validation and failure modes
How a DGX H100 looks to software: eight H100 GPUs on four NVSwitch chips, one 400 Gb/s ConnectX-7 rail per GPU, two NUMA domains, a RAID 0 NVMe cache …
Read article →Distillation for Cost Savings, in depth: GPU-seconds per request, the one-time ledger, fallback and upkeep, and the break-even volume
When does distilling a large model into a small one actually save money? Measuring GPU-seconds per request at the SLO, pricing data, training, evaluat…
Read article →Distributed Checkpointing, in depth: global-offset metadata, the save and load planners, resharding by chunk intersection, deduplication and format conversion
How PyTorch Distributed Checkpoint makes checkpoints independent of the parallel layout that wrote them: the .metadata index and .distcp files, the pl…
Read article →Dragonfly Network Topology, in depth: groups and global links, the radix arithmetic, minimal versus Valiant and UGAL routing, and placing training jobs
The dragonfly interconnect from first principles: p, a and h, router radix and maximum size, minimal local-global-local routes, the group-shift patter…
Read article →CUDA Dynamic Parallelism, in depth: CDP2 tail launch and fire-and-forget streams, memory visibility, launch-pool limits and migrating from CDP1
How CUDA Dynamic Parallelism works since CUDA 12 (CDP2): parent and child grids, device streams including tail launch and fire-and-forget, memory visi…
Read article →AWS EFA, in depth: interface types, the SRD transport, the libfabric stack, network rules, counters for triage and EFA on EKS
AWS Elastic Fabric Adapter from the software side: ENA, EFA and EFA-only interfaces, why SRD is reliable but unordered and multipath, the NCCL to libf…
Read article →Fat-Tree Network Topology, in depth: radix math, bisection, oversubscription, routing collisions and job placement for GPU clusters
How fat-tree (folded Clos) networks for GPU clusters are sized from switch radix, what bisection bandwidth and oversubscription mean for all-reduce an…
Read article →FLOPS Budget for LLM Training, in depth: counting 6N plus attention, MFU versus HFU, goodput and a worked GPU-hour ledger
Plan the compute for an LLM training run: derive FLOPs per token from 6N plus the attention term, use dense rather than sparsity peaks, convert with m…
Read article →FluidStack, in depth: what a dedicated neocloud GPU cluster is, how its fabric shapes training, acceptance burn-in, topology-aware Slurm, checkpoint cadence and the contract questions that matter
Fluidstack and the dedicated-cluster neocloud model, explained for training teams: node and fabric layout, a reproducible acceptance test with nvidia-…
Read article →GB200 Grace Blackwell, in depth: the compute tray as a Linux host, Arm porting, NUMA affinity and tray acceptance testing
GB200 from the host's point of view: what a compute tray exposes to Linux, porting a training stack to aarch64, 4 KB versus 64 KB pages, NUMA and…
Read article →GKE with TPU, in depth: slice node pools, labels and chip requests, Indexed Jobs, Multislice with JobSet, capacity options and failure modes
How TPUs work on Google Kubernetes Engine: single-host and multi-host slice node pools and their atomic behaviour, the accelerator and topology labels…
Read article →Google Cloud TPU Pods, in depth: pods, slices and cubes, torus wraparound, a collective cost model and mapping JAX meshes onto ICI and DCN
Cloud TPU pods as a network: pod, slice, cube, host and DCN vocabulary, v5p, v6e and Ironwood pod figures, when torus wraparound exists, twisted topol…
Read article →Google TPU v4 + v5 (v5e + v5p), in depth: megacore, the optically switched torus, SparseCores, and moving v4 jobs to v5p and v5e
TPU v4 as the reference design for v5e and v5p: megacore and TensorCore-counted names, 4x4x4 cubes joined by optical circuit switches, twisted tori, S…
Read article →Ada Lovelace / Blackwell architecture
Deep-dive on NVIDIA Hopper→Blackwell arc: SM redesign, FP8/FP4, memory bandwidth, NVLink generations, GB200 Grace-Blackwell.
Read article →Admission Control for LLM Serving
How admission control protects p99 latency in LLM serving: why unbounded queues fail, sizing a queue-depth cap from the SLO with Little's Law, lo…
Read article →AI Gateway Overview
How an AI gateway works as a policy plane in front of LLM backends: virtual API keys and identity, per-key and per-team budget enforcement, rate limit…
Read article →Alignment Data, in depth: SFT and preference formats, chat templates, loss masks, packing, decontamination and GPU cost
Engineering alignment data for training: the three record shapes, chat templates and loss masking, packing with cu_seqlens, cached reference log-probs…
Read article →AMD Instinct GPUs
AMD Instinct and ROCm as the practical alternative to CUDA: how HIP maps onto the CUDA programming model, what hipify translates and what it silently …
Read article →Anthropic Workbench + Claude, in depth: what replaced the Console Workbench, and how to rebuild its prompt-evaluation loop in code
Workbench (legacy) in the Claude Console is retired and replaced by a stateless playground. What the playground does, what was lost, and how to rebuil…
Read article →Apigee AI Gateway, in depth: token quotas, prompt spike limits, semantic caching and Model Armor in front of GPU-backed models
How to turn an Apigee proxy into an AI gateway that protects GPU serving capacity: LLMTokenQuota enforce and count pairs, PromptTokenLimit spike contr…
Read article →Apple Silicon, in depth: unified memory, the GPU and its Neural Accelerators, the Neural Engine, and how MLX, PyTorch and llama.cpp use them for local inference and fine-tuning
How Apple Silicon works as a machine-learning computer: unified memory and why it changes what fits, the CPU, GPU, M5 Neural Accelerators and Neural E…
Read article →GPU Architecture Explained: Streaming Multiprocessors and Warps
How GPU architecture works: streaming multiprocessors (SMs), warps of 32 threads and SIMT execution, the memory hierarchy, occupancy and tensor cores.
Read article →ASR (Speech-to-Text) Serving, in depth: request pipelines, long-form chunking, batching, RTFx capacity planning and failure guards
How to serve speech-to-text on GPUs: batch versus streaming workloads, the decode-VAD-chunk-batch-stitch pipeline, where Whisper spends GPU time, a VA…
Read article →GPU asynchronous copy and software pipelining
Deep-dive on the asynchronous global-to-shared copy pipeline that fast GPU kernels are built on: cp.async streams tiles directly into a multi-stage sh…
Read article →Audio Encoders, in depth: log-mel and learned front ends, Whisper and wav2vec frame math, GPU cost, batching and streaming
How audio encoders such as Whisper, wav2vec 2.0 and HuBERT turn waveforms into frames, how to estimate their GPU cost from frame counts, and how fixed…
Read article →Axolotl, in depth: config-driven LLM fine-tuning, from YAML to packed batches, QLoRA memory, multi-GPU sharding and the mistakes that waste a run
How Axolotl turns one YAML file into a fine-tuning run: what each config block does at run time, chat_template datasets and label masking, the prepare…
Read article →Azure AI Foundry, in depth: resources, projects, deployment types and planning GPU capacity
How Azure AI Foundry (now Microsoft Foundry) works: the resource and project model, RBAC, how Standard, Provisioned, Batch and managed compute deploym…
Read article →NVIDIA B100, in depth: Blackwell at Hopper's 700 W, what a power-capped GPU means for training and inference, and how to plan around it
The B100 was announced as the 700 W Blackwell GPU for existing HGX H100 power and cooling. What its launch figures mean, why a power cap hurts compute…
Read article →NVIDIA B200, in depth: dual-die Blackwell, FP4 and microscaling, tensor memory, and what changes for training and inference
A practitioner's guide to the NVIDIA B200: the dual-die package, HBM3e capacity and bandwidth, fifth-generation tensor cores with tensor memory, …
Read article →GPU Batching Strategies Overview
A decision guide to GPU batching for inference: static, dynamic, continuous and chunked prefill compared on the workload signals that actually select …
Read article →Arena-Hard, in depth: BenchBuilder prompts, the pairwise judging protocol, scoring code, serving a candidate on your GPUs and gating releases
How Arena-Hard-Auto works and how to run it: v0.1 and v2.0 prompt sets, judges and baselines, the five-verdict two-game protocol, weighted bootstrap s…
Read article →Custom LLM Benchmarks, in depth: item sets from your own traffic, programmatic scorers, paired statistics and the GPU noise floor
Build a custom LLM benchmark that decides between models, quantisations and serving configs: items from production traffic, stratification, programmat…
Read article →LLM Eval Frameworks, in depth: how lm-evaluation-harness turns benchmarks into GPU work, from request types and batching to reproducible scores
How LLM eval frameworks work on the GPU: the request-centric pipeline, loglikelihood versus generate requests, boundary tokenisation, acc versus acc_n…
Read article →HELM Benchmark for Serving, in depth: wiring the harness to your inference server, reading its outputs and gating quantization and engine changes with paired statistics
How to use Stanford CRFM HELM against your own vLLM-style endpoint: scenarios, adapters and run entries, model_deployments.yaml with VLLMClient, the p…
Read article →LiveBench, in depth: contamination-limited releases, objective scorers, running it against a model on your own GPUs and comparing serving configurations
How LiveBench works and how to run it on your own GPUs: its categories and tasks, dated releases and private questions, the answer-judge-show pipeline…
Read article →LLMPerf Benchmark, in depth: how the load test measures, the inter-token metric that misleads, recomputing TPOT and goodput, and a worked concurrency sweep
How Ray's archived LLMPerf load generator measures TTFT, inter-token latency and throughput, read from its source; why its inter-token figure inc…
Read article →LMSYS Chatbot Arena, in depth: anonymous battles, Bradley-Terry scores, bootstrap ranks, style control and running a private arena
How Chatbot Arena (LMArena) turns anonymous pairwise votes into scores: the battle protocol, Bradley-Terry instead of online Elo, bootstrap intervals …
Read article →MLPerf Benchmark, in depth: how Training and Inference results are produced, LoadGen scenarios, LLM latency limits, divisions and how to read a result
How MLPerf measures GPUs and other accelerators: Training time-to-quality with olympic scoring and reference convergence points, Inference LoadGen sce…
Read article →SWE-bench for Serving Model Choice, in depth: resolve rate, cost and latency on your own endpoint
Use SWE-bench to choose a serving configuration, not a leaderboard winner: the serving knobs that move resolve rate, a pinned-scaffold rig, harness co…
Read article →vLLM Benchmark Suite, in depth: bench latency, throughput and serve, arrival processes, concurrency caps, goodput and a reproducible sweep
How to use vLLM's own benchmarks: offline latency and throughput versus online bench serve, TTFT, TPOT, ITL and goodput, request rate, burstiness…
Read article →BentoML for LLM Serving, in depth: services, vLLM integration, composition, adaptive batching, concurrency autoscaling and cold starts
A practical guide to serving LLMs with BentoML: what the framework does above the inference engine, a verified vLLM service using __command__, service…
Read article →Bricktree / Traefik AI, in depth: Traefik Hub's AI Gateway middlewares as a capacity control for GPU inference
Traefik Hub AI Gateway in front of self-hosted GPU inference: enabling it, middleware order, Chat Completion, token rate limits and quotas, semantic c…
Read article →GPU Capacity Planning, in depth: demand ledgers, stranded GPUs, spare pools and the buy trigger for a shared fleet
Capacity planning for a running, shared GPU fleet: a GPU-hour demand ledger with per-class utilisation targets, a placement simulation showing strande…
Read article →GPU Carbon Footprint, in depth: measuring energy per job, operational and embodied emissions, and the software levers that cut them
How to measure and reduce the carbon footprint of GPU training and inference: what is counted, reading the NVML energy counter, node overhead and PUE,…
Read article →GPU Checkpointing Deep Dive, in depth: bytes per parameter, the save path, async sharded saves and commits that survive crashes
How large-model training checkpoints work end to end: what state costs per parameter, the GPU-to-storage save path and which hop blocks training, a wo…
Read article →Chunked Prefill in Serving
How chunked prefill splits a long prompt into decode-sized slices so it stops blocking other requests' decode: the head-of-line stall it removes,…
Read article →CLIP Training on GPU, in depth: the contrastive loss, why batch size is everything, distributed local loss, memory arithmetic, precision and the data pipeline
How contrastive image-text pre-training actually runs on GPUs: the symmetric InfoNCE loss, why the negative set makes batch size a quality knob, the q…
Read article →GPU Cluster Bandwidth, in depth: the HBM-to-NIC ladder, what each parallelism strategy sends, a per-step communication budget, and how to measure it
The bandwidth ladder of a GPU training cluster with checked H100 and B200 figures, algbw versus busbw, bytes per step for DDP, FSDP and tensor paralle…
Read article →GPU Collective Operations, in depth: all-reduce, all-gather, reduce-scatter and all-to-all mapped to DDP, FSDP, tensor and expert parallelism
GPU collectives from first principles: what each operation does, which parallelism strategy issues it, an alpha-beta cost model, algorithm versus bus …
Read article →Collective Communication Overlap
How data-parallel training hides gradient all-reduce behind backward compute: gradient bucketing and why buckets fire during backward instead of after…
Read article →GPU-Enabled Containers
How to run GPU workloads in Docker containers via nvidia-container-toolkit.
Read article →Continuous batching
Deep-dive on continuous (in-flight) batching for LLM inference: the scheduling technique that forms a fresh batch every decode iteration, letting fini…
Read article →GPU Data Center Cooling: Air, Direct-to-Chip Liquid, Immersion
How GPU server cooling is chosen by kW per rack: air, rear-door heat exchangers, direct-to-chip liquid cooling, immersion, and the CDU coolant loop be…
Read article →Core ML for LLM, in depth: stateful KV cache, flexible shapes, int4 weights and the on-device generation loop
How to run a decoder language model through Core ML on Apple devices: why LLMs fit Core ML awkwardly, wrapping the KV cache as model state, converting…
Read article →GPU Cost Optimization
An ordered lever list for cutting GPU spend: utilization first, then right-sizing, batching, cache hit rate, quantization, request routing, context tr…
Read article →GPU CUDA Programming
The CUDA programming model: how kernels launch threads organized into blocks and grids, and the __global__/__device__ function distinction.
Read article →CUDA streams and graphs -- overlap and launch-overhead elimination
Deep-dive on CUDA streams and graphs: streams as ordered work queues enabling concurrency and copy-compute overlap, events for cross-stream synchroniz…
Read article →CUTLASS, in depth: the GEMM hierarchy, CuTe layouts, Hopper kernel schedules and the Python DSL
How NVIDIA CUTLASS builds matrix multiplications: tiling from first principles, the device, kernel, collective and atom layers, CuTe layouts, a comple…
Read article →GPU Data Pipeline, in depth: measuring data stalls, DataLoader workers, pinned memory, GPU decode and sharded storage
How to keep GPUs fed during training: the input pipeline as a five-stage throughput budget, measuring data wait against GPU time, PyTorch DataLoader w…
Read article →GPU Datacenter Deployment
Deploying GPU capacity as a build project: turning a megawatt envelope into a GPU count, floor loading and freight paths, the cable plant, commissioni…
Read article →GPU Datacenter Power Requirements
How power is delivered to a GPU hall: MW-per-rack math, why synchronized training looks like a square wave to the grid, step loads and breaker coordin…
Read article →GPU DataLoader Bottleneck, in depth: inside the PyTorch DataLoader, proving the loader is the limit, and fixing it
How the PyTorch DataLoader really works (workers, index queues, shared memory, pin thread, reorder buffer), three measurements that prove an input bot…
Read article →NVIDIA DCGM, in depth: the host engine, profiling metrics, health watches, diagnostics as a node gate and dcgm-exporter
NVIDIA Data Center GPU Manager as a fleet system: the host engine and its watch cache, groups and fields, what the profiling metrics say about a train…
Read article →Decode Compute Math, in depth: a per-operator FLOP and byte ledger for one LLM decode step
The arithmetic of one LLM decode step, operator by operator: FLOPs and HBM bytes for QKV, attention, MLP and LM head, why attention intensity equals t…
Read article →DeepInfra, in depth: the four ways to run a model, the per-model concurrency limit as a throughput ceiling, custom deployments, the Batch API and break-even math
DeepInfra for engineers: OpenAI-compatible and native APIs, the 200-concurrent-requests-per-model limit and Little's law, a client limiter with b…
Read article →DeepSpeed, in depth: the engine, the launcher, custom ops, pipeline and MoE engines, and when to pick it over FSDP
How the DeepSpeed library is put together and how to run it: what deepspeed.initialize wraps, the batch-size triangle, backward and step at accumulati…
Read article →GPU Depreciation Timeline, in depth: book life versus economic life, what useful-life disclosures mean, the training-to-inference cascade and when to replace a fleet
GPU depreciation as two clocks: straight-line book value against front-loaded economic value, hyperscaler useful-life changes, the five drivers of val…
Read article →Diffusion Model Serving, in depth: the cost model, shape buckets, step-level batching, prompt caching, compilation, LoRA handling and a split decode stage
How to serve text-to-image diffusion models on GPUs: a cost model built from steps, resolution and guidance, a staged request path, resolution buckets…
Read article →Direct Liquid Cooling, in depth: the heat path from die to coolant, inside a cold plate, flow and pressure drop, the residual air load and what a training job sees
How direct-to-chip liquid cooling works for GPU servers: the thermal resistance stack from junction to coolant, cold plate construction, flow per plat…
Read article →Disaggregated LLM Serving, in depth: the request path through a real proxy, vLLM and SGLang launch, routing, failure handling, and splits beyond prefill and decode
An operator's guide to disaggregated LLM serving: the kv_transfer_params handshake through a proxy, launching vLLM with NixlConnector and SGLang …
Read article →DPO Training on GPU, in depth: four forward passes, the reference-model bill, logits memory and a run that fits
The systems side of Direct Preference Optimization: what one DPO step computes on the GPU, where memory goes for an 8B model, the four ways to host th…
Read article →Dropless MoE, in depth: sort and permute dispatch, grouped and block-sparse GEMM, variable all-to-all, worst-case memory and stragglers
How dropless mixture-of-experts works on GPUs: why capacity factors drop and pad, the sort, permute and unpermute data flow, a PyTorch reference, grou…
Read article →PyTorch Dynamo, in depth: PEP 523 frame hooks, symbolic bytecode interpretation, sources and guards, dynamic shape symbols, graph breaks as continuations and writing your own backend
How TorchDynamo, the graph-capture front end of torch.compile, really works: the PEP 523 frame evaluation hook, InstructionTranslator and VariableTrac…
Read article →Edge Inference for LLM, in depth: bandwidth-bound decode, memory and KV budgets, runtimes, thermals and a hybrid local-cloud router
How to run language models on edge GPUs and SoCs: why decode is bandwidth-bound, sizing weights and KV cache, ceilings for Jetson Orin and Apple M4 pa…
Read article →Edge NPUs, in depth: how on-device neural accelerators execute models, and how to quantize, compile and deploy for them
How edge neural processing units work and how software uses them: MAC arrays, scratchpad SRAM and DMA tiling, why integer quantization is the entry fe…
Read article →Embedding Model Serving, in depth: token-budget batching, pooling and prefixes, TEI configuration, capacity math and versioned embedding contracts
How to serve embedding models on GPUs: what the encoder forward pass costs, CLS, mean and last-token pooling, instruction prefixes, token-budget batch…
Read article →GPU ECC and Memory Reliability, in depth: SECDED from first principles, containment, row remapping budgets and a drain-reset-replace policy
How GPU error-correcting codes work and how to operate them: a runnable SECDED encoder and decoder, where ECC and parity sit on the die and in HBM, vo…
Read article →GPU Experiment Tracking, in depth: stack fingerprints, honest step timing, MFU and goodput, cost ledgers and fair run comparisons
The efficiency and cost half of GPU experiment tracking: per-node stack fingerprints as comparison keys, step timing on an asynchronous device, MFU wi…
Read article →Expert Parallelism
Expert parallelism on the training side: how EP, tensor and data parallelism compose into one device mesh, which parameters shard by expert and which …
Read article →DoRA Fine-Tuning, in depth: magnitude and direction, a from-scratch layer, PEFT internals and GPU cost
DoRA from first principles: decomposing weights into a trainable magnitude and a LoRA-updated direction, a from-scratch PyTorch layer with merging, ho…
Read article →DPO Fine-Tuning, in depth: what beta controls, sweeping beta with learning rate on a GPU budget, measuring drift and choosing a checkpoint with real confidence intervals
Running a DPO fine-tune as an experiment: beta as the KL price with a worked toy, the gradient weight that couples beta and learning rate, a LoRA swee…
Read article →Full Fine-Tuning in Depth: Memory Budget, Data Preparation, Training Recipe, and When It Beats LoRA, with Pseudocode
How to fully fine-tune a pretrained language model on GPUs: the 16 bytes per parameter model-state budget and which ZeRO or FSDP stage fits 13B and 70…
Read article →GRPO Fine-Tuning, in depth: a run plan from task fit and difficulty filtering to rewards, LoRA or vLLM server layouts, sizing and checkpoint choice
A practical GRPO fine-tuning runbook: when GRPO fits, filtering prompts by pass rate so groups carry signal, testable reward functions, TRL configs fo…
Read article →LoRA Fine-Tuning, in depth: where GPU memory and time go when you train adapters on an 8B model
What a LoRA training step does on the GPU: why frozen weights still cost an input-gradient GEMM, 4N against 6N FLOPs per token, a memory budget for an…
Read article →Fine-Tuning Ops on GPU, in depth: a runbook for memory planning, launching, checkpointing, monitoring and shipping adapters
An operations runbook for fine-tuning LLMs on GPUs: reproducible run specs, a worked LoRA and QLoRA memory plan for an 8B model, the fp32 logits trap,…
Read article →ORPO Fine-Tuning, in depth: building preference pairs, templates and masking, memory planning, a lambda sweep and promotion gates
A practitioner's recipe for ORPO fine-tuning: when one-stage alignment fits, building preference pairs from your own model, chat templates and tr…
Read article →PPO Fine-Tuning for LLMs, in depth: a run plan from method choice and reward-model audits to memory sizing, reward-hacking guards and checkpoint selection
A practical PPO fine-tuning plan for LLMs: when PPO beats DPO or GRPO, the 2026 tooling change after TRL removed PPOTrainer, auditing the reward model…
Read article →SFT on GPUs, in depth: length distributions, padding versus packing, padding-free attention, logit memory and token-weighted loss
The GPU side of supervised fine-tuning: why SFT batches are mostly padding, length grouping versus BFD packing, varlen attention with cu_seqlens and r…
Read article →GPU Fine-Tuning Costs
Estimate fine-tuning spend before you launch: GPU-hour arithmetic, the bytes-per-parameter memory floor, LoRA and QLoRA as cost decisions, MFU, spot e…
Read article →Fireworks AI, in depth: deployments, precision and KV memory, speculative decoding with Predicted Outputs, structured output and a measurement harness
How to run open-weight models on Fireworks AI well: model strings, firectl deployments with shapes, accelerators, FP8 and replica limits, the GPU memo…
Read article →FlashAttention architecture
The GPU-side view of FlashAttention: the memory-traffic arithmetic that puts attention on the wrong side of the roofline, why a generic fusion compile…
Read article →FlashInfer, in depth: paged-KV attention kernels, the plan and run split, load-balanced scheduling and how serving engines use it
How FlashInfer accelerates LLM serving on NVIDIA GPUs: why serving attention differs from training attention, the block-sparse page-table format with …
Read article →FLUX on the GPU, in depth: tokens per image, the double-stream transformer, FLOP and memory budgets, offload, quantization and serving
How FLUX.1 image generation uses a GPU: the pipeline stages, how resolution becomes latent tokens, double-stream and single-stream transformer blocks,…
Read article →PyTorch FSDP, in depth: sharding parameters, gradients and optimizer state with fully_shard
PyTorch Fully Sharded Data Parallel explained from first principles: the per-block all-gather and reduce-scatter lifecycle, memory arithmetic for a 7B…
Read article →NVIDIA GB200, in depth: the Grace Blackwell superchip, the NVL72 NVLink domain and how software uses them
How the NVIDIA GB200 is built and how training and inference software uses it: the Grace CPU and two Blackwell GPUs on NVLink-C2C, HBM3E versus LPDDR5…
Read article →NVIDIA GH200, in depth: the Grace Hopper superchip, NVLink-C2C coherence and how software uses a CPU memory tier
What is on the GH200 module, how ATS and NVLink-C2C give the GPU coherent access to CPU memory, first touch and migration, a malloc-based CUDA example…
Read article →GPU Sharing Strategies Overview
A decision matrix for GPU sharing: MPS, time-slicing, MIG and vGPU scored on isolation strength, achievable utilization, blast radius, failure contain…
Read article →GPUDirect
A map of the GPUDirect family: which path a transfer actually takes -- P2P between GPUs, RDMA to the NIC, Storage from NVMe, or Async -- what each one…
Read article →GPUDirect Storage architecture
Deep-dive on GPUDirect Storage: the cuFile API and kernel driver, peer-to-peer PCIe DMA from NVMe and NVMe-oF into GPU HBM, the CPU bounce buffer it e…
Read article →GRPO, in depth: the group-relative objective from scratch, where the GPU time and memory go, and the Dr. GRPO and DAPO fixes
Group Relative Policy Optimization for LLMs, implemented and budgeted: the objective with group-normalised advantages, a PyTorch loss with three aggre…
Read article →NVIDIA H100 Architecture: Hopper GPU Explained
NVIDIA H100 Hopper architecture explained: the GH100 die and SMs, 80 GB HBM3, L2, NVLink 4, SXM5 vs PCIe, the FP8 Transformer Engine, MIG, and the A10…
Read article →NVIDIA H100 Architecture: Inside the Hopper SM, Tensor Cores
NVIDIA H100 (Hopper) GPU architecture at the SM level: four warp schedulers, register file limits, L1 and shared memory, tensor cores, TMA async copy,…
Read article →NVIDIA H200, in depth: 141 GB of HBM3e, why memory sets LLM decode speed, KV-cache sizing, training gains and deployment pitfalls
What the NVIDIA H200 changes and what it keeps from the H100: 141 GB HBM3e at 4.8 TB/s with the same Hopper compute, roofline reasoning for prefill an…
Read article →GPU Hardware Faults, in depth: how ECC errors, Xids, NVLink failures and silent corruption reach a training job, and how the job survives them
How GPU hardware faults show up to software: the failure arithmetic of large jobs, a taxonomy from correctable ECC to fallen-off-the-bus with NVIDIA&#…
Read article →GPU HBM architecture
How GPU HBM actually works: 3D stacking and TSVs, channels and banks, row buffers and the activate/precharge cycle, refresh, ECC, and why achieved ban…
Read article →HCCL, in depth: how Huawei Ascend collectives bootstrap, choose algorithms, use buffers and fail, with a PyTorch benchmark
Huawei HCCL explained for engineers training on Ascend NPUs: host NIC versus device RoCE links, rank tables and root info, the hierarchical algorithms…
Read article →Hugging Face Ecosystem, in depth: one checkpoint from Hub to GPU and back, through safetensors, device maps, attention kernels, adapters and serving
Follow a model through the Hugging Face stack on GPUs: pinned downloads, the safetensors format, meta-device loading and device_map placement, 4-bit l…
Read article →GPU Hyperparameter Tuning, in depth: systems knobs vs optimisation knobs, batch and learning-rate coupling, precision traps and budget-fair search
Tune training hyperparameters the GPU-aware way: separate micro-batch, accumulation and checkpointing from the learning rate, benchmark throughput and…
Read article →Image Classification Serving, in depth: GPU decode, TensorRT engines, Triton ensembles, dynamic batching and fleet sizing
How to serve image classifiers on GPUs: where the time goes, nvJPEG and DALI preprocessing, ONNX to TensorRT FP16 and INT8 engines, a Triton ensemble …
Read article →Immersion Cooling for GPU, in depth: single- and two-phase physics, sizing fluid flow, what changes inside the server, GPU telemetry and scheduler integration
Immersion cooling for GPU clusters from first principles: single-phase versus two-phase, a worked flow-sizing example, server changes for fans, heatsi…
Read article →GPU Incremental Training, in depth: continuing from checkpoints, LR re-warming, replay, sharded state and forgetting
How to add data to a trained model without retraining from scratch on GPUs: resume versus continual pre-training versus incremental fine-tuning, what …
Read article →PyTorch Inductor, in depth: lowering, the loop-level IR, fusion scheduling, Triton code generation, autotuning and AOT packaging
How PyTorch Inductor turns an ATen graph into GPU kernels: decompositions, the Pointwise and Reduction IR, vertical and horizontal fusion, generated T…
Read article →GPU Inference Latency, in depth: TTFT, TPOT and end-to-end time from first principles, with a runnable latency model
Where the milliseconds of an LLM request go on a GPU: TTFT, TPOT and end-to-end latency defined, prefill as a compute floor, decode as a memory-bandwi…
Read article →LLM Inference Optimization Overview
A triage guide to LLM inference optimization on GPUs: the levers in the order you should try them, from measuring whether you are prefill- or decode-b…
Read article →InfiniBand + NVLink architecture
How the scale-out fabric works for GPU training: RDMA and kernel bypass, GPUDirect RDMA into HBM, InfiniBand HCAs, switches and the subnet manager, fa…
Read article →GPU Infrastructure Planning, in depth: from workload demand to GPU count, fabric, storage bandwidth and a failure budget
How to plan a GPU cluster from the workload down: training compute from 6ND and MFU, inference capacity from measured throughput and headroom, memory …
Read article →GPU Infrastructure Cost, in depth: building the fully loaded cost of a useful GPU-hour from capital, power, facility and people
How to compute what a GPU-hour really costs on owned or colocated infrastructure: capital recovery with cost of capital, fabric and storage, power tim…
Read article →Intel Gaudi + GPU Max, in depth: two different architectures, how PyTorch drives each, porting from CUDA and running existing fleets
A practical guide to Intel's two data-center AI accelerator lines: Gaudi's graph-compiled matrix engines, tensor processor cores and integra…
Read article →Iterative DPO, in depth: on-policy rounds, pair building, the moving reference and scheduling GPUs between generation and training
Iterative DPO as a GPU workload: why one offline round goes stale, the generate-score-pair-train round loop, a worked budget where generation costs as…
Read article →ITL, in depth: inter-token latency as a distribution, the stalls behind its tail, the p99 cliff and ITL SLOs
Inter-token latency for LLM serving: ITL versus TPOT, the decode-step floor, prefill interference, preemption, speculative bursts and proxy buffering,…
Read article →GPU kernel fusion -- fewer kernels, less memory traffic
Deep-dive on GPU kernel fusion: the memory-bound HBM-traffic problem, fusing operations to keep intermediates on-chip (registers/SRAM), elementwise/ep…
Read article →Kong AI Gateway, in depth: ai-proxy plugins, balancing across vLLM GPU pools, token budgets, guards, caching and failure modes
Running Kong Gateway as an AI gateway in front of self-hosted vLLM GPU pools and hosted LLM APIs: the request path through the ai-* plugin chain, ai-p…
Read article →KServe for LLMs, in depth: InferenceService vs LLMInferenceService, model loading, KEDA autoscaling on engine metrics, multi-node serving and safe rollouts
How KServe serves LLMs on Kubernetes: the Hugging Face runtime with vLLM, OpenAI-style routes, LLMInferenceService for multi-node and disaggregated se…
Read article →KTO Training, in depth: Kahneman-Tversky Optimization from thumbs-up data, the in-batch KL reference point, GPU cost, imbalance weights and reading the metrics
A practical guide to KTO (Kahneman-Tversky Optimization) for aligning LLMs with binary desirable/undesirable feedback: the loss and its prospect-theor…
Read article →Kubeflow, in depth: Pipelines, Trainer v2 TrainJobs, Katib and Kueue on a shared GPU cluster
How Kubeflow turns Python into GPU pods: KFP components, artifacts and caching, Trainer v2 TrainJobs and runtimes, Kueue admission that prevents parti…
Read article →Kueue, in depth: job-level admission, quota, cohort borrowing, preemption and readiness for GPU training on Kubernetes
How Kueue queues GPU jobs on Kubernetes: ResourceFlavor, ClusterQueue, LocalQueue and Workload, v1beta2 YAML, queueing strategies, cohort borrowing an…
Read article →KV Cache Disk Offload, in depth: break-even arithmetic, chained chunk keys, NVMe layout, LMCache configuration and failure modes
When reading LLM KV cache from NVMe beats recomputing it: bytes per token, a break-even calculator, safe cache keys, O_DIRECT layout, a reference disk…
Read article →KV Cache Sizing for Deployments, in depth: bytes per token, the per-GPU budget, tensor parallelism, concurrency and worked examples
How to size the KV cache when deploying an LLM: the bytes-per-token formula for MHA, GQA, MLA and sliding-window layers, how the per-GPU budget is lef…
Read article →Liquid Cooling for GPU Datacenters, in depth: sizing the loop, the seconds between a pump fault and a throttled GPU, CDU telemetry and wiring cooling alarms into the scheduler
An operator's guide to liquid-cooled GPU clusters: the secondary loop as a system, heat-balance and flow sizing with worked numbers, capture rati…
Read article →LiteLLM, in depth: one OpenAI-compatible gateway over vLLM GPU pools and hosted models, with routing, fallbacks, virtual-key budgets and supply-chain hygiene
LiteLLM as GPU-serving infrastructure: SDK versus proxy, a config for two vLLM pools and a hosted fallback, routing strategies, retries, cooldowns and…
Read article →LLM A/B Testing, in depth: experimenting with quantisation, engines and GPU changes without fooling yourself
How to A/B test LLM serving changes on GPUs: user-level randomisation, per-arm replica pools that avoid batching interference, guardrail and quality m…
Read article →LLM Active Learning, in depth: acquisition signals from logprobs, LoRA ensembles, GPU selection and the cost of rescoring the pool
Active learning for LLM tasks as a GPU workload: label margin and truncated top-k entropy, multi-LoRA disagreement, a vLLM scoring sketch, two-stage u…
Read article →LLM Annotation Tools, in depth: task design, GPU pre-annotation, active learning, agreement and exporting SFT and DPO data
How human-in-the-loop annotation tools fit the GPU training pipeline: task shapes for LLM data, the six-part architecture, Label Studio, Argilla and P…
Read article →LLM Bill of Materials, in depth: the five layers of a served model, runtime capture on a GPU node, CycloneDX output, diffs and fleet queries
What a bill of materials for a served LLM must contain, from system prompt and chat template to kernels, CUDA libraries, driver and GPU; capturing it …
Read article →LLM Blue-Green Deployment, in depth: environment parity, an atomic switch, cold caches after cutover and the price of the rollback window
How to run blue-green releases for GPU-served LLMs: a hashed release manifest, parity checks for tokenizer, chat template and generation defaults, loa…
Read article →LLM Bottleneck Analysis, in depth: naming the bound with MFU and MBU, reading the right counters, and confirming with perturbation tests
A procedure for diagnosing slow LLM inference and training: the five bounds (compute, memory bandwidth, communication, host launch, capacity), MFU and…
Read article →LLM Canary Deployment
How to run a canary rollout for a model or inference-engine change: why the quality signal lags, sticky per-conversation traffic assignment, which met…
Read article →LLM Capacity Analysis, in depth: goodput at SLO from load-test sweeps, naming the binding resource, model ceilings and an honest replica count
How to analyse the capacity of an LLM serving deployment from evidence: open-loop sweeps, goodput at SLO, finding the knee, vLLM saturation metrics, f…
Read article →LLM Capacity Forecasting, in depth: forecasting tokens not requests, quantile peaks, backtesting and converting demand into replicas
How to forecast LLM serving capacity: measuring input and output tokens, decomposing growth and weekly shape, overlays for launches and mix shifts, a …
Read article →LLM Change Management, in depth: release manifests, risk classes, gates and rollback for GPU-served models
Change management for GPU-served LLMs: the layers that change outputs and capacity, a content-addressed release manifest, automatic risk classificatio…
Read article →LLM Chargeback, in depth: cost pools, billable units, budget rates, idle capacity policy, an append-only ledger and the monthly close
How to build internal chargeback for shared LLM GPU fleets: cost pools and overhead uplift, reserved GPU-hours versus weighted token units, budget rat…
Read article →LLM CI/CD Pipeline, in depth: the serving release unit, GPU test tiers, numerical parity and performance gates
How to build CI/CD for LLM serving on GPUs: a release manifest pinning image digest, engine, CUDA and host driver, weights and serving config; CPU and…
Read article →LLM Committed Use Discounts, in depth: commitment instruments, hourly matching, the quantile sizing rule, ladders and GPU-specific risks
How to buy and manage GPU commitment discounts for LLM fleets: resource-based and spend-based instruments on Google Cloud and AWS, how discounts are m…
Read article →LLM Content Moderation, in depth: a GPU classifier cascade for user content, one-token logprob scoring, threshold selection, capacity math and policy backfills
How to moderate user-generated content with LLMs at platform scale: a three-stage cascade of embedding heads, an LLM safety classifier scored from one…
Read article →LLM Cost Analysis, in depth: from a GPU-hour to cost per million tokens, per request and per break-even
How to turn a GPU-hour price into cost per million input and output tokens for LLM inference: bandwidth-bound decode, compute-bound prefill, KV cache …
Read article →LLM Cost Attribution, in depth: splitting shared GPU steps among batched requests, KV memory-time, prefix cache credit and reconciliation
How to attribute GPU inference cost to individual requests under continuous batching: a calibrated step-time model, compute and KV memory shares, a co…
Read article →LLM Data Curation Pipelines, in depth: running dedup, embedding and quality classifiers on GPUs at billion-document scale
How large-scale LLM data curation runs as a GPU job: stage placement, MinHash and LSH fuzzy deduplication with buckets-to-edges and connected componen…
Read article →LLM Data Flywheel, in depth: production signals, preference pairs, GPU budgets per turn and the gates that keep the loop honest
How to run an LLM data flywheel as an engineered loop: which user signals to capture and how they are biased, an event schema with consent, building p…
Read article →LLM Deployment Pattern Overview
The five LLM deployment topologies compared: dedicated single-tenant, pooled multi-tenant, serverless scale-to-zero, on-prem and hybrid burst. What is…
Read article →LLM Distillation Data, in depth: the teacher scoring pass as a GPU workload
How to produce token-level distillation data on GPUs: why teacher scoring is a prefill-only, compute-bound job, the logits memory wall and chunked on-…
Read article →LLM Disaster Recovery Drill, in depth: per-asset RTO and RPO, the weights and GPU bottleneck, restore validation and a scripted drill
How to run disaster recovery drills for LLM serving: scenarios from regional loss to account compromise, per-asset RTO and RPO, weight transfer arithm…
Read article →LLM Error Budgets, in depth: budget arithmetic, a written policy and cause attribution for GPU serving fleets
How to run error budgets for GPU-backed LLM serving: request versus token weighting, availability, latency and quality budgets, a three-state policy, …
Read article →LLM FinOps
Cost attribution for a shared GPU fleet: the GPU-hour to cost-per-1k-tokens unit model and why utilization sits in the denominator, the request-time m…
Read article →Flame Graphs for LLMs, in depth: CPU, wall and GPU-weighted stacks, folding Kineto traces, and finding hidden syncs in decode
Build flame graphs that answer LLM questions: py-spy for host and wait time, PyTorch export_stacks and a correlation-based trace fold for GPU time, a …
Read article →LLM Gameday, in depth: rehearsing GPU serving and training failures with hypotheses, safe fault injection and a scorecard
How to run gamedays for GPU-backed LLM systems: a failure catalogue, steady-state hypotheses as code, a replica-loss scenario with capacity and cold-s…
Read article →GitOps for LLM Deployments, in depth: pinning the release tuple, moving weights, slow-start health checks, drift and rollback
Run LLM model servers with Argo CD or Flux: a deploy repo layout, pinning engine digest and model revision, prefetching weights before rollout, startu…
Read article →LLM Guardrails
LLM guardrails viewed as a serving component rather than a policy document: where to place the input and output classifiers, what each placement costs…
Read article →LLM System Health Scoring, in depth: turning GPU, runtime and service signals into scores that route traffic, place jobs and trigger repair
How to design a health score for LLM serving replicas and training nodes: deriving it from the decisions it drives, signal categories, normalisation w…
Read article →Helm for LLM Deployments, in depth: model profiles and schemas, GPU-aware templates, probes sized to load time, weights outside the chart and safe upgrades
How to build a Helm chart for LLM serving: what a release stores, model profiles with a values schema and fail guards, a vLLM Deployment template with…
Read article →LLM Human-in-the-Loop, in depth: confidence signals from the serving stack, review routing, reviewer capacity and turning decisions into training data
How to run human review for LLM serving: logprob, verifier and self-consistency signals and their GPU cost, calibrated routing with an audit sample, r…
Read article →HITL Review Patterns, in depth: pre-delivery gates, judge pre-screens, stratified audits, reviewer agreement and shadow mode on a GPU serving stack
A catalogue of human review patterns for LLM output and what each costs the serving stack: buffered gates and regeneration with prefix caching, batche…
Read article →LLM Incident Communication Channels, in depth: channel topology, telemetry snapshots and model-fallback notices for GPU serving incidents
How to run communication for LLM serving incidents on GPU fleets: a channel topology including upstream providers and API consumers, a declare bot and…
Read article →LLM KPI Dashboards, in depth: a four-layer metric tree from cost per token down to SM activity, with PromQL, recording rules and the traps that make panels lie
How to build an LLM serving KPI dashboard that answers questions instead of decorating a wall: a four-layer metric tree from business KPIs to service …
Read article →LLM Individual Contributor KPIs, in depth: MFU with the attention term, goodput, cost per token at the SLO, and guard metrics that stop gaming
The few metrics an engineer on an LLM training, inference or eval team should own: MFU and HFU derived from first principles, goodput, cost per millio…
Read article →LLM Leadership KPIs, in depth: a six-number GPU fleet scorecard, ratio-of-sums rollups, commitment coverage and decision thresholds
The KPIs a platform leader needs to run an LLM GPU fleet: cost per qualified task, busy fraction of paid GPU-hours, commitment coverage, SLO attainmen…
Read article →LLM North Star Metric, in depth: qualified successful tasks, the factor tree that links model, serving and GPU capacity, and the counter-metrics that stop gaming
How to choose and govern one North Star metric for a GPU-served LLM product: why tokens and utilisation fail, a precise definition of qualified succes…
Read article →Kustomize for LLM Deployments, in depth: GPU components, patching server args safely, generators and pinned images, and render-time checks
Kustomize for GPU inference fleets: base, components and overlays, why strategic merge replaces vLLM args, JSON 6902 appends, label selector immutabil…
Read article →LLM Data Labeling, in depth: prefill-bound batch inference, prefix caching, constrained labels with logprobs, calibration and cascades
Labeling data with LLMs as a GPU workload: why it is prefill-bound, prefix caching, constrained single-token labels with vLLM structured outputs and l…
Read article →LLM Load Testing
How to load test a token-streaming LLM service correctly: why closed-loop generators can never overload a server, Little's Law and coordinated om…
Read article →LLM Serving Math 101, in depth: five numbers, a Little's law calculator and a full Llama 3.1 70B sizing example
The back-of-envelope math of LLM serving from first principles: weight bytes, KV bytes per token, prefill FLOPs, decode bytes per step and Little'…
Read article →LLM Memory Profiling, in depth: predict the GPU budget, measure each phase, and attribute every gigabyte before the OOM
A practical method for LLM GPU memory profiling: allocated vs reserved vs device memory, a five-term budget for training and inference, per-phase peak…
Read article →LLM Metrics Deep Dive, in depth: TTFT, ITL and queue histograms, KV cache saturation, DCGM and cardinality for GPU serving
How to measure LLM serving in production: vLLM and OpenTelemetry GenAI metric names, TTFT versus ITL versus TPOT, correct histogram_quantile aggregati…
Read article →NVIDIA Nsight for LLMs, in depth: NVTX for prefill and decode, capture windows, CUDA graphs, kernel families, decode GEMMs and tensor-parallel all-reduce
A working method for profiling LLM training and serving with Nsight Systems and Nsight Compute: annotating steps and layers with NVTX, capturing stead…
Read article →LLM On-Call Playbook, in depth: the first fifteen minutes, a five-layer triage tree, a mitigation ladder and sustainable GPU serving shifts
A responder's playbook for LLM serving on GPUs: the first fifteen minutes, a five-layer triage tree from gateway to model output, a Prometheus tr…
Read article →LLM Output Validation, in depth: constrained decoding on the GPU, token masks, retry economics and semantic validators
How LLM output validation works in a GPU serving stack: grammar-constrained decoding and token bitmasks, why mask computation overlaps the forward pas…
Read article →LLM Ownership Matrix, in depth: a RACI over the GPU serving stack, ownership as code with one Accountable per row, and wiring it into reviews, alerts and incidents
How to build an ownership matrix for a self-hosted LLM serving stack: RACI rules, a matrix from weights to quotas, a YAML ownership file with a CI val…
Read article →LLM On-Call Rotation, in depth: page-load arithmetic, rotation shapes, escalation and handoffs for GPU serving teams
Designing on-call for GPU LLM serving: what makes it different, sizing with page-load arithmetic, rotation shapes compared, a rotation generator and b…
Read article →LLM Penetration Testing, in depth: assessing the self-hosted GPU inference stack below the prompt
A defensive penetration-testing plan for self-hosted LLM infrastructure: scoping and rules of engagement, mapping the serving stack, auditing exposed …
Read article →LLM Performance Profiling, in depth: decode and prefill budgets, MBU and MFU, torch.profiler, Nsight timelines and a worked example
A top-down method for profiling LLM inference: compute bandwidth and FLOP floors for decode and prefill, measure TTFT and inter-token latency correctl…
Read article →LLM Pricing Models, in depth: input and output meters, cache writes and reads, reasoning tokens, batch and flex tiers, length tiers and committed capacity
How LLM API prices are actually computed: why output costs more than input, cache write surcharges and read discounts with break-even hit rates, hidde…
Read article →LLM Health Probes, in depth: startup, readiness and liveness for GPU inference servers without restart storms
How to design Kubernetes startup, readiness and liveness probes for LLM inference servers: sizing startup budgets from measured load times, keeping lo…
Read article →LLM Prompt Injection Defense, in depth: serving-side controls with windowed scanners, spotlighting cost, constrained tool calls, quarantined pools and cache isolation
Prompt injection defences that live in the GPU serving stack: provenance tags, windowed Prompt Guard 2 scanning for its 512-token limit with capacity …
Read article →LLM Python Frameworks, in depth: how LangChain, LlamaIndex and DSPy drive a self-hosted GPU server, and how to stop them wasting it
What Python LLM application frameworks make the GPU do: connecting LangChain, LlamaIndex and DSPy to an OpenAI-compatible vLLM or SGLang server, mappi…
Read article →PyTorch Profiler for LLMs, in depth: schedules, key_averages, traces, memory snapshots and multi-GPU profiling
A working manual for torch.profiler on LLM training and inference: what it records, the step schedule, reading key_averages by self time and input sha…
Read article →LLM Ranking and Comparison UI, in depth: blind side-by-side serving on GPUs, vote logging, Bradley-Terry ranks with intervals, style control and anti-gaming
How to build a blind side-by-side LLM comparison system: the four guarantees, concurrent generation with identical parameters and a parity buffer, GPU…
Read article →LLM Performance Regression Analysis, in depth: noise floors, bootstrap comparisons, bisecting and kernel profile diffs
How to detect and localise LLM serving performance regressions: mapping TTFT and TPOT to prefill and decode, the usual causes from drivers to chat tem…
Read article →LLM Reserved Capacity, in depth: capacity guarantees versus discounts, sizing from demand, filling the pool and surviving the end date
How reserved GPU capacity works for LLM training and serving: capacity reservations versus billing commitments, Capacity Blocks and calendar-mode rese…
Read article →LLM Rollback Strategies, in depth: where rollback time goes on GPUs, the serving tuple, automatic triggers, adapter rollback and the cost of a warm pool
How to roll back an LLM serving release on GPUs: weight transfer, engine startup and cold prefix caches, rolling back the full serving tuple, an autom…
Read article →LLM Routing Strategies, in depth: model-tier routing, cascades and KV-cache-aware replica routing
The two LLM routing problems and how to solve each: rules, classifiers, preference-trained routers such as RouteLLM and cascades for choosing a model;…
Read article →LLM Runbook Index, in depth: a symptom-keyed catalogue for GPU fleets, alert wiring, CI lint and the entries that matter
How to build a runbook index for GPU LLM serving and training: runbooks as data in git, an entry schema, symptom-keyed catalogue with Xid 79, 48, 63, …
Read article →GPU Serving Architecture for LLMs: The Full Stack
An architecture map of a GPU LLM serving stack: ingress and admission, the scheduler, the model executor and its tensor/pipeline layout, the KV cache …
Read article →LLM Serving Stacks
How to choose an LLM serving engine: the axes that actually differentiate vLLM, TGI, TensorRT-LLM and SGLang - scheduling model, KV-cache management, …
Read article →LLM Shadow Deployment, in depth: mirroring live traffic to a candidate model without side effects, wasted GPUs or misleading diffs
How to shadow-deploy an LLM: gateway mirroring, inert tool calls, comparing nondeterministic outputs, cache and batching confounders on GPUs, a worked…
Read article →LLM SLO Burn Rate Alerts, in depth: TTFT and inter-token SLIs, multiwindow rules for GPU serving, and what really burns the budget
How to page on error-budget burn for LLM inference: defining good and bad events for streaming responses, TTFT and inter-token latency SLIs from vLLM …
Read article →LLM Spend Alerts, in depth: gateway metering, multiwindow burn rates in dollars, idle GPU spend and an enforcement ladder
Designing spend alerts for LLM APIs and self-hosted GPU fleets: why billing data is too late, meter events with a versioned price table, threshold, bu…
Read article →Python LLM Stack Overview, in depth: from pip install to GPU kernels, the version contract, memory arithmetic, fine-tuning and serving layers
A first-principles map of the Python LLM stack: how a PyTorch call becomes a GPU kernel, the driver and CUDA wheel version contract, what transformers…
Read article →LLM Status Page, in depth: components customers can act on, mapping GPU and serving signals to status, quality degradation, and a feed that survives the outage
Designing a status page for GPU-backed LLM serving: components by surface, model and region, thresholds for errors, TTFT, capacity rejections and qual…
Read article →Synthetic Data Generation for LLMs on GPUs, in depth: cost per accepted sample, parallel layout, offline vLLM batches, judging, dedup and resumable jobs
Running LLM synthetic data generation as a GPU workload: why decode and acceptance rate drive cost, a GPU-hours budget model, tensor-parallel versus r…
Read article →LLM Synthetic Monitoring, in depth: outside-in probes for GPU serving, cache-proof prompts, correctness checks and quorum alerts
How to build synthetic monitoring for an LLM serving fleet: client-side TTFT and decode timing, probe prompt design that defeats prefix caching, a Pyt…
Read article →LLM TCO, in depth: a three-year total cost of ownership model for API, cloud and owned GPU serving
How to build a three-year total cost of ownership model for an LLM workload: demand growth, measured GPU throughput, capacity for peak, API versus on-…
Read article →LLM Trace Analysis, in depth: reading PyTorch profiler traces to find idle GPUs, exposed communication and launch overhead
How to capture and analyse PyTorch profiler (Kineto) traces of LLM training and inference: the trace format, a capture schedule that measures steady s…
Read article →LLM Uptime Calculation, in depth: good-event definitions, three measurement methods, series and k-of-n composition, and GPU failure maths
How to calculate availability for an LLM endpoint: defining a good request with TTFT and output checks, time vs request vs probe availability, series,…
Read article →LMDeploy, in depth: TurboMind, blocked KV cache sizing, AWQ and KV quantization, and serving it in production
How LMDeploy serves LLMs: TurboMind and PyTorch engines, persistent batching, the cache_max_entry_count share of free memory, worked KV sizing for an …
Read article →Lookahead Decoding
Draft-model-free speculative decoding on the GPU: Jacobi-style parallel decoding, lookahead branches and the n-gram pool, prompt-lookup drafting that …
Read article →LoRAX, in depth: serving hundreds of LoRA adapters on one base model with SGMV, exchange scheduling and tiered weight caching
How LoRAX serves many LoRA fine-tunes from one GPU deployment: router and per-adapter queues, heterogeneous continuous batching with SGMV kernels, ada…
Read article →Megatron-LM, in depth: how tensor, pipeline, data and context parallelism compose, sizing a 70B run on 512 GPUs, and the flags that matter
Megatron-LM and Megatron-Core explained for practitioners: the parallel dimensions and their communication, the tp-cp-ep-dp-pp rank layout, a 512-GPU …
Read article →GPU Memory Hierarchy
The GPU memory hierarchy as one system: register file, shared memory and L1, L2, and HBM. How capacity rises while bandwidth falls at each level, why …
Read article →CUDA Memory Pool, in depth: the stream-ordered allocator, release thresholds, cross-stream reuse, CUDA graphs and PyTorch's caching allocator
How GPU memory pools work: why cudaMalloc and cudaFree are slow, the CUDA stream-ordered allocator with cudaMallocAsync and explicit pools, the releas…
Read article →AMD MI300X, in depth: eight XCDs, private 4 MB L2 slices, a 256 MB Infinity Cache and how kernels should use them
MI300X from the kernel's point of view: the XCD and I/O die hierarchy, round-robin workgroup dispatch and a program-id remap for L2 reuse, wave64…
Read article →AMD MI325X, in depth: 256 GB of HBM3E on MI300X compute, KV-cache planning, FNUZ FP8, partitioning and porting traps
What the AMD Instinct MI325X changes for software: the same CDNA 3 compute as MI300X with 256 GB HBM3E at 6 TB/s and 1,000 W, a worked KV-cache and de…
Read article →NVIDIA MIG architecture
How Multi-Instance GPU partitions a datacenter GPU: SM slices, L2 slices and the memory-controller path that makes isolation a hardware property rathe…
Read article →Mixed Precision Training, in depth: autocast, GradScaler, BF16 versus FP16 on GPUs, FSDP policies and debugging NaNs
How mixed precision training actually runs on NVIDIA GPUs with PyTorch: what autocast casts and caches, when you need GradScaler, choosing FP16, BF16 …
Read article →ML Framework Comparison, in depth: PyTorch, JAX and TensorFlow by execution model, compilation and scaling
PyTorch, JAX and TensorFlow compared the way that matters on GPUs: eager dispatch versus traced functions, torch.compile versus jax.jit versus tf.func…
Read article →GPU ML Ops, in depth: running a GPU fleet as a control loop, from node lifecycle and failure triage to checkpoint intervals and goodput
How to operate a GPU fleet for training and serving: the node lifecycle state machine, burn-in gates, classifying faults into drain or ignore, the ari…
Read article →ML Reproducibility on GPUs, in depth: floating-point order, PyTorch determinism controls, seeded data pipelines, distributed reductions and a bitwise replay test
Why identical GPU training runs diverge and how to stop it: non-associative float sums, atomics, cuDNN autotuning, TF32, seeded DataLoaders, RNG state…
Read article →MLflow for GPU Training Ops, in depth: server topology, rank-0 logging, metric volume, system metrics, checkpoints and the model registry
How to run MLflow underneath multi-GPU and multi-node training without slowing it down or losing data: tracking server topology, which rank logs, why …
Read article →Multimodal LLM Training, in depth: the encoder-connector-decoder pipeline, staged freezing, packing and the GPU load-balancing problem
How vision-language models are trained on GPUs: vision encoder, connector and LLM decoder, how many tokens an image becomes at fixed and dynamic resol…
Read article →MLX, in depth: unified-memory arrays, lazy evaluation, composable transforms, quantisation, mlx-lm fine-tuning and custom Metal kernels
How Apple MLX works and how to use it well: arrays without devices, lazy evaluation and mx.eval, grad, vmap and compile, a training loop, affine quant…
Read article →MoE All-to-All Communication
The MoE dispatch and combine all-to-all as a communication primitive: why every expert-parallel layer pays for two of them, what each one moves, why t…
Read article →MoE Expert Parallelism Deployment, in depth: sizing, multi-node launch, network prerequisites, failure domains and rollout
How to deploy a large mixture-of-experts model with wide expert parallelism across nodes: sizing experts and KV memory per GPU, the divisibility rule …
Read article →Grouped GEMM for MoE, in depth: jagged expert batches in one launch, persistent tile scheduling, alignment, the backward pass and choosing a kernel
How grouped GEMM runs every MoE expert's matmul in a single kernel: offsets and problem descriptors, persistent tile scheduling, alignment and ti…
Read article →MoE Kernels, in depth: the six kernels of an expert layer, block alignment, fused expert GEMMs, the decode memory wall and profiling
How a mixture-of-experts layer actually runs on a GPU: fused router top-k, block-aligned sorting with padding, gathers folded into a grouped expert GE…
Read article →MoE Load Balancing, in depth: measuring imbalance, capacity factors, auxiliary losses, DeepSeek-V3's bias controller and expert replication
How mixture-of-experts load balancing works in practice: why routing collapses, how to measure imbalance per expert and per GPU, capacity factors and …
Read article →MoE Routing Math, in depth: scoring recipes, where the gradient flows, capacity and drop arithmetic, balance losses and the dispatch index math on the GPU
The arithmetic of mixture-of-experts routing as GPUs execute it: softmax and sigmoid scoring in Switch, Mixtral and DeepSeek-V3, gradients through gat…
Read article →Mixture-of-Experts Serving Architecture in Depth
How mixture-of-experts models actually behave on serving GPUs: expert placement and expert parallelism, the dispatch and combine all-to-all that domin…
Read article →MSCCL, in depth: custom GPU collective algorithms with MSCCLang, the XML IR and the selection rules that decide whether they run
How MSCCL runs custom collective schedules on top of NCCL: MSCCLang programs, the compiler and XML IR, synthesis, the runtime selection predicate behi…
Read article →Multi-Instance GPU Deployment, in depth: profiles, mig-parted geometries, Kubernetes strategies, bin-packing, monitoring and safe reconfiguration
How to deploy NVIDIA MIG in production: GPU and compute instances, reading the driver's profile tables, manual setup with nvidia-smi, declarative…
Read article →Multi-LoRA Serving, in depth: what the GPU does when one batch carries many adapters, gathered shrink and expand kernels, adapter memory arithmetic and the decode cost model
A GPU-level guide to serving many LoRA adapters on one base model: the unmerged LoRA computation, why grouping by adapter fails, gathered shrink and e…
Read article →Multi-Provider LLM Strategy, in depth: task-named APIs, adapters, safe failover, capacity arithmetic and self-hosted GPUs as a provider
How to run LLM workloads across several providers and your own GPUs: goals and their designs, a task-named internal API, adapters and circuit breakers…
Read article →Multi-Stream Execution, in depth: stream semantics, events, copy-compute overlap, PyTorch side streams, allocator hazards and proving overlap in a trace
A practical guide to CUDA multi-stream execution: what streams guarantee, ordering with events, a double-buffered CUDA pipeline, a PyTorch prefetcher …
Read article →Multi-Turn KV Cache, in depth: keeping a conversation's attention state alive across turns, think time, routing and prompt rendering
How multi-turn KV cache reuse works in LLM serving: why each chat turn re-sends the whole history, the prefill arithmetic of reuse versus recompute, K…
Read article →Multimodal Model Serving, in depth: the request path, media token accounting, CPU preprocessing, encoder scheduling and caching, KV budgets and encoder disaggregation
How to serve vision and audio language models in production: the request path, computing image and video token costs, safe CPU preprocessing, encoder …
Read article →NCCL Collectives Explained: All-Reduce, Ring vs Tree, Topology
How NCCL runs multi-GPU collectives: all-reduce and other ops, ring vs tree algorithms, NVLink, NVSwitch and InfiniBand topology, GPUDirect RDMA, tuni…
Read article →NVIDIA Nemotron, in depth: lineage, hybrid Mamba-Transformer MoE layers, KV memory math, FP8 and NVFP4, pruning and serving
A technical guide to NVIDIA Nemotron models: the lineage from Nemotron-4 340B to Nemotron 3, the hybrid Mamba-Transformer MoE layer pattern, KV cache …
Read article →nvidia-smi
How nvidia-smi provides quick GPU status, and the key views for debugging.
Read article →NVLink and NVSwitch Explained: The Multi-GPU Scale-Up Fabric
How NVLink and NVSwitch connect GPUs: memory-semantic links, the NVSwitch crossbar, peer-to-peer access, NVLink vs PCIe bandwidth, topology and collec…
Read article →NVLink Switch, in depth: rack-scale NVLink domains, in-switch reduction and the software that makes them work
How the NVLink Switch turns NVLink into a rack-scale fabric: the GB200 NVL72 wiring of 72 GPUs to 18 switch chips, NVLink SHARP and NCCL's NVLS a…
Read article →Object Detection Serving, in depth: GPU decode, letterboxing, TensorRT engines with NMS, Triton dynamic batching and capacity planning
How to serve object detectors on GPUs: what YOLO-style heads output, confidence filtering and NMS, letterbox coordinate mapping, GPU decode with nvJPE…
Read article →GPU Occupancy: Warps per SM, Registers, Shared Memory
What GPU occupancy measures: resident warps per SM against the maximum, the register, shared memory and block-slot limits, a worked example, when low …
Read article →Ollama on the GPU, in depth: VRAM budgets, layer offload, multi-GPU placement, KV cache precision and measuring decode speed
How Ollama uses a GPU: the VRAM budget of weights, KV cache and compute buffers, why a partial CPU offload collapses decode speed, how models are plac…
Read article →Online DPO, in depth: on-policy pairs every step, reward models and LLM judges in the loop, vLLM colocation, weight sync and the GPU budget
Online DPO (online AI feedback) on GPUs: how it differs from offline and iterative DPO, the loss with a worked number, a minimal PyTorch step, reward …
Read article →ONNX Runtime, in depth: execution providers, graph partitioning, IOBinding, CUDA graphs, TensorRT and quantization
How ONNX Runtime executes a model: graph optimisation levels, how execution providers claim nodes, memcpy at device boundaries, CUDA EP options, IOBin…
Read article →OpenAI Realtime API, in depth: transports, events, turn detection, barge-in and tool calls for voice agents
How the OpenAI Realtime API works: WebRTC, WebSocket and SIP transports, ephemeral client secrets, the session and event model, server and semantic VA…
Read article →OpenRouter, in depth: model and provider routing, the provider object, fallbacks, streaming errors and production operation
How OpenRouter routes a request to a model and a GPU provider, the documented provider fields for data policy, precision and price, fallback behaviour…
Read article →OpenVINO for LLMs, in depth: export and weight compression, a memory budget for an 8B model, and running on Intel CPUs, GPUs and NPUs
How OpenVINO runs language models: Optimum Intel export, stateful IR, int4 compression flags, an 8B memory and bandwidth budget, LLMPipeline code, CPU…
Read article →Hugging Face Optimum, in depth: the package split, what the ONNX exporter does, O1-O4 graph optimization, int8 quantization and when not to use it
A practical guide to Hugging Face Optimum: which package targets which hardware, how the ONNX exporter infers tasks and validates graphs, ORTModel and…
Read article →GPU Orchestration Overview, in depth: inventory, quota, gang and topology-aware placement, failure handling, and Kubernetes versus Slurm
How GPU clusters decide which jobs run where: device plugins and Kubernetes DRA, Kueue and Slurm gang admission, topology-aware placement with tested …
Read article →ORPO, in depth: the odds-ratio objective, one training stage without a reference model, and what it costs on the GPU
Odds Ratio Preference Optimization from first principles: how ORPO folds preference alignment into supervised fine-tuning, the loss and its gradient, …
Read article →P2P KV Cache Transfer
How the KV cache actually moves from a prefill GPU to a decode GPU: bytes-per-token arithmetic under GQA, layer-wise streaming that overlaps transfer …
Read article →GPU P2P Memory Access, in depth: enabling peers, copies versus direct loads, cross-device ordering, CUDA IPC and proving the fast path
How GPU peer-to-peer memory access works in software: UVA and peer mappings, cudaDeviceEnablePeerAccess per direction, cudaMemcpyPeerAsync versus dire…
Read article →Page-Locked (Pinned) Host Memory, in depth: why DMA needs it, the CUDA allocation and registration APIs, PyTorch pin_memory and non_blocking, and pinning without starving the node
How pinned host memory works and how to use it: why DMA cannot read pageable pages, cudaHostAlloc and cudaHostRegister flags, mapped and write-combine…
Read article →Paged KV cache architecture
Paged KV cache as GPU memory management: why contiguous per-sequence allocation fragments HBM, how fixed-size blocks and a block-table indirection fix…
Read article →PCIe host interconnect architecture
How the PCIe host interconnect really behaves for GPUs: lanes and generation scaling, the host-to-device transfer path, why pinned memory matters, cop…
Read article →Prefill/Decode Disaggregation Architecture in Depth
Why LLM prefill and decode have opposite hardware profiles, how they interfere when they share a GPU, what splitting them into separate pools actually…
Read article →Pipeline Parallelism: GPipe vs 1F1B and the Pipeline Bubble
How pipeline parallelism trains big models across GPUs: layer stages, microbatches, GPipe vs 1F1B schedules, the bubble fraction, activation stashing,…
Read article →GPU Datacenter Placement Strategy, in depth: matching training and inference workloads to sites by power, latency, data and risk
How to decide where GPU capacity should live and which workloads go where: workload classes and their constraints, power availability as the gating fa…
Read article →GPU Pod Network, in depth: scale-up and scale-out domains, rail-optimised fabrics, mapping parallelism and pod bring-up
How a GPU pod network is designed and validated: the four networks in a pod, NVLink scale-up versus 400 Gb/s scale-out bandwidth, rail-optimised fat-t…
Read article →Portkey AI Gateway, in depth: config trees, fallback and load-balance semantics, retries, guardrail status codes and routing to self-hosted GPU pools
How the open-source Portkey AI Gateway evaluates its routing config: strategy modes, inheritance, the exact fallback stop condition, weighted load bal…
Read article →PPO for LLMs, in depth: implementing the training step, from tensor masks and per-token rewards to the diagnostics that tell you a run is going wrong
An implementer's guide to PPO for language models on GPUs: tensor shapes and response masks, per-token KL rewards with the score on the last toke…
Read article →GPU Preemption Handling
Taking work back from a GPU that is already running it: request-level preemption inside an inference server, victim selection and the swap-versus-reco…
Read article →Prefill vs Decode Split, in depth: measuring each phase, a fitted decode cost model, and choosing a split from your own traffic
How to decide between colocated, chunked and disaggregated LLM serving from evidence: an open-loop streaming harness for TTFT and inter-token latency,…
Read article →Prefix Caching in Depth: What the GPU Skips, What It Still Pays For, and How Cached KV Blocks Live and Die in HBM
Prefix caching from the GPU's point of view: prefill FLOPs and what a cache hit removes, the full-block rule, reference-counted KV blocks and the…
Read article →GPU Profiling
How the Nsight tools actually work: what each Nsight Systems timeline row records, how correlation links a CUDA API call to its kernel, reading gaps, …
Read article →LLM Provider Failover, in depth: error classification, health-scored routing, deadline budgets, model equivalence, mid-stream failures and a warm backup
How to build failover across LLM providers that works under a real outage: classifying errors into retry, fail over or fail fast, per-target circuit b…
Read article →PUE, in depth: meter boundaries, partial-load behaviour of GPU halls, pPUE and WUE, the reporting rules, and charging facility energy to a training job
Power usage effectiveness for GPU datacenters, measured properly: where the meters sit, interval versus annual PUE, why lightly loaded AI halls score …
Read article →PyTorch Lightning, in depth: the module and Trainer split, hook order, FSDP strategies, memory arithmetic, checkpoints and Fabric
A practical guide to PyTorch Lightning 2.x for GPU training: what the LightningModule and Trainer each own, a full fine-tuning module, the hook order,…
Read article →QPS Math for LLM Serving, in depth: from measured slots and service time to queueing tails, pooled routing, goodput and fleet size
Turn measured TTFT and TPOT into requests per second: slots and service time, Little's law, mixed workloads, Erlang C tail waits in code, pooled …
Read article →Ray for GPU Workloads, in depth: tasks, actors, GPU scheduling, placement groups, Ray Data and Ray Train
How Ray runs GPU work from first principles: logical GPU resources and CUDA_VISIBLE_DEVICES, fractional GPUs, placement groups for gang scheduling, th…
Read article →Ray Serve for LLM, in depth: from LLMConfig to GPU placement groups, replica sizing, multi-LoRA and failure modes
How Ray Serve LLM turns one LLMConfig into replicas on GPUs: placement group bundles and strategies for tensor and pipeline parallelism, a worked KV c…
Read article →GPU Regulatory Landscape, in depth: how export controls measure chips, how AI laws measure training compute, and the compliance engineering that keeps a GPU team shippable
An engineer's guide to AI hardware regulation at the end of September 2026: 3A090 total processing performance and performance density, the Janua…
Read article →Replicate, in depth: predictions, Cog packaging, deployments, signed webhooks, cold starts and cost control
How Replicate works as production infrastructure: models, versions and the prediction lifecycle, packaging with Cog, the three billing models and cold…
Read article →Response Caching for LLM APIs, in depth: exact-match keys, determinism, streaming, invalidation and stampede control
How to build an application-side response cache in front of an LLM API: what must go into an exact-match cache key, request normalisation, the determi…
Read article →GPU Pricing: Retail vs Cloud vs Enterprise, in depth: what each channel's price buys, licence and capacity constraints, hidden costs, and normalising to dollars per useful GPU-hour
How to compare GPU prices across retail cards, cloud instances and enterprise servers: what each price includes, the GeForce datacenter licence clause…
Read article →Reward Model Training on GPU, in depth: the paired training step, last-token pooling, sizing, evaluation and the bugs that silently ruin a reward model
How to implement and run reward model training on GPUs: architecture, the pairwise loss with worked numbers, a PyTorch training step, the padding and …
Read article →Ring Attention
Ring attention explained from the GPU side: sharding the sequence across devices, rotating KV blocks around a ring so every query sees every key, merg…
Read article →RLHF Pipeline on GPU
RLHF and post-training as a GPU systems problem: holding a policy, reference, reward and critic model in one memory budget, the generation-versus-trai…
Read article →RoCE (RDMA over Ethernet), in depth: verbs, queue-pair states, go-back-N retransmission, timeouts and debugging a hung all-reduce
How the RoCE reliable transport behaves under GPU training traffic: verbs objects and GPU memory registration, the RC queue-pair state machine with re…
Read article →RoCE v2, in depth: running GPU training traffic over routed Ethernet with GID selection, PFC, ECN, DCQCN and NCCL tuning
How RDMA over Converged Ethernet v2 carries GPU collectives: the UDP 4791 encapsulation and ECMP entropy, GID selection, lossless versus lossy designs…
Read article →Run:ai, in depth: quotas, over-quota borrowing, preemption, gang scheduling and fractional GPUs on Kubernetes
How NVIDIA Run:ai and the open-source KAI Scheduler share GPUs between teams: projects and queues, fair share, priorities, podgroups, fractions, a wor…
Read article →GPU scheduling architecture
How work actually gets onto a GPU and shares it: the hardware work distributor and warp scheduler, stream priorities, why concurrent kernel execution …
Read article →GPU Scheduling With Slurm
How Slurm schedules GPUs on an HPC cluster: GRES configuration, request flags, cgroup enforcement of CUDA_VISIBLE_DEVICES, affinity, partitions and Qo…
Read article →Segmentation Model Serving, in depth: output arithmetic, GPU post-processing, mask encoding, promptable encoder caches and tiling
How to serve semantic, instance and promptable segmentation models on GPUs: output size arithmetic, GPU upsample and argmax, RLE and PNG mask encoding…
Read article →GPU Selection in Depth: A Decision Method Built on Memory, Bandwidth, Compute and Cost per Unit of Work
How to choose a GPU for training or inference from first principles: memory-capacity math for weights, KV cache and optimizer state, why decode speed …
Read article →Active-Passive LLM Serving, in depth: standby tiers, cold-start budgets, fenced promotion and in-flight streams
Active-passive LLM serving on GPUs: hot, warm, pilot-light and cold standby tiers, vLLM sleep mode, a stage-by-stage cold-start budget, generating pro…
Read article →Anycast for LLM Serving, in depth: BGP catchments, long-lived token streams, the PoP-terminated front door, health-driven announcements and gentle traffic engineering
How to put an LLM API behind BGP anycast: catchments and ECMP, why long token streams break, anycast at PoPs with unicast GPU regions, resumable strea…
Read article →GPU Serving Capacity Estimation, in depth: memory budgets, decode bandwidth, prefill cost and the latency-bounded request rate
How to estimate the request rate one GPU replica can serve for an LLM: KV pool and concurrency, decode step time from HBM bandwidth, prefill cost from…
Read article →DNS Failover for LLM Serving, in depth: health-checked records, the real failover timeline, GPU-aware health probes and warm standby
How DNS failover works for LLM inference: failover record pairs, the detection, TTL and client-cache timeline, a cached synthetic-generation health pr…
Read article →LLM Serving DR, in depth: DR tiers costed in GPUs, the release manifest and parity check, where standby GPUs come from, a degraded-mode ladder and the declare-disaster rule
Disaster recovery for LLM inference fleets: events to plan for, backup, pilot light, warm standby and active-active priced in GPUs, digest-based repli…
Read article →Edge PoP LLM Serving, in depth: what the edge saves, small models and caches near users, capacity fragmentation and model rollout
How to serve LLM traffic from edge points of presence: TTFT versus decode latency, which work belongs at the edge, a vendor-neutral edge router with t…
Read article →Geographic Routing for LLM Serving, in depth: residency filters, TTFT-based region scoring, cache affinity and spillover with hysteresis
How to choose a serving region per LLM request: why nearest-region routing fails under GPU queueing, the TTFT budget, residency and model filters, loa…
Read article →Multi-Region LLM Serving
Why and when to run LLM inference in more than one region: latency, data residency, GPU capacity access and blast radius; the cost of idle standby acc…
Read article →LLM Serving Reliability, in depth: failure domains, truthful health probes, preemption storms, retry budgets, graceful drains and stuck GPUs
How to keep LLM serving reliable: a failure-domain map, liveness versus readiness versus deep probes for vLLM, KV pressure and preemption alerts, per-…
Read article →LLM Serving SLOs, in depth: choosing TTFT, TPOT, ITL and E2E targets, joint attainment versus percentiles, goodput and the load sweep
How to define LLM serving SLOs that match what users feel: per-class SLIs and targets, why marginal p90s overstate attainment (worked example), goodpu…
Read article →SGLang, in depth: RadixAttention, cache-aware scheduling, structured decoding and running it in production
How SGLang serves LLM programs: the process architecture, the RadixAttention prefix tree and eviction, scheduling policies, the frontend language, con…
Read article →GPU shared memory architecture
Deep-dive on GPU shared memory: why the memory hierarchy makes on-SM shared memory the linchpin of performance, the 32-bank structure and how warp-wid…
Read article →Shared Prefix KV Caching, in depth: reading the prefix once per batch with cascade attention, the bandwidth arithmetic and when it pays
Shared-prefix attention for LLM decode: why prefix caching stores the prefix once but reads it per sequence, the log-sum-exp split behind Hydragen and…
Read article →SLO-Aware Scheduling
Deadline-aware ordering of already-admitted LLM requests: deriving a per-request deadline from a TTFT or TPOT budget, earliest-deadline-first and why …
Read article →S-LoRA, in depth: serving thousands of LoRA adapters with unified paging, adapter prefetch, early abort and a LoRA-aware tensor-parallel layout
How S-LoRA serves thousands of LoRA adapters on one base model: why unmerged serving wins, unified paging of adapters and KV cache with worked memory …
Read article →Speculative Decoding Architecture in Depth
The GPU-execution view of speculative decoding: why decode is memory-bandwidth bound, how verifying k tokens raises arithmetic intensity and turns a G…
Read article →GPU Spot Instances, in depth: how clouds sell and reclaim spare GPUs, the real price per useful hour, and making training and inference survive it
How AWS, Google Cloud and Azure sell spot GPU capacity and reclaim it, the notice signals and metadata endpoints, why GPU spot differs from CPU spot, …
Read article →Stable Diffusion Training + Inference, in depth
Latent diffusion on GPUs: VAE, text encoders, UNet vs MMDiT, epsilon, v and rectified-flow objectives, a training step with min-SNR, memory budgets, g…
Read article →Static Batching
Static batching on the GPU: why fixed-shape batches lose badly for online serving but still win for offline and batch jobs. The slowest-sequence probl…
Read article →CUDA Stream Capture, in depth: recording stream work into graphs, fork and join, capture modes, invalidation, memory and graph updates
How CUDA stream capture turns stream code into a graph: what gets frozen, cross-stream fork and join, what Global, ThreadLocal and Relaxed modes reall…
Read article →StreamingLLM
StreamingLLM keeps a transformer decoding forever in constant KV memory by pinning a handful of attention-sink tokens and rolling a recent window over…
Read article →Streaming LLM Serving, in depth: SSE token delivery, TTFT and inter-token latency, cancellation, backpressure and proxies
How token streaming really works between the GPU and the user: TTFT versus inter-token latency, the SSE wire format, incremental detokenisation with s…
Read article →GPU Supply Chain, in depth: from wafer to rack, where the bottlenecks sit, and how to engineer training around the capacity you actually get
How data-center GPUs are made and delivered: logic die, HBM, advanced packaging, substrates, modules, servers, networking and power; why yield compoun…
Read article →Switch Transformer Architecture
The architectural lineage of sparse mixture-of-experts models: what GShard's top-2 routing established, why Switch argued top-1 is enough and wha…
Read article →GPU Tail Latency Management, in depth: where p99 comes from in LLM serving, measuring it honestly and the levers that move it
Why tail latency dominates LLM and agent workloads, the queueing math, nine causes of slow requests on GPU serving replicas with their signatures, an …
Read article →What Is a Tensor Core? GPU Tensor Core Architecture Explained
What GPU tensor cores are: warp-level matrix multiply-accumulate, WMMA and MMA, narrow inputs with wide accumulators, fp16, bf16, fp8, 2:4 sparsity.
Read article →GPU Tensor Cores: How GEMM Tiling Keeps Them Fed
Why tensor cores make most GPU kernels memory-bound and how GEMMs keep them fed: arithmetic intensity, the warp tile hierarchy, wave quantization and …
Read article →Tensor Parallelism in Depth: Splitting Transformer Layers Across GPUs, the Communication It Costs, and How to Run It
How tensor parallelism splits individual transformer layers across GPUs: the column and row matrix splits, Megatron's f and g operators, attentio…
Read article →TensorRT-LLM, in depth: the PyTorch-based runtime, in-flight batching, paged KV cache, quantization, parallelism and how to run it
A practical guide to NVIDIA TensorRT-LLM as it works today: why the engine-build workflow is gone, how the scheduler and paged KV cache use GPU memory…
Read article →TensorFlow Lite for LLM, in depth
How LLMs run on TensorFlow Lite, now LiteRT: flatbuffers and delegates, prefill and decode signatures with an explicit KV cache, conversion with liter…
Read article →GPU Thermal Management, in depth: heat as power, the clock-management loop, reading throttle reasons, and why one hot GPU slows a whole training job
How GPU temperature turns into lost training throughput: dynamic and leakage power, the thermal resistance chain and its time constants, how firmware …
Read article →GPU Throughput Math for LLM, in depth: three ceilings, decode batch curves, Little's law and cost per million tokens
A first-order throughput model for LLM serving and training: compute, bandwidth and capacity ceilings, why decode stays bandwidth-bound, a runnable es…
Read article →GPU Time Slicing
GPU time slicing on Kubernetes from a cluster-operations point of view: how the device plugin advertises one physical GPU as N schedulable replicas, w…
Read article →GPU Tensor Memory Accelerator (TMA): Architecture Deep-Dive
How Hopper's Tensor Memory Accelerator moves tensor tiles asynchronously via descriptors and mbarriers, freeing compute warps and enabling deep p…
Read article →Together AI, in depth: serverless, dedicated endpoints, batch and fine-tuning, with integration code and a break-even model
An engineer's guide to Together AI: how serverless, dedicated endpoints, batch and fine-tuning map onto GPU economics, OpenAI-compatible client c…
Read article →GPU Topology Awareness, in depth: reading the node graph, how NCCL detects it, binding processes to GPUs, CPUs and memory, and making schedulers respect it
How to make GPU jobs topology-aware: reading nvidia-smi topo -m and NVML, how NCCL detects and dumps topology and the variables that steer it, binding…
Read article →torch.compile, in depth: Dynamo graph capture, guards and graph breaks, AOTAutograd, Inductor code generation, modes, dynamic shapes and compile-time caching
A first-principles guide to torch.compile: how TorchDynamo turns Python bytecode into FX graphs with guards, why graph breaks and recompilations happe…
Read article →Google TPU
TPU as an architectural contrast to the GPU: what a systolic array does well and badly, XLA ahead-of-time compilation against static shapes versus CUD…
Read article →Google TPU v5, in depth
Running training on Google TPU v5e and v5p in depth: accelerator naming and provisioning, one process per host, compile-once train steps, input pipeli…
Read article →Google TPU v6 Trillium, in depth: the 256x256 MXU, shapes that fill it, collectives on a 16x16 torus, and measuring a JAX step
Working on Google TPU v6e (Trillium): documented chip and pod specifications, how the 256x256 MXU tiles matrix multiplications, a JAX benchmark and pr…
Read article →LLM Training Checkpointing, in depth: choosing the interval, pricing failures, exact-resume state, restarting on a different GPU count and tiered retention
Checkpoint policy for large training runs: the Young and Daly interval, a waste model priced with Llama 3's interruption rate, every piece of sta…
Read article →LLM Training Cluster Design, in depth: sizing GPUs, fabric, storage, power and restarts from the 6ND budget
Design an LLM training cluster from the job backwards: the 6ND compute budget and MFU, a worked 70B on 15T tokens example on 4,096 H100s, parallel lay…
Read article →LLM Training Data Pipeline, in depth: deterministic token indexes, mixture blending, rank partitioning, exact resumption and replayable batches
How the online data pipeline for LLM pretraining works: memory-mapped token shards, packed sample indexes, deterministic mixture blending, data-parall…
Read article →LLM Training Debugging, in depth: expected curves, a symptom triage table, the shrink-the-problem ladder, data and NaN localisation, gradient flow and distributed-only bugs
How to diagnose LLM training bugs: what a healthy curve predicts (ln V start loss), a symptom-to-suspect table, overfitting one batch, decoding real b…
Read article →LLM Training Experiment Tracking, in depth: metrics, token axes, run lineage and spike alerts
How to track LLM training runs: a metric taxonomy with cadences, tokens-seen as the x-axis, rank-0 async logging without host syncs, run lineage acros…
Read article →LLM Training Health Checks, in depth: preflight gates, in-loop numerics, loss-spike rollback, NCCL hang detection and rank-agreed responses
How to keep a distributed LLM training run healthy: DCGM and nccl-tests preflight gates, a matmul canary, global nonfinite checks, a robust loss-spike…
Read article →Hyperparameter Tuning in Practice: Grid, Random, Bayesian, ASHA, and PBT, with Pseudocode
A practical guide to hyperparameter search on GPUs: what is worth tuning for neural network and LLM training, why random search beats grid, Bayesian o…
Read article →LLM Training Orchestration, in depth: launch and rendezvous, preflight checks, hang detection, a restart supervisor with spares, and goodput
How to orchestrate a long LLM training run: Slurm plus torchrun launch, node preflight checks, detecting crashes, hangs and stragglers, a restart supe…
Read article →LLM Training Reproducibility, in depth: run manifests, exact resume, data order, parallel layouts, seed variance and bisecting divergent runs
Run-level reproducibility for multi-GPU LLM training: levels of guarantee, a run manifest, checkpoints complete enough for exact resume, sample order …
Read article →Anatomy of One GPU Training Step: Forward, Backward, Optimizer Update, and Where the Memory Goes, with Pseudocode
A timeline of one training step on a GPU: the data copy, the forward pass that saves activations, the backward pass that frees them and produces gradi…
Read article →Triton Inference Server for LLMs, in depth: decoupled streaming, vLLM and TensorRT-LLM backends, KV cache sizing and multi-GPU modes
How to serve large language models on NVIDIA Triton Inference Server: why LLM backends own batching, decoupled streaming and the generate endpoints, a…
Read article →Triton, in depth: the block programming model, the compiler pipeline, autotuning and debugging GPU kernels in Python
A first-principles guide to writing GPU kernels in Triton: programs and blocks instead of threads, masks and constexprs, a worked fused softmax kernel…
Read article →TRL Deep Dive, in depth: trainers, data formats, memory and scaling
A practical deep dive into Hugging Face TRL: the config and trainer architecture, the v1.0 stable and experimental tiers, dataset formats, GPU memory …
Read article →TTFT, in depth: a full time-to-first-token ledger with prefill FLOPs, the attention term, prefix caching, tensor parallelism and queueing
Time to first token from first principles: non-embedding prefill FLOPs, the causal attention term, the HBM floor, prefix-cache arithmetic, tensor-para…
Read article →TTS (Text-to-Speech) Serving, in depth: time to first audio, real-time factor, chunked decoding, batching and barge-in
Running text-to-speech as a GPU service: TTFA and per-stream real-time factor as SLOs, streaming text segmentation, codec-token generation with contin…
Read article →UCX, in depth: Unified Communication X layers, GPU data paths and debugging transport selection
How UCX moves data for MPI, UCC, NIXL and RAPIDS: UCP, UCT and UCS layers, workers, endpoints and progress, eager versus rendezvous, GPU memory paths …
Read article →Unsloth, in depth: where fine-tuning memory goes, Triton kernels, offloaded checkpointing, an end-to-end QLoRA run and export paths
How Unsloth speeds up LLM fine-tuning: a worked memory budget for QLoRA on an 8B model, the Triton kernels, hand-derived LoRA backward and offloaded g…
Read article →Vertex AI Gemini, in depth: capacity lanes, dynamic shared quota, Provisioned Throughput sizing and a per-request router
Gemini on Google Cloud (now the Gemini Enterprise Agent Platform) as a capacity system: Standard, Priority and Flex PayGo, batch inference and Provisi…
Read article →NVIDIA vGPU
NVIDIA vGPU explained as a hypervisor-mediated sharing model: mediated passthrough and what the guest driver actually talks to, time-sliced engine sha…
Read article →Video Diffusion Models, in depth: 3D VAEs, spacetime tokens, diffusion transformers and what quadratic attention over video does to GPU training and inference
How modern video diffusion models work and what they cost on GPUs: causal 3D VAE compression, patchified spacetime tokens, diffusion transformers with…
Read article →Video Generation Serving, in depth: async jobs, shape buckets, sequence-parallel GPU groups and split decode stages
How to serve text-to-video diffusion models: an asynchronous idempotent job API, a token-count cost model by shape bucket, scheduling onto sequence-pa…
Read article →Video LLM, in depth: frames, token budgets, encoder cost and KV memory on the GPU
How video language models turn a clip into tokens and what that costs on a GPU: decode, frame sampling, 3D patches and token merging, time-aware posit…
Read article →Vision Encoders, in depth: patch math, token budgets, FLOPs by resolution, native-resolution packing and serving ViT, CLIP, SigLIP and DINOv2 on the GPU
Vision encoders as a GPU workload: patchify as a matmul, token count and FLOP formulas, ViT-L/14 worked from 224 to 672 pixels, tiling versus native r…
Read article →Vision Transformer Serving, in depth: the quadratic crossover, fused attention, resolution buckets with CUDA graphs, token merging and quantization
How to serve a standalone ViT for classification, embeddings and dense prediction: a per-resolution cost model, where attention starts to dominate, SD…
Read article →vLLM on GPU, in Depth: How the Engine Spends GPU Memory, Schedules Every Step, and What to Tune When It Misbehaves
Run vLLM on GPUs with intent: the process layout, how the startup memory profile becomes a KV-cache token budget (worked for Llama-3.1-8B and 70B), ho…
Read article →Vision-Language Model Architectures, in depth: projectors, learned queries, cross-attention, early fusion and what each costs on a GPU
How vision-language models connect an image encoder to a language model: LLaVA projectors, Q-Former and Perceiver queries, Flamingo and Llama 3.2 cros…
Read article →GPU vs CPU for Deep Learning: Latency Machines, Throughput Machines, and Why Training Runs on GPUs
Why deep-learning training runs on GPUs, worked from the workload: training-step FLOP counts, one roofline for a CPU and an H100, a time-to-train esti…
Read article →GPU warp specialization architecture
Deep-dive on GPU warp specialization: splitting a thread block's warps into producer warps that drive cp.async/TMA copies and consumer warps that…
Read article →WebAssembly for LLM Inference, in depth: SIMD kernels, threads, memory limits and when to move to WebGPU
How LLM inference runs in WebAssembly: why decode is memory-bandwidth bound, 128-bit SIMD int8 dot-product kernels, threads with SharedArrayBuffer and…
Read article →WebGPU for LLM Inference, in depth: buffers and limits, f16 and subgroups, a quantised matvec kernel and the decode budget
How WebGPU runs language models in the browser: adapters, devices, buffers and pipelines, default limits that force weight sharding, shader-f16 and su…
Read article →GPU Workload Types, in depth: classifying training, inference, rendering and HPC jobs by the resource they saturate
A practical taxonomy of GPU workloads by bottleneck: arithmetic intensity and the ridge point computed in code, why training and prefill are tensor-bo…
Read article →ZeRO Optimizer, in depth: what each stage adds to a training step, a 7B memory budget, and running it with DeepSpeed
How to operate the ZeRO optimizer in practice: the collectives stages 1, 2 and 3 add to every step, a worked memory budget for a 7B model on eight GPU…
Read article →ZeRO sharding architecture
Deep-dive on ZeRO: the 16-bytes-per-parameter memory math, stages 1/2/3, reduce-scatter and all-gather volume, prefetch overlap, hybrid sharding, ZeRO…
Read article →Gradient Clipping, in depth: norm versus value, the AMP and accumulation order, global norms under sharding, Adam's moments and threshold choice
How gradient norm clipping really works: the clamped coefficient, where it sits relative to loss scaling and accumulation, computing the global norm u…
Read article →Gradient Checkpointing (Activation Recomputation), in depth: activation memory arithmetic, full versus selective recompute, PyTorch and Megatron switches, and reading MFU afterwards
Activation recomputation from first principles: why activations dominate at long context, the sbh(34 + 5as/h) memory formula, a 7B worked example, sqr…
Read article →Graphcore IPU, in depth: 1,472 tiles, bulk synchronous execution, Poplar and PopTorch, and the memory arithmetic that decides what fits
How the Graphcore GC200 IPU works and how software uses it: tiles with private SRAM, BSP compute-sync-exchange, a Poplar vertex example, PopTorch pipe…
Read article →Grid Connection Challenges for AI, in depth: how a training site gets power, why connections take years, what utilities now require of large loads, and the software controls that make a cluster connectable
How AI data centers connect to the grid: load versus generator interconnection, the large-load study process, what training workloads do to the grid, …
Read article →Groq LPU, in depth: SRAM-resident weights, the tensor streaming processor, compiler-scheduled determinism and sizing a 70B deployment
How Groq's LPU works and when to use it: why decode is bandwidth-bound, the tensor streaming processor layout, static compiler scheduling, multi-…
Read article →H100 Availability History (2023-2025), in depth: from scarcity to commodity, quota versus capacity, and a procurement model for the next GPU generation
How H100 availability changed from 2023 to 2025 and what it means for engineers: a dated timeline, the three phases, why quota is not capacity, Capaci…
Read article →H100-Hours per Foundation Model, in depth: reading published disclosures, back-solving utilisation, goodput and estimating your own run
What published GPU-hour figures for Llama 3.1, Llama 2 and DeepSeek-V3 actually count, how to back-solve delivered throughput and MFU from them, why g…
Read article →H100 System Network Breakdown, in depth: the hop-by-hop bandwidth budget, which parallelism rides which link, and a ladder for finding the slow one
The network of an H100 system broken into hops: HBM, NVLink and NVSwitch, PCIe Gen5, ConnectX-7 rails, leaf, spine and storage. Per-direction bandwidt…
Read article →H100 Pricing History, in depth: reading tiered price data, normalising quotes, cost per useful hour and per token, and pricing a contract
H100 rental prices as data: a tiered index from 2023 to 2025, why hyperscaler, neocloud and marketplace prices diverge, unit traps, a script that conv…
Read article →H200 vs H100 Economics, in depth: cost per token, the break-even price ratio and when extra HBM pays
When an H200 is worth its premium over an H100: the break-even rule in cost per token, why freed KV-cache capacity can raise throughput faster than ba…
Read article →Immersion Cooling, in depth: choosing a fluid, qualifying servers and optics, floor loading, fluid health monitoring and servicing submerged GPU servers
Running an immersion cooling programme for GPU servers: fluid families and PFAS exposure, material compatibility soak tests, optics and warranty, floo…
Read article →Inference Cost per Query, in depth: pricing the whole request graph, the cost distribution and cost per successful answer
How to price one user query end to end: token, GPU-time and fixed-fee spans, a runnable cost model, worked retrieval and agent examples, why caching c…
Read article →InfiniBand NDR (400G) + XDR (800G), in depth: lanes, radix, fabric size, host PCIe limits and collective-time math
What changes between InfiniBand NDR and XDR and why training jobs care: per-lane rates and port split, Quantum-2 versus Quantum-X800 radix, fat-tree s…
Read article →Intel Gaudi 2 + Gaudi 3, in depth: the generations compared, the in-box Ethernet mesh, scale-out arithmetic and porting between them
Gaudi 2 against Gaudi 3 for the people who run software on them: published figures, the all-to-all RoCE mesh and why it favours tensor parallelism of …
Read article →Kubernetes GPU Operator, in depth: operands and their order, ClusterPolicy, per-node control, sharing hooks and driver upgrades without lost jobs
How the NVIDIA GPU Operator turns a Kubernetes node into a GPU node: driver, toolkit, validator, device plugin, feature discovery and DCGM exporter in…
Read article →KV Cache Reuse and Prefix Caching, in depth: hash chains versus radix trees, prompt layouts that earn hits, cache-aware routing across a fleet, and isolation
How KV cache reuse works beyond a single engine: hash-chain and radix-tree indexes, prompt layouts that keep prefixes stable, why round robin destroys…
Read article →Lambda Labs, in depth: instances, the Cloud API lifecycle, region-locked filesystems, 1-Click Clusters and the billing traps
How to operate on the Lambda GPU cloud: on-demand instances, 1-Click Clusters and private cloud, a controller that finds capacity and launches within …
Read article →Latitude.sh
Bare metal GPU cloud with reserved and on-demand capacity. Predictable pricing, physical servers, minimal managed services. Allocation shape, contract…
Read article →Meta AI Research SuperCluster (RSC), in depth: DGX A100 nodes on a non-blocking InfiniBand Clos, the AIRStore data path, isolation and reliability at scale
Meta's AI Research SuperCluster as a design case study: 6,080 A100 GPUs at launch on a non-blocking two-level InfiniBand Clos, tiered flash and A…
Read article →Meta MTIA, in depth: the processing-element grid and memory hierarchy, why recommendation models shaped it, the PyTorch-to-Triton path, and what MTIA 300 and 400 change
Meta's MTIA accelerator from a software engineer's view: generation names, MTIA 200's PE grid, SRAM and LPDDR5, roofline arithmetic for…
Read article →Azure AI Supercomputer for OpenAI, in depth: five generations from 2020 to Fairwater, NVLink domains, fat trees, checkpoints, failure rates and multi-site training
The Microsoft systems built to train OpenAI's models, from the 2020 10,000-GPU cluster through Eagle to Fairwater's GB200 NVL72 halls and AI…
Read article →MIG, in depth: what a CUDA process sees inside a slice, sizing weights and KV cache to an instance, MPS on MIG, training limits and measuring isolation
MIG from the application's side: GPU and compute instances, CUDA enumeration before and after R570, IPC and P2P rules, sizing an 8B model and its…
Read article →MIG for Multi-Tenant GPU Sharing, in depth: tenant classes, layout planning, quotas, re-layout drains and chargeback
Running MIG as a shared service: mapping tenants to H100 profiles, GPU Operator single and mixed strategies, mig-parted layouts, ResourceQuota per pro…
Read article →Modal Labs, in depth: serverless GPUs in Python, cold-start anatomy and memory snapshots, autoscaling knobs, batch fan-out, preemption-safe training and multi-node clusters
How to build on Modal: apps, images and GPU functions, GPU type strings and fallbacks, what a cold start costs and how min_containers, scaledown_windo…
Read article →Model Lifetime Utility, in depth: amortising training GPU-hours over a serving life, demand decay, the replica floor and when to retire a model
A GPU-hour ledger for the life of a model version: fixed versus recurring hours, ramp, plateau and decay of demand, the replica floor, a runnable simu…
Read article →Modular Datacenters for AI, in depth: interface contracts, modules as failure domains, factory and site testing, and wiring modules into the scheduler
How prefabricated power, cooling and IT modules work for AI datacenters: interface contracts with transients, modules as failure and maintenance domai…
Read article →Small Modular Reactors (SMR) for AI DCs, in depth: the designs, connection models, sizing around refuelling, training load swings and realistic timelines
An engineering view of feeding an AI data center from small modular reactors: what counts as an SMR, the designs with real licensing progress, behind-…
Read article →MoE Model Inference Cost, in depth: capacity, bandwidth and compute bills, experts touched per batch, and a dollars-per-million-tokens model
A first-principles cost model for serving mixture-of-experts LLMs: why capacity scales with total parameters, decode bandwidth with experts touched an…
Read article →Multi-GPU Peer-to-Peer Access, in depth: the server as a graph, NCCL transport choice, one-shot all-reduce over peer pointers and placing groups on P2P cliques
Peer access across four to eight GPUs: modelling a server as a connectivity graph and finding cliques, how NCCL chooses P2P, SHM or network transport …
Read article →Multi-Year GPU Capacity Commitments, in depth: a locked rate against falling prices, sizing under demand risk, contract clauses and the fleet lifecycle
How to evaluate a two-to-five-year GPU commitment: a present-value model against a falling market price for the same work, break-even decline, term co…
Read article →All-Reduce, in depth: the ring and tree algorithms, NCCL protocols, busbw versus algbw, the APIs, a worked cost estimate and how training jobs fail
A first-principles guide to all-reduce for GPU training: the alpha-beta cost model, ring reduce-scatter plus all-gather and why it is bandwidth-optima…
Read article →NCCL, in depth: communicators, topology search, transports, buffer registration, fault recovery and a debugging runbook
NCCL as a runtime you operate: communicator init and split, topology detection and transports, channels and the proxy thread, buffer registration and …
Read article →NCCL Topology-Aware Communication, in depth: the path-type ladder, topology files for VMs, NIC selection and diagnosing what NCCL detected
How NCCL turns a node's topology into decisions: the NVL, PIX, PXB, PHB and SYS path types and the variables that gate on them, NVB and PXN, topo…
Read article →Nebius AI Cloud, in depth: regions and InfiniBand fabrics, building a GPU cluster, P-Key isolation, Soperator's shared root and acceptance testing
A practitioner's runbook for Nebius AI Cloud: the region, fabric, GPU cluster and VM model, the current fabric and platform table, CLI cluster cr…
Read article →NCCL + NIXL, in depth: lock-step collectives versus one-sided transfers, the NIXL agent lifecycle, and moving KV cache between instances
When to use NCCL and when to use NIXL: communicators versus agents, memory sections, backend plug-ins and metadata, the NIXL Python lifecycle, prepare…
Read article →NVIDIA Nsight Systems, in depth: capture windows for training jobs, nsys stats reports, the SQLite export, scripted idle-gap analysis and recipes
Using nsys as an analysis pipeline for GPU training: NVTX and profiler start-stop capture windows, the stats reports, the SQLite export schema, a merg…
Read article →NVIDIA Rubin, in depth: HBM4 and NVLink 6 through a software lens, the NVL72 domain, NVFP4 training and how to get ready
What NVIDIA Rubin changes for people who train and serve models: verified specifications, roofline and decode arithmetic, the 72-GPU NVLink 6 domain a…
Read article →NVIDIA Spectrum-X, in depth: why ECMP fails AI traffic, per-packet adaptive routing, out-of-order placement, NIC-side congestion control and operating it
How NVIDIA Spectrum-X changes RoCE fabrics for AI training: an ECMP collision worked example, per-packet adaptive routing, SuperNIC reordering and dir…
Read article →GB200 NVL72 Rack, in depth: the rack as a physical and failure unit, link arithmetic, power and cooling, acceptance testing, spare trays and checkpoint cadence
The GB200 NVL72 rack from an operator's point of view: 18 compute and 9 NVLink switch trays, why 1,296 links make 18 planes, facility power and l…
Read article →GPU Occupancy Calculator, in depth: the exact arithmetic, a CI-ready calculator, the CUDA occupancy API and reading the result
How a GPU occupancy calculator works: kernel and device inputs from ptxas and the compute capability table, register and shared memory allocation gran…
Read article →OCI GPU SuperCluster, in depth: bare-metal GPU shapes, the RoCEv2 cluster network, compute clusters, bandwidth arithmetic and bring-up
How Oracle's GPU SuperCluster works for training jobs: bare-metal H100, H200, B200, B300 and GB200 shapes, per-GPU RDMA bandwidth, the RoCEv2 clu…
Read article →Optical Transceivers for AI Networking, in depth: the signal path, reach classes, linear and co-packaged optics, FEC budgets and fleet telemetry
How optical transceivers work in GPU training fabrics: inside a pluggable module, VR, SR, DR, FR and LR reach classes, DSP versus linear and co-packag…
Read article →Optimizer State Offloading to CPU/NVMe
Optimizer state offloading as an engineering decision: the per-parameter byte cost of Adam, why optimizer state is the first tier to move off the GPU,…
Read article →PCIe Gen4 + Gen5 for GPUs, in depth: what the generation changes, why links fall back, how to prove your link speed, and which workloads feel the doubling
PCIe Gen4 and Gen5 for GPU systems: per-direction bandwidth derived from the line rate, which GPUs use which generation, signal integrity and retimers…
Read article →Power Backup for AI Datacenters, in depth: the ride-through timeline, why GPU loads stress generators, a checkpoint budget and wiring UPS events into the scheduler
How backup power protects AI training: PSU holdup, rack BBUs, UPS, transfer switches and generators in time order, why synchronized GPU loads stress g…
Read article →Power Density in AI Datacenters, in depth: from GPU watts to rack kilowatts, the airflow and coolant physics, why AI racks are built dense, and density-aware scheduling
Why AI racks run from 40 kW to over 120 kW: how per-GPU power compounds, the airflow and coolant flow arithmetic, electrical and floor-loading limits,…
Read article →PUE for AI Datacenters, in depth: a loss model for liquid-cooled GPU halls, facility water temperature, heat reuse and the levers that move the number
A design-side guide to PUE for AI datacenters: building an hour-by-hour loss model of a liquid-cooled GPU hall, electrical conversion losses, liquid c…
Read article →Qualcomm Cloud AI 100, in depth: AI cores, SRAM versus LPDDR, compiling to a QPC and sizing LLM inference by bandwidth
The Qualcomm Cloud AI 100 as a software target: SKU specifications, the tensor, vector and scalar units of an AI core, the scratchpad memory model, th…
Read article →Rail-Aligned Topology for AI Clusters, in depth: why collectives follow rails, how NCCL and PXN use them, and how placement keeps them
Rail-aligned (rail-optimized) GPU fabrics from the software side: the NIC-i-to-rail-i rule, hierarchical all-reduce on rails, NCCL_CROSS_NIC and PXN f…
Read article →Ray on GPU and Distributed Training, in depth: FSDP inside Ray Train, NCCL setup, data shards, sharded checkpoints and elastic recovery
How Ray Train runs multi-node GPU training: controller and worker group, ScalingConfig and TorchConfig, FSDP with Distributed Checkpoint, Ray Data ing…
Read article →Rear-Door Heat Exchangers, in depth: passive versus active doors, sizing the energy balance, dew point, telemetry and GPU throttling
Rear-door heat exchangers for GPU racks from first principles: passive and active doors, air and water energy balance, effectiveness, a worked 40 kW s…
Read article →GPU Register Pressure
How the CUDA compiler allocates registers and what drives per-thread demand up: live ranges, loop unrolling, inlining and large local arrays. Why spil…
Read article →Reserved GPU Capacity vs On-Demand, in depth: the run-level decision, gang acquisition, sizing a reservation window from a duration distribution and placing each workload
Reserved versus on-demand GPUs for a finite training run: why on-demand fails for gangs of nodes, a Monte Carlo duration model, choosing a reservation…
Read article →Ring All-Reduce, in depth: the chunk schedule step by step, a working implementation, pipelining, multi-ring and hierarchical variants
How the ring all-reduce algorithm works: reduce-scatter and all-gather schedules, a traced four-GPU example, a verified simulator and a torch.distribu…
Read article →SambaNova RDU
SambaNova RDU reconfigurable dataflow architecture for inference: how spatial mapping of computation graphs onto distributed compute and memory units …
Read article →Second-Hand GPU Market
Enterprise churn feeds resale.
Read article →SHARP
How SHARP moves reduction arithmetic into the switch ASIC: aggregation-tree construction and the finite switch state that backs it, which collectives …
Read article →SIMT Execution: How Thousands of GPU Threads Compute in Parallel, with Worked Examples
How the SIMT model maps scalar-looking CUDA threads onto 32-wide warps: active masks, branch divergence and its cost in worked numbers, predication, i…
Read article →DGX SuperPOD Topology, in depth: scalable units, four fabrics, the switch arithmetic, the UFM node and how Slurm and NCCL use the layout
The DGX SuperPOD as a system: the 32-node scalable unit and why pods have 127 nodes, the compute, storage and two management fabrics, non-blocking lea…
Read article →TensorRT, in depth: the builder, tactic selection, strong typing and ModelOpt precision in TensorRT 11, optimization profiles, the runtime API and engine portability
How TensorRT compiles and runs inference: graph fusion, tactic timing and the timing cache, strongly typed precision with ModelOpt AutoCast and Q/DQ q…
Read article →Foundation Model Project Cost Breakdown, in depth: the full ledger beyond the final run, compute from 6ND and MFU, a runnable cost model, a worked 30B example, sensitivity and tracking actuals
What a foundation-model project really costs: reading published figures for scope, a ten-line ledger covering research, failures, post-training, evalu…
Read article →Training Cost Trajectory 2024-2027, in depth: what the 2.4x-a-year trend measures, a cost identity, calibration and a forecast you can run
Frontier training cost from 2024 to 2027: final-run versus programme cost, Epoch AI's 2.4x-per-year estimate and its uncertainty, a FLOPs times p…
Read article →Training on Thousands of GPUs, in depth: composing TP, PP and DP, mapping the mesh to the network, and a full step budget for 70B on 1,024 GPUs
How to lay out an LLM training job across thousands of GPUs: the parallelism axes, building a device mesh, mapping groups to NVLink and the scale-out …
Read article →Tree All-Reduce, in depth: NCCL's double binary tree built bit by bit, a simulator, the pipelined cost model and when rings win
Tree all-reduce from the inside: NCCL's rank bit arithmetic for the binary tree, the mirror and shift rules for the second tree, a Python port an…
Read article →OpenAI Triton Language, in depth: blocks, constexpr and specialization, masks, tl.dot, and a fused matmul written line by line
Triton as a language: program instances and blocks, shape and dtype rules, constexpr and integer specialization, masked loads, tl.dot and pipelined lo…
Read article →Ultra Ethernet Consortium, in depth: the UET transport layer by layer, packet spraying, trimming, NSCC and RCCC, LLR and CBFC, profiles, and what training traffic gains
How the Ultra Ethernet specification works and what it changes for GPU training collectives: SES, PDS, CMS and TSS sublayers, job-based addressing, co…
Read article →Vast.ai, in depth: the GPU marketplace model, choosing offers, interruptible bidding, and training jobs that survive a kill
How Vast.ai works for ML: offers, hosts and the three billing meters, on-demand versus reserved versus interruptible bidding, the search query languag…
Read article →Voltage Park, in depth: a foundation-owned H100 cloud, its bare-metal on-demand and reserved offers, the Lightning AI merger, and how to validate and run training on it
Voltage Park explained with dated facts: Navigation Fund ownership, about 24,000 H100s at launch, HGX nodes on Quantum-2 InfiniBand and VAST storage, …
Read article →Warp Scheduling on GPU
How a GPU SM warp scheduler picks which warp issues each cycle: eligible versus merely resident warps, the scoreboard, and the stall-reason taxonomy (…
Read article →Waste Heat Reuse from Datacenters, in depth: heat grade, heat pumps, ERF and why training load shapes heat supply
Reusing GPU datacenter heat: temperature grade and liquid cooling, the heat path to a district network, heat-pump COP in kelvin, the Energy Reuse Fact…
Read article →Water Usage in AI Datacenters, in depth: withdrawal versus consumption, evaporation physics, cycles of concentration, source water and a per-job water model
How AI datacenters use water, built from first principles: withdrawal, consumption and discharge, about 1.5 litres evaporated per kWh of heat, blowdow…
Read article →xAI Colossus, in depth: the 64-GPU rack and 512-GPU array, Spectrum-X Ethernet for collectives, power transients and failure math at 100,000 GPUs
xAI's Memphis training cluster read as engineering: the reported server, rack and array building blocks, rail-optimised Spectrum-X Ethernet and w…
Read article →XLA, in depth: tracing to StableHLO, HLO passes, fusion, layout and buffer assignment, SPMD partitioning and avoiding recompilation
How the XLA compiler turns a JAX function into device code: jaxpr, StableHLO, HLO fusion, layout and buffer assignment, Shardy partitioning, GPU code …
Read article →Zero-Copy Memory (Pinned Host), in depth: what a kernel load over PCIe costs, the break-even against copying, sparse gathers and host-visible results
How CUDA zero-copy memory works: mapped pinned host memory with cudaHostAllocMapped and cudaHostRegisterMapped, the cost of a kernel load over PCIe, c…
Read article →DeepSpeed ZeRO Stages, in depth: choosing stage 0 to 3 by memory and traffic, gradient accumulation, the stage ladder and moving checkpoints between stages
How to choose a DeepSpeed ZeRO stage: per-stage sharding and traffic per optimizer step, why gradient accumulation makes stage 2 and 3 move more bytes…
Read article →