Serving an image classifier looks like the easiest job in production machine learning: one image in, one label out, a fixed-size network, no sequence lengths, no KV cache. In practice the model is often the cheapest part of the request. Teams buy a GPU, deploy a ResNet or a ViT behind a Python web server, and find the GPU at 15 percent utilization while CPU cores decoding JPEGs sit at 100 percent, or find the accuracy in production a few points below the offline number because the server resizes images differently from the training loader.
This article builds a serving path from first principles: what each stage costs, how to move decode and preprocessing onto the GPU, how to build a TensorRT engine, how to configure Triton Inference Server with dynamic batching and an ensemble, how to size a fleet against a latency target, and how to catch training-serving skew before it costs accuracy. Numbers in worked examples are illustrative; the method for measuring your own is the point.
Where a classification request spends its time
A classification request passes through six stages. The client sends compressed bytes, usually JPEG. The server decodes them to an RGB array, resizes (commonly the short side to 256 pixels), center-crops to the network input size (224 by 224 for many ImageNet models), converts to float, subtracts the per-channel mean and divides by the standard deviation, and reorders the array to the layout the network expects, usually NCHW. Then the network runs, and a top-k over the output vector yields the class indices that are mapped to labels.
The split of cost surprises people. A ResNet-50 forward pass is roughly four billion multiply-accumulates per image, which a modern data-center GPU running FP16 tensor cores finishes in well under a millisecond when batched. Decoding a typical 500 KB photo on one CPU core takes several milliseconds. So a server that decodes on the CPU needs many cores per GPU just to keep the GPU fed. Time decode, preprocess and inference separately under load before optimizing, because the slowest stage sets throughput.
Decode and preprocess on the GPU
Moving decode onto the GPU removes the CPU bottleneck. NVIDIA nvJPEG decodes JPEG on the GPU, and several data-center GPUs since the A100 also include dedicated hardware JPEG decode engines that nvJPEG can use, which keeps decode off the streaming multiprocessors entirely. Check the decode capability of your exact GPU before planning around it. NVIDIA DALI wraps decode, resize, crop and normalize into a pipeline, and Triton has a DALI backend, so the preprocessing can be a model in the same server as the network.
The DALI pipeline below is the GPU twin of the classic ImageNet evaluation transform. The device="mixed" decoder takes CPU bytes and produces GPU pixels; everything after that stays on the device.
import nvidia.dali as dali
import nvidia.dali.fn as fn
import nvidia.dali.types as types
@dali.pipeline_def(batch_size=64, num_threads=4, device_id=0)
def preprocess():
raw = fn.external_source(device="cpu", name="RAW_IMAGE")
img = fn.decoders.image(raw, device="mixed", output_type=types.RGB)
img = fn.resize(img, resize_shorter=256)
img = fn.crop_mirror_normalize(
img, crop=(224, 224), dtype=types.FLOAT, output_layout="CHW",
mean=[0.485 * 255, 0.456 * 255, 0.406 * 255],
std=[0.229 * 255, 0.224 * 255, 0.225 * 255])
return img
preprocess().serialize(filename="model_repository/preprocess/1/model.dali")The mean and standard deviation are scaled by 255 because the decoder produces 0 to 255 values while torchvision-style training normalizes 0 to 1 tensors. Getting that factor wrong does not crash anything; it just makes every prediction worse.
Building the engine
The network itself should run as an optimized engine, not as eager framework code. The usual route is to export the trained model to ONNX with a dynamic batch dimension, then build a TensorRT engine that fuses layers, chooses kernels for the target GPU and runs in reduced precision.
import torch, torchvision
model = torchvision.models.resnet50(weights="IMAGENET1K_V2").eval()
dummy = torch.randn(1, 3, 224, 224)
torch.onnx.export(model, dummy, "resnet50.onnx", opset_version=17,
input_names=["input"], output_names=["logits"],
dynamic_axes={"input": {0: "batch"}, "logits": {0: "batch"}})trtexec --onnx=resnet50.onnx --saveEngine=model.plan --fp16 \
--minShapes=input:1x3x224x224 \
--optShapes=input:32x3x224x224 \
--maxShapes=input:64x3x224x224The optimization profile tells TensorRT which batch sizes to expect; the opt shape is where kernel selection is tuned, so set it to the batch size dynamic batching will actually form. FP16 is usually close to lossless for classifiers but must still be validated on a held-out set. INT8 can roughly double throughput again, but only with proper quantization: either calibration data or a quantization-aware trained model with explicit quantize and dequantize nodes. Running trtexec --int8 without either gives you a speed number and nothing else; never ship an INT8 engine whose top-1 accuracy you have not measured.
An engine is tied to the GPU architecture and the TensorRT version that built it. Build engines in CI per target GPU type, store them as artifacts keyed by both, and fail the deployment if a node's GPU or TensorRT version does not match the artifact.
Wiring it up in Triton
Triton Inference Server loads models from a repository directory, one folder per model with a config.pbtxt and numbered version subfolders. Three models make the pipeline: the DALI preprocessor, the TensorRT classifier and an ensemble that connects them.
# model_repository/classifier_trt/config.pbtxt
name: "classifier_trt"
platform: "tensorrt_plan"
max_batch_size: 64
input [ { name: "input" data_type: TYPE_FP32 dims: [ 3, 224, 224 ] } ]
output [ { name: "logits" data_type: TYPE_FP32 dims: [ 1000 ] } ]
dynamic_batching {
preferred_batch_size: [ 16, 32 ]
max_queue_delay_microseconds: 2000
}
instance_group [ { count: 2 kind: KIND_GPU } ]
# model_repository/classifier/config.pbtxt (the ensemble clients call)
name: "classifier"
platform: "ensemble"
max_batch_size: 64
input [ { name: "RAW_IMAGE" data_type: TYPE_UINT8 dims: [ -1 ] } ]
output [ { name: "LOGITS" data_type: TYPE_FP32 dims: [ 1000 ]
label_filename: "labels.txt" } ]
ensemble_scheduling {
step [
{ model_name: "preprocess" model_version: -1
input_map { key: "RAW_IMAGE" value: "RAW_IMAGE" }
output_map { key: "DALI_OUTPUT_0" value: "pixels" } },
{ model_name: "classifier_trt" model_version: -1
input_map { key: "input" value: "pixels" }
output_map { key: "logits" value: "LOGITS" } }
]
}The preprocess model uses backend: "dali" with the serialized pipeline as its model file, and its output name must match the one in the ensemble map. Encoded images differ in length, so its input needs allow_ragged_batch: true to batch at all; confirm the setting in the DALI backend documentation for your release. Two instances of the engine on one GPU let one batch copy data while another computes; more instances rarely help and cost memory. The ensemble keeps the pixels on the GPU between steps, so the only PCIe traffic is compressed bytes in and a short result out.
Clients can ask Triton for top-k directly through the classification extension. Each returned element is a string of the form value:index:label. The value is whatever the output holds, which here is a logit, not a probability; top-k order is the same either way, but if callers need calibrated confidence, add a softmax to the model graph.
import numpy as np
import tritonclient.http as httpclient
client = httpclient.InferenceServerClient(url="localhost:8000")
data = np.frombuffer(open("cat.jpg", "rb").read(), dtype=np.uint8)[None, :]
inp = httpclient.InferInput("RAW_IMAGE", list(data.shape), "UINT8")
inp.set_data_from_numpy(data)
out = httpclient.InferRequestedOutput("LOGITS", class_count=5)
result = client.infer("classifier", [inp], outputs=[out])
print(result.as_numpy("LOGITS")) # e.g. [[b'14.2:281:tabby', ...]]
Dynamic batching from first principles
A GPU runs one image and thirty-two images in nearly the same wall time when the batch is small, because a single image cannot fill the tensor cores. Dynamic batching exploits that: the scheduler holds arriving requests for at most max_queue_delay_microseconds and dispatches as soon as it has a preferred batch size or the window expires.
The cost model is simple. Per-image latency is queue wait plus preprocessing plus the compute time of the whole batch plus network. Throughput is batch size divided by batch compute time, summed over the instances that can overlap. At low traffic the window adds pure latency because batches never fill; at high traffic it adds almost none. That is why the delay should be small, one to a few milliseconds, and the preferred sizes should match the engine's opt profile. For a deeper treatment of these policies see GPU batching strategies.
Worked example: sizing for a latency target
Suppose the target is 4,000 images per second at p99 under 80 ms end to end, and requests arrive independently. The procedure, with illustrative numbers:
- Measure one GPU with
perf_analyzer -m classifier --concurrency-range 8:128:8 --percentile=99 --input-data=samples.jsonusing real JPEGs of production size, not random tensors. Record throughput and p99 at each concurrency. - Read the knee: say throughput plateaus at 2,400 images per second with p99 of 35 ms at concurrency 64, while concurrency 96 adds 2 percent throughput and doubles p99.
- Plan to run at about 65 percent of the plateau so bursts and a node loss do not push p99 past the budget: 2,400 times 0.65 is about 1,560 images per second per GPU.
- Divide: 4,000 / 1,560 is 2.6, so three GPUs carry the load and a fourth gives N+1 redundancy.
- Confirm by replaying a recorded traffic trace against the four-GPU deployment, because synthetic constant concurrency hides burst behavior.
Little's law ties this together: requests in flight equal arrival rate times time in system. At 4,000 per second and 30 ms average, about 120 requests are in flight across the fleet, which tells you how much client-side concurrency and connection pooling you need. Latency accounting in more detail is in GPU inference latency.
Training-serving skew
The most expensive failure in image serving is silent: the server preprocesses differently from the training loader and accuracy drops a few points with no error anywhere. The usual culprits are a different resize library or interpolation (bilinear versus bicubic, with or without antialiasing), RGB versus BGR channel order, normalization in the wrong value range, NHWC versus NCHW, and EXIF orientation applied in one pipeline but not the other, which presents phone photos rotated by 90 degrees.
Guard against it with a parity test in CI: run a fixed set of a few hundred images through the training-time transform and through the deployed server, and compare both the predicted labels and the logits.
def parity(images, train_transform, model_fp32, triton_client, tol=0.02):
mismatched = 0
for path in images:
ref = model_fp32(train_transform(load(path))[None]).argmax(1).item()
got = top1_from_triton(triton_client, path) # via the ensemble
mismatched += (ref != got)
rate = mismatched / len(images)
assert rate <= tol, f"top-1 disagreement {rate:.1%} exceeds {tol:.0%}"Some disagreement is expected from FP16 and GPU resize, so set the tolerance from a measured baseline and alert on change, not on zero. Also track the production distribution of top-1 confidence: a sudden shift is often the first visible symptom of skew or of a new camera.
Operating it
Triton exports Prometheus metrics on port 8002. The ones to chart for a classifier:
nv_inference_queue_duration_usdivided by request count: average time waiting in the batcher; if it rises while GPU utilization is low, the window is too long or preprocessing is starving the engine.nv_inference_compute_infer_duration_us: model compute time, which should stay flat for a given batch size.nv_inference_countdivided bynv_inference_exec_count: the average batch size actually formed, the single best check that dynamic batching is working.- GPU utilization and memory from DCGM, and decode engine utilization if you use hardware decode.
For several small classifiers on one large GPU, consider partitioning it; see Multi-Instance GPU for isolation trade-offs, and Triton for LLMs for the server features this article does not cover.
Failure modes
- CPU-bound decode. GPU utilization low, CPU high, throughput flat as you add GPUs. Move decode to the GPU or add decode workers.
- Batches never form. Average batch size near 1 because clients send sequentially with concurrency 1. Raise client concurrency or batch on the client.
- Engine mismatch. A plan built on one GPU type fails to load, or loads with warnings, on another after an autoscaler adds a different node pool.
- Oversized uploads. A 40-megapixel image decodes to more than 100 MB of pixels. Cap upload size and dimensions at the edge, before decode.
- Corrupt or hostile files. Truncated JPEGs and decompression bombs. Return a 4xx for the one request; never let one bad file fail a whole batch.
- INT8 regression. Throughput up, accuracy down on a minority class that the calibration set barely contained. Evaluate per class, not only overall top-1.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| GPU decode with DALI | Frees CPU, higher throughput | Another pipeline to keep in parity with training |
| FP16 engine | Large speedup, small accuracy risk | Per-GPU builds, validation run |
| INT8 engine | More throughput again | Calibration or QAT, per-class evaluation |
| Longer batch window | Bigger batches under light load | Added latency exactly when traffic is low |
| More instances per GPU | Overlaps copy and compute | Memory, diminishing returns past two |
| Client-side batching | Fewer requests, less overhead | Head-of-line blocking inside the client |
What to do next
- Time decode, preprocess and inference separately under load before changing anything.
- Export to ONNX with a dynamic batch axis and build an FP16 TensorRT engine per GPU type in CI.
- Put preprocessing in a DALI model and connect it with a Triton ensemble.
- Enable dynamic batching with a 1 to 2 ms window and preferred sizes equal to the opt profile.
- Add the training-versus-serving parity test and a per-class accuracy check to CI.
- Size the fleet from a perf_analyzer knee at about 65 percent of plateau, plus one spare.
- Chart queue time, compute time and average formed batch size, and alert on drift.