Serving an object detector looks easy on a slide: an image goes in, a list of boxes comes out. In production the model is only one of six stages, and it is frequently not the slow one. JPEG decoding, resizing, the transfer to the GPU, the decoding of raw prediction tensors and non-maximum suppression (NMS) together can cost more than the convolution layers, and most of them default to running on the CPU in Python. A detector that benchmarks at a few milliseconds per batch can deliver a few dozen frames per second end to end because the GPU spends most of its time waiting.

This article builds a detection service from first principles: what the network actually outputs, why NMS exists, how letterboxing changes coordinates, how dynamic batching trades latency for throughput, and how to keep the whole pipeline on the GPU with TensorRT and Triton Inference Server. It ends with a worked capacity plan, the failure modes that show up in real deployments, and a checklist. The examples use a YOLO-style anchor-free detector exported to ONNX because that is the most common shape, but the structure applies to any single-stage or DETR-style model.

What a detector actually outputs

A single-stage detector predicts on a grid. For a 640 by 640 input, a YOLOv8-style model uses three feature maps with strides 8, 16 and 32, which gives 80 by 80, 40 by 40 and 20 by 20 cells: 6,400 + 1,600 + 400 = 8,400 candidate predictions per image. Each candidate carries four box values and one score per class, so with the 80 COCO classes the raw output tensor is [batch, 84, 8400]. None of it is a detection yet: most candidates score low, and confident ones cluster around each object across cells and scales.

Turning candidates into detections takes three steps. First, a confidence filter discards candidates whose best class score falls below a threshold such as 0.25. Second, the box values are converted from centre, width and height into corner coordinates. Third, NMS walks the survivors in descending score order and removes any box whose intersection over union (IoU) with an already kept box of the same class exceeds a threshold such as 0.45 to 0.7. A final top-k cap bounds the output size so one crowded frame cannot produce a megabyte response.

Two consequences matter for serving. The output size is data-dependent, which is awkward for batched GPU execution and for fixed-shape engines. And NMS is inherently sequential within a class, so a naive Python loop over candidates costs milliseconds per image when a scene is crowded. Some newer designs remove NMS: DETR-family models and YOLOv10 use a one-to-one prediction head at inference, so each object gets a single prediction. They still need a score threshold and a top-k, but the data-dependent suppression step disappears, which simplifies the serving graph.

import torch
from torchvision.ops import batched_nms

def postprocess(raw, conf_thres=0.25, iou_thres=0.6, max_det=300):
    """raw: [B, 4 + C, N] on the GPU -> list of [K, 6] (x1, y1, x2, y2, score, cls)."""
    raw = raw.transpose(1, 2)                      # [B, N, 4 + C]
    boxes_cxcywh, cls_scores = raw[..., :4], raw[..., 4:]
    scores, classes = cls_scores.max(dim=-1)       # best class per candidate
    out = []
    for b in range(raw.shape[0]):                  # loop over images, not candidates
        keep = scores[b] > conf_thres
        if not keep.any():
            out.append(raw.new_zeros((0, 6)))
            continue
        cx, cy, w, h = boxes_cxcywh[b, keep].unbind(-1)
        xyxy = torch.stack((cx - w / 2, cy - h / 2, cx + w / 2, cy + h / 2), -1)
        s, k = scores[b, keep], classes[b, keep]
        idx = batched_nms(xyxy, s, k, iou_thres)[:max_det]   # per-class NMS, GPU kernel
        out.append(torch.cat((xyxy[idx], s[idx, None], k[idx, None].float()), 1))
    return out

Preprocessing and the letterbox inverse

Detectors are trained on square inputs, but cameras produce 16:9 frames. Stretching the frame distorts object shapes the model learned, so the standard approach is letterboxing: scale the image by r = min(640 / width, 640 / height), place it in the centre of a 640 by 640 canvas and fill the remainder with a constant grey. A 1920 by 1080 frame scales by one third to 640 by 360 with 140 pixels of padding above and below.

Every box the model returns is in letterboxed coordinates, so the service must invert the transform: subtract the padding, divide by r and clip to the frame. Getting this wrong is the most common silent bug: boxes look right on square test images and drift by up to 140 pixels on real camera frames. The preprocessing used at serving time must also match training exactly: channel order (RGB or BGR), normalization (divide by 255 or mean and standard deviation), interpolation method and pad value. A mismatch here does not crash anything; it just costs several points of accuracy.

Letterbox: resize with aspect ratio, pad the restsource 1920 x 1080scale r = 1/3640 x 360 imagepad top 140 pxpad bottom 140 pxmodel input 640 x 640Inverse map: x_src = (x - pad_x) / r, y_src = (y - pad_y) / r, then clip to the source frame.
Letterboxing preserves aspect ratio; the padding and scale must be stored per image to map boxes back.
def letterbox_params(w, h, size=640):
    r = min(size / w, size / h)
    new_w, new_h = round(w * r), round(h * r)
    pad_x, pad_y = (size - new_w) / 2, (size - new_h) / 2
    return r, pad_x, pad_y

def unletterbox(det, w, h, size=640):
    """det: [K, 6] in model space -> source pixel space, clipped."""
    r, pad_x, pad_y = letterbox_params(w, h, size)
    det[:, [0, 2]] = ((det[:, [0, 2]] - pad_x) / r).clamp(0, w)
    det[:, [1, 3]] = ((det[:, [1, 3]] - pad_y) / r).clamp(0, h)
    return det

Keeping every stage on the GPU

The architecture that keeps a GPU busy moves every stage onto it. Images arrive as compressed bytes; the GPU decodes them with nvJPEG for stills or NVDEC for H.264 and HEVC video, so the only PCIe transfer is the small compressed payload rather than a 6 MB raw frame. NVIDIA DALI wraps this: its image decoder with device="mixed" parses on the CPU and decodes on the GPU, and its resize and pad operators produce the letterboxed tensor without the data leaving device memory. CUDA kernels or torchvision on CUDA tensors work equally well; what matters is never round-tripping through host memory.

The detector itself should run as a compiled engine rather than eager PyTorch. Exporting to ONNX and building a TensorRT engine fuses convolutions with their activations, selects kernels for the specific GPU and enables FP16, which for convolutional detectors usually costs no measurable accuracy. INT8 roughly halves the work again but needs a calibration set drawn from production-like images, and it hurts small objects first, so validate it on per-size mAP rather than the headline number. Engines are built for a range of shapes through an optimization profile; build with the batch sizes you will actually serve.

NMS can be folded into the engine. TensorRT ships an EfficientNMS_TRT plugin that takes boxes and scores and returns four fixed-shape tensors: the number of detections per image and padded arrays of boxes, scores and classes sized to a maximum such as 100. Fixed shapes are exactly what batching wants, and moving NMS into the engine removes the last Python loop from the hot path. The cost is flexibility: thresholds are baked into the engine at export time, so changing the IoU threshold means rebuilding.

Object detection request path on one GPUClientsJPEG / video framesDecodenvJPEG / NVDECPreprocessletterbox, normalizeDynamic batchermax batch, queue delayDetector engineTensorRT FP16 / INT8Decode headsboxes, scoresNMS + top-kon GPUUn-letterboxto source pixelsKeep every stage on the GPUa CPU hop between stages costs a PCIe copy and a Python loopThe model is often less than half the latency; decode, resize and NMS decide whether the GPU stays busy.
A GPU-resident detection pipeline: only compressed bytes cross PCIe on the way in and only small box arrays on the way out.
# Build an engine with a batch range that matches the batcher's settings.
trtexec --onnx=detector.onnx --saveEngine=detector_fp16.plan --fp16 \
        --minShapes=images:1x3x640x640 \
        --optShapes=images:8x3x640x640 \
        --maxShapes=images:16x3x640x640

Serving with Triton and dynamic batching

Triton Inference Server turns these pieces into a service. Each stage becomes a model in the repository, and an ensemble wires them together so a client sends one request with encoded bytes and receives boxes. A typical layout is a DALI preprocessing model, the TensorRT detector and a small Python or C++ backend for coordinate mapping. Because the ensemble runs inside the server, intermediate tensors never cross the network.

Dynamic batching is the most important setting. The scheduler holds incoming requests for up to max_queue_delay_microseconds to assemble a larger batch, up to max_batch_size. Batching raises throughput because a batch of eight costs far less than eight batches of one, but every request now waits for the window. instance_group controls how many copies of the engine run concurrently on each GPU; two instances let one batch run while the next is being formed and copied.

# model_repository/detector/config.pbtxt
name: "detector"
platform: "tensorrt_plan"
max_batch_size: 16
input  [ { name: "images" data_type: TYPE_FP16 dims: [ 3, 640, 640 ] } ]
output [
  { name: "num_dets"   data_type: TYPE_INT32 dims: [ 1 ] },
  { name: "det_boxes"  data_type: TYPE_FP32  dims: [ 100, 4 ] },
  { name: "det_scores" data_type: TYPE_FP32  dims: [ 100 ] },
  { name: "det_classes" data_type: TYPE_INT32 dims: [ 100 ] }
]
dynamic_batching {
  preferred_batch_size: [ 8, 16 ]
  max_queue_delay_microseconds: 2000
}
instance_group [ { count: 2 kind: KIND_GPU } ]

Tensor names and dtypes must match your exported engine; read them from the ONNX graph. Use Triton's perf_analyzer against the deployed ensemble, not just the detector, and sweep concurrency so you see the latency and throughput curve rather than a single point.

Worked example: a 48-camera capacity plan

Suppose a retailer runs 48 cameras at 15 frames per second and needs results within 150 ms of capture: 720 frames per second in total. Profiling one GPU with the full ensemble gives illustrative numbers like these (measure your own; they vary with model size and GPU): GPU decode and letterbox take 1.2 ms per 1080p frame, the FP16 engine takes 9 ms for a batch of 16 including NMS, and network plus serialization add 4 ms.

At batch 16 one engine instance completes about 1,000 / 9 = 111 batches per second, or roughly 1,780 frames per second of model capacity, so the detector is not the limit. Decode is: at 1.2 ms per frame on a shared decode engine, the GPU handles about 830 frames per second, only about 16 percent above demand. The plan is therefore one GPU with headroom thin enough that a second is needed for failover and growth, and the first optimization is to reduce decode cost by having cameras stream at 720p, which the model downsamples anyway.

Latency per frame is queue delay (up to 2 ms) plus the batch wait, the 9 ms execution, the decode and the network, around 20 to 30 ms at the median. The 150 ms budget is safe; the risk is the tail when a frame burst arrives, which is why the batcher's queue delay stays small and the service sheds load rather than letting the queue grow without bound.

Video streams

Video changes the problem. Consecutive frames are highly correlated, so many systems run the detector on every second or third frame and use a lightweight tracker in between, which cuts GPU cost proportionally with a modest loss of responsiveness. Requests for a single stream should arrive in order, so route each stream to a consistent instance and drop stale frames when the queue for that stream exceeds one or two entries: a detection for a frame that is already 500 ms old is usually worth less than nothing. Keep decode next to the detector.

Failure modes

SymptomLikely causeFix
GPU utilization 30 percent, CPU saturatedJPEG decode, resize or NMS in Python on the CPUMove decode to nvJPEG or DALI, NMS into the engine or torchvision on CUDA
Boxes offset vertically on camera frames onlyLetterbox padding not removed or applied twiceStore r and padding per image; unit-test round trips on non-square frames
Accuracy drop after deployment, no errorsBGR versus RGB, different normalization or interpolationGolden-image test comparing serving output to training-time inference
Small objects vanish after INT8Calibration set lacks small objects; quantization errorCalibrate on production frames; gate on per-size mAP; keep sensitive layers FP16
p99 latency spikes under loadQueue delay plus unbounded queueSmall max queue delay, queue limits, load shedding, more instances
Huge responses on crowded scenesNo top-k capFixed max detections per image, documented in the API
Engine fails to load after driver upgradeTensorRT plans are tied to TensorRT version and GPUBuild engines in CI per target; never copy plans between machines

Trade-offs

Larger batches raise throughput and latency together; real-time video favors small batches and more instances, offline indexing favors the largest batch that fits. Baking NMS into the engine wins speed and loses runtime tunability. INT8 doubles capacity at a small accuracy cost that lands unevenly on small objects. NMS-free models simplify serving but may lag the best NMS-based models on your data, so compare on your own validation set. For the general theory of batch sizing, see GPU batching strategies and static batching; for server configuration in depth, Triton Inference Server; and for how engine compilation works for transformers rather than CNNs, TensorRT-LLM.

What to do next

  1. Profile the current service per stage (decode, preprocess, transfer, model, NMS, postprocess) and find which one is actually limiting throughput.
  2. Write a golden-image test that compares serving boxes against training-time inference on non-square frames, including the inverse letterbox.
  3. Export to ONNX, build a TensorRT FP16 engine with an optimization profile matching your batch range, and compare accuracy on your validation set.
  4. Move decode and resize onto the GPU with DALI or CUDA tensors, and NMS into the engine or a GPU kernel.
  5. Deploy behind Triton with dynamic batching, start with a 1 to 2 ms queue delay and two instances, and sweep concurrency with perf_analyzer.
  6. Only then evaluate INT8 with a production calibration set and per-size metrics.
  7. Add queue limits, stale-frame dropping for video, top-k caps and per-stage latency metrics before going live.
Key takeaway: Object detection serving is a pipeline problem before it is a model problem. Understand that the network emits thousands of candidates that need filtering, NMS and a coordinate inverse; keep decode, preprocessing, inference and NMS on the GPU; compile the detector with TensorRT and serve it behind a dynamic batcher with a small queue delay; and plan capacity from per-stage measurements, because decode or NMS often runs out before the model does.