TensorRT is NVIDIA's ahead-of-time compiler and runtime for neural network inference. You give it a model graph, usually as ONNX; it fuses layers, chooses the fastest available kernel for every layer on the exact GPU it is running on, plans memory, and writes the result to a plan file. At serving time a small runtime loads that plan and executes it on a CUDA stream. Everything expensive happens once at build time, which is why TensorRT engines are fast and also why they are fragile: a plan is specialized to a TensorRT version, a GPU and a range of input shapes.

This article explains the build and the runtime from the software side, for TensorRT 11.x. Release 11.0 changed the precision workflow: strong typing is mandatory and the builder flags that used to enable FP16 or INT8 are gone, so a lot of older tutorials no longer run. It covers the architecture, what happens during a build, precision under strong typing, building and running an engine in Python, dynamic shapes, measuring with trtexec, engine portability, a worked example for an embedding encoder, and the failure modes that show up in production. For large language models use the dedicated runtime covered in TensorRT-LLM; for TensorRT inside ONNX Runtime see ONNX Runtime.

The architecture

Two phases, two sets of objects. The build phase uses a Builder, an INetworkDefinition populated by the ONNX parser, and an IBuilderConfig carrying memory limits, optimization profiles and the timing cache. Its output is a serialized plan. The runtime phase deserializes the plan into an ICudaEngine, which holds the weights and the chosen kernels, and creates one or more IExecutionContext objects, each with its own activation memory and input shapes. A server typically loads one engine and creates a context per concurrent stream.

TensorRT 11: offline precision, an ahead-of-time build, a small runtimeONNX exportfrom PyTorchModelOptAutoCast or Q/DQ quantizeONNX parserINetworkDefinitionBuilderconfig, profilesBuild: graph optimizerfuse layers, remove no-ops, plan memorytime candidate kernels (tactics) per layerpick fastest legal tactic for each shape profileTiming cachereused across buildslookup / storePlan file (serialized engine)tied to TRT version and GPUdeserializeRuntime: ICudaEngineone per processExecution contextsset shapes, bind addresses, enqueueexecute_async_v3 on a CUDA stream
The TensorRT 11 pipeline. Precision is decided before the build, in the ONNX graph; the build is a search over kernels; the runtime only binds memory and launches work.

What happens during a build

A build does three kinds of work, and knowing them explains most surprises.

Graph optimization. TensorRT folds constants, removes no-op layers such as identity reshapes, and fuses sequences of layers into single kernels: a convolution with its bias, activation and residual add; a matrix multiply with its bias and activation; normalization and attention patterns where a fused kernel exists. Fusion saves the round trip of each intermediate tensor through GPU memory, which for many layers is a larger cost than the arithmetic. The general principle is covered in GPU Kernel Fusion.

Tactic selection. For each fused layer there are usually several implementations, called tactics, with different tilings, data layouts and math instructions. The builder runs candidates on the actual GPU with the profile's optimization shapes and keeps the fastest legal one. This is why builds take minutes, why they depend on the exact GPU model and clocks, and why two builds of the same model can pick different tactics when timings are close.

Memory planning. The builder assigns activations to a shared region, reusing memory between tensors whose lifetimes do not overlap, and sizes it for the largest shapes allowed. The workspace pool limit caps the scratch memory any tactic may request, so too tight a limit silently removes fast tactics from the search.

The timing cache records tactic timings keyed by layer configuration and GPU. Feeding it back into later builds skips re-timing layers it has seen, which cuts rebuild time sharply and makes tactic choices more repeatable. Keep one cache per GPU model and TensorRT version, and store it next to your build artifacts.

Precision in TensorRT 11

Precision is where TensorRT 11 differs most from what older guides describe. In 10.x and earlier, a network was weakly typed by default: you set builder flags such as FP16 or INT8 and the builder was free to run any layer in any enabled precision, with INT8 scales coming from a runtime calibrator. In 11.0, NVIDIA removed weak typing, all the precision-enabling builder flags, the INT8 calibrator classes and the matching trtexec options. Every network is strongly typed: each tensor's type in the graph is the type it runs in, and TensorRT will not change it.

So precision becomes a property of the ONNX file, decided before the build with NVIDIA Model Optimizer (ModelOpt). For mixed FP16 or BF16, the migration guide points to AutoCast; for INT8, FP8 or INT4, to ModelOpt quantization, which inserts explicit QuantizeLinear and DequantizeLinear (Q/DQ) nodes with calibrated scales.

# mixed precision: rewrite the ONNX graph to FP16 where it is safe
python -m modelopt.onnx.autocast --onnx_path model.onnx --output_path model_fp16.onnx

# post-training quantization with calibration data (npz/npy batches)
python -m modelopt.onnx.quantization \
    --onnx_path=model.onnx \
    --quantize_mode=int8 \
    --calibration_data_path=calib.npz \
    --output_path=model_int8.onnx

This is a better contract even though it is a migration cost. The precision of every tensor is visible in the artifact you version and review, accuracy can be checked on the ONNX model before any engine exists, and the build will not move a tensor to FP16 or INT8 on its own. One exception remains: the TF32 builder flag still exists and lets FP32 matrix multiplies and convolutions use TF32 Tensor Core math, so check its state in your build configuration and clear it when you need strict FP32. If you maintain 10.x builds, the old flags still work there, but plan the move: code that sets BuilderFlag.FP16 or assigns an int8_calibrator will fail on 11.

Building and running an engine in Python

The Python build is short. The parts that matter in production are reporting parser errors, persisting the timing cache, and declaring an optimization profile for every dynamic input.

import os
import tensorrt as trt

LOGGER = trt.Logger(trt.Logger.WARNING)

def build_engine(onnx_path, plan_path, cache_path, shapes, workspace_gib=4):
    builder = trt.Builder(LOGGER)
    network = builder.create_network(0)            # strongly typed in 11.x
    parser = trt.OnnxParser(network, LOGGER)
    with open(onnx_path, "rb") as f:
        if not parser.parse(f.read()):
            errors = [str(parser.get_error(i)) for i in range(parser.num_errors)]
            raise RuntimeError("ONNX parse failed:\n" + "\n".join(errors))

    config = builder.create_builder_config()
    config.set_memory_pool_limit(trt.MemoryPoolType.WORKSPACE, workspace_gib << 30)

    blob = open(cache_path, "rb").read() if os.path.exists(cache_path) else b""
    cache = config.create_timing_cache(blob)
    config.set_timing_cache(cache, ignore_mismatch=False)

    profile = builder.create_optimization_profile()
    for name, (lo, opt, hi) in shapes.items():     # one entry per dynamic input
        profile.set_shape(name, lo, opt, hi)
    config.add_optimization_profile(profile)

    plan = builder.build_serialized_network(network, config)
    if plan is None:
        raise RuntimeError("build failed; rerun with a VERBOSE logger")
    with open(plan_path, "wb") as f:
        f.write(plan)
    with open(cache_path, "wb") as f:
        f.write(config.get_timing_cache().serialize())

build_engine("encoder_fp16.onnx", "encoder.plan", "h100.cache", {
    "input_ids":      ((1, 16), (16, 128), (32, 512)),
    "attention_mask": ((1, 16), (16, 128), (32, 512)),
})

At runtime, the context must be told the actual input shapes before execution, and every input and output tensor must be bound to a device address. Using PyTorch tensors for the buffers keeps the example short; any CUDA allocation works.

import torch
import tensorrt as trt

runtime = trt.Runtime(LOGGER)
with open("encoder.plan", "rb") as f:
    engine = runtime.deserialize_cuda_engine(f.read())
context = engine.create_execution_context()

def embed(input_ids, attention_mask):
    context.set_input_shape("input_ids", tuple(input_ids.shape))
    context.set_input_shape("attention_mask", tuple(attention_mask.shape))
    out = torch.empty(tuple(context.get_tensor_shape("embeddings")),
                      dtype=torch.float16, device="cuda")   # match the graph's output type
    context.set_tensor_address("input_ids", input_ids.data_ptr())
    context.set_tensor_address("attention_mask", attention_mask.data_ptr())
    context.set_tensor_address("embeddings", out.data_ptr())
    stream = torch.cuda.current_stream()
    if not context.execute_async_v3(stream.cuda_stream):
        raise RuntimeError("enqueue failed")
    stream.synchronize()
    return out

Two rules prevent most runtime bugs. Inputs must be contiguous and of the dtype the engine expects, which under strong typing is exactly the dtype in the ONNX graph. And a context is not thread-safe: give each concurrent request stream its own context, created from the shared engine.

Dynamic shapes and optimization profiles

An optimization profile declares, for each dynamic input, a minimum, an optimum and a maximum shape. The engine accepts any shape inside the range; tactics are timed at the optimum shape; and activation memory is sized for the maximum. Those three facts give the design rules. A very wide range costs memory for every context and may run small inputs on tactics tuned for large ones. A range that is too narrow rejects real traffic at set_input_shape time.

When traffic has two distinct regimes, use two profiles in the same engine rather than one wide profile. Each context selects one profile, so a server keeps a pool of contexts per profile and routes requests by size. trtexec takes profile shapes on the command line, which is the quickest way to measure the effect before changing server code.

trtexec --onnx=encoder_fp16.onnx --saveEngine=encoder.plan \
        --minShapes=input_ids:1x16,attention_mask:1x16 \
        --optShapes=input_ids:16x128,attention_mask:16x128 \
        --maxShapes=input_ids:32x512,attention_mask:32x512 \
        --timingCacheFile=h100.cache

trtexec builds the engine, runs it repeatedly and reports latency percentiles and throughput, so it is also the baseline to compare your server against. If the server is much slower than trtexec for the same shapes, the gap is in your code: host copies, synchronization or Python overhead, which a profiler such as Nsight Systems will show on the timeline.

Worked example: an embedding encoder

Take a sentence-embedding encoder serving a search product. Logged traffic shows sequence lengths with a median of 64 tokens, a 99th percentile of 384 and a hard cap of 512, and the batcher forms batches of up to 32 requests. A single profile from 1x16 to 32x512 with the optimum at 16x128 works, but it sizes every context's activation memory for 32x512 and tunes tactics for medium inputs.

A better layout is two profiles. A short profile covers batch 1 to 32 and length 16 to 128, optimum 32x64, and serves almost all traffic. A long profile covers length 129 to 512, optimum 8x384, with a smaller maximum batch of 16 because long requests are rare. The router pads each request to the next bucket inside its profile and sends it to a context of that profile. Memory per long context drops because its maximum batch halved, and the short contexts run tactics timed at the shapes they actually see.

Validate the build before shipping it. Polygraphy, which ships with TensorRT, runs the same inputs through ONNX Runtime and TensorRT and compares outputs with a tolerance; for an embedding model also compare cosine similarity of embeddings on a held-out set, since small element-wise differences matter less than ranking changes.

polygraphy run encoder_fp16.onnx --trt --onnxrt \
    --trt-min-shapes input_ids:[1,16] attention_mask:[1,16] \
    --trt-opt-shapes input_ids:[32,64] attention_mask:[32,64] \
    --trt-max-shapes input_ids:[32,128] attention_mask:[32,128] \
    --atol 1e-2 --rtol 1e-2

Engines are not portable

A plan file is a compiled binary. By default it runs only with the TensorRT version that built it, and NVIDIA documents that it is only compatible with the type of device it was built on. Two relaxations exist, exposed in trtexec as --versionCompatible and --hardwareCompatibilityLevel and in the Python API through the builder configuration. The hardware level SAME_COMPUTE_CAPABILITY allows any GPU with the same compute capability, and AMPERE_PLUS allows Ampere and newer. Both can cost throughput or latency because the builder excludes tactics that are not valid everywhere the engine may run, and NVIDIA notes SAME_COMPUTE_CAPABILITY generally performs better than AMPERE_PLUS. Most teams instead treat engines as per-target build artifacts: build in CI for each GPU model in the fleet, name the plan by model hash, TensorRT version and GPU, and refuse to load a mismatched plan.

When only the weights change, for example a nightly fine-tune of the same architecture, an engine built with the REFIT flag can have new weights loaded through the refitter API without a full rebuild. That keeps the tactic choices fixed, which is usually what you want for stable latency.

Failure modes

  • Parser rejects an operator. Unsupported or exotic ONNX ops fail at parse time. Simplify the graph, change the export opset, rewrite the op in supported primitives, or implement a plugin as a last resort.
  • Accuracy loss after AutoCast or quantization. Overflow in reductions, softmax or normalization shows up as NaNs or drift. Compare against the original FP32 model, not the converted ONNX file, keep the offending nodes in FP32 in the ONNX graph, and re-check task metrics, not only tensors.
  • Shape out of profile. A request longer than the maximum fails at set_input_shape. Enforce the cap at the API edge and log rejections.
  • Version or GPU mismatch. Deserialization fails after a driver, container or node-type change. Version plans and check compatibility at startup.
  • Slow or different builds. Builds on a busy or throttled GPU time tactics badly. Build on an idle GPU, reuse the timing cache, and record which tactics changed when latency moves between releases.
  • Out of memory with many contexts. Each context reserves activation memory for its profile's maximum; check device_memory_size_v2 and size the context pool from it.

Trade-offs

TensorRT trades flexibility for speed. You give up eager execution, arbitrary shapes and portable binaries, and you take on a build step, a precision workflow and per-GPU artifacts. In return you get fused kernels selected for your hardware and a runtime with almost no overhead. For a stable model that serves heavy traffic on a known fleet, that trade is usually worth it. For models that change weekly, or graphs full of unsupported ops, ONNX Runtime with the TensorRT provider or torch.compile may be the better first step. In a multi-model server, Triton can host TensorRT plans next to other backends.

What to do next

  1. Export your model to ONNX and confirm it parses with the build function above, printing every parser error.
  2. Run ModelOpt AutoCast, compare outputs against the FP32 model, and pin any sensitive nodes to FP32.
  3. Derive optimization profiles from logged input shapes, using two profiles if traffic has two regimes.
  4. Build with trtexec and a persisted timing cache, and record latency percentiles at your real shapes.
  5. Add a CI job that builds one plan per GPU model and TensorRT version and names it by all three.
  6. Benchmark your server against trtexec and profile any gap before tuning the model further.
Key takeaway: TensorRT does its work at build time: it fuses the graph, times real kernels on your GPU and freezes the result into a plan. In TensorRT 11 precision lives in the ONNX graph, so decide it with ModelOpt, size optimization profiles from real traffic, keep a timing cache, and treat every plan as a per-GPU, per-version build artifact.