ONNX Runtime (ORT) is Microsoft's open-source inference engine for models in the ONNX format. You export a model from PyTorch or another framework to an ONNX graph, and ORT runs that graph on CPUs, NVIDIA and AMD GPUs, NPUs, Apple silicon and browsers through pluggable execution providers (EPs). The appeal is one artifact and one API across very different hardware. The trap is that the same flexibility lets a model run correctly but slowly, partly on the CPU, without any error.
This article explains what happens inside a session, from graph optimisation through partitioning to execution, and how GPU memory, copies and kernels behave. It covers provider configuration, IOBinding and CUDA graphs, the TensorRT EP, quantization, a worked example of finding a hidden CPU fallback, failure modes, trade-offs and a checklist. API names below come from the ORT documentation; check the version you deploy, since providers are added and removed between releases. For example, release 1.23 removed the ROCm EP from the source tree and pointed AMD GPU users to the MIGraphX or Vitis AI EPs, and added an NVIDIA TensorRT RTX EP.
Inside a session
Creating an InferenceSession does most of the work up front. ORT parses the protobuf graph, applies graph transformations, then asks each registered EP, in the order you listed them, which nodes it can run. An EP claims maximal subgraphs it supports; the TensorRT EP compiles each claimed subgraph into one engine, while the CUDA EP keeps per-operator kernels. Whatever no accelerator claims lands on the CPU EP, which supports every standard operator. ORT then inserts memory copy nodes wherever a tensor crosses a device boundary, plans buffer reuse, and creates allocators. By default GPU memory comes from an arena that grows and is then reused rather than returned to the driver.
At run() time ORT walks the plan. On the GPU path each kernel is a CUDA launch on a stream, so small operators cost more in launch overhead than in arithmetic. That is why graph fusion and CUDA graphs matter so much for latency-bound models, and why a single unsupported operator in the middle of a graph hurts: it splits one GPU region into two and adds a round trip through host memory.
Configuring providers you can trust
A production session should state its providers explicitly, assert that the ones you expect actually loaded, and keep profiling one flag away. The Python package for NVIDIA GPUs is onnxruntime-gpu; installing it alongside the CPU-only onnxruntime package in the same environment is a common source of confusion about which build is imported.
import onnxruntime as ort
so = ort.SessionOptions()
so.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL
so.intra_op_num_threads = 4 # CPU threads used inside one operator
so.log_severity_level = 2 # 0 = verbose, shows node placement
# so.enable_profiling = True # writes a JSON trace with per-node timing
providers = [
("CUDAExecutionProvider", {
"device_id": 0,
"arena_extend_strategy": "kSameAsRequested",
"gpu_mem_limit": 6 * 1024**3,
"cudnn_conv_algo_search": "HEURISTIC",
}),
"CPUExecutionProvider",
]
sess = ort.InferenceSession("model.onnx", sess_options=so, providers=providers)
active = sess.get_providers()
if active[0] != "CUDAExecutionProvider":
raise RuntimeError(f"CUDA EP did not load, running on {active}")The assertion matters because when the CUDA EP fails to initialise, for example because the CUDA or cuDNN libraries it was built against are missing, ORT logs a warning and continues with the remaining providers. A service then passes every functional test while running many times slower. cudnn_conv_algo_search accepts EXHAUSTIVE, the default, which benchmarks convolution algorithms on first use, HEURISTIC or DEFAULT. Exhaustive search gives the fastest steady state but a slow first request for every new input shape. kSameAsRequested makes the arena grow by exactly what is requested instead of the default kNextPowerOfTwo, trading some allocation speed for less wasted memory when several models share a GPU.
Graph optimisation levels
Graph optimisation has four levels. ORT_ENABLE_BASIC does semantics-preserving rewrites such as constant folding, removing redundant Identity and Dropout nodes, and folding BatchNorm into Conv. ORT_ENABLE_EXTENDED adds fusions with provider-specific kernels: GELU, LayerNorm, Attention and SkipLayerNorm fusions, which are where transformer models gain the most. ORT_ENABLE_ALL adds layout optimisations such as the CPU NCHWc layout. ORT_DISABLE_ALL turns everything off, which is useful for bisecting a numerical difference.
Optimisation runs online at every session creation by default. For large models you can run it once offline by setting optimized_model_filepath and loading the saved result later with optimisation disabled. The documentation is explicit that the offline model is tied to the options, providers and hardware used to produce it; a model whose layout was optimised for AVX2 needs a CPU with AVX2. Treat an optimised model as a build artifact for a specific target, never as a portable model file.
IOBinding and CUDA graphs
By default sess.run() takes NumPy arrays, copies inputs to the GPU and copies outputs back on every call. For a pipeline whose data already lives on the GPU, or whose output feeds the next GPU stage, those copies are pure waste. IOBinding lets you hand ORT device buffers directly. It is also a prerequisite for CUDA graphs, which record the whole sequence of kernel launches once and replay it with a single launch, removing per-kernel CPU overhead.
import numpy as np
import onnxruntime as ort
sess = ort.InferenceSession(
"model.onnx",
providers=[("CUDAExecutionProvider", {"enable_cuda_graph": True})],
)
x_host = np.zeros((8, 3, 224, 224), dtype=np.float32)
x_dev = ort.OrtValue.ortvalue_from_numpy(x_host, "cuda", 0)
y_dev = ort.OrtValue.ortvalue_from_shape_and_type((8, 1000), np.float32, "cuda", 0)
binding = sess.io_binding()
binding.bind_ortvalue_input("input", x_dev)
binding.bind_ortvalue_output("logits", y_dev)
sess.run_with_iobinding(binding) # first run captures the graph
def infer(batch):
x_dev.update_inplace(batch) # same address, new contents
sess.run_with_iobinding(binding) # replays the captured graph
return y_dev.numpy()CUDA graphs come with strict conditions, all from the CUDA EP documentation: every node must be assigned to the CUDA EP, the model cannot contain control-flow operators (If, Loop, Scan), input and output shapes and addresses must not change between calls, and you may not call run on the same session from several threads at once. That is why the example preallocates fixed buffers and updates them in place. Variable batch sizes need one captured graph per shape or padding to a fixed shape.
The TensorRT execution provider
The TensorRT EP hands supported subgraphs to TensorRT, which fuses layers, picks tactics per GPU and can run in FP16 or INT8. List it first with CUDA and CPU behind it, so unsupported nodes fall back. Engine builds can take minutes, so enable the engine cache, and the timing cache that speeds up rebuilds.
trt_options = {
"trt_fp16_enable": True,
"trt_engine_cache_enable": True,
"trt_engine_cache_path": "/var/cache/ort-trt",
"trt_timing_cache_enable": True,
"trt_max_workspace_size": 2 * 1024**3,
}
providers = [
("TensorrtExecutionProvider", trt_options),
"CUDAExecutionProvider",
"CPUExecutionProvider",
]
sess = ort.InferenceSession("model.onnx", providers=providers)The documentation tells you to delete cached engines and profiles whenever the model, the ORT version or the TensorRT version changes. Key your cache directory on all three plus the GPU model, and build it in your deployment pipeline rather than on the first production request. With dynamic input shapes, TensorRT builds an optimisation profile per shape range; inputs outside the range trigger a rebuild, which shows up as a multi-second latency spike.
Quantization
ORT ships a quantization toolkit in onnxruntime.quantization. Dynamic quantization stores weights as INT8 and computes activation scales at run time; it needs no calibration data and suits transformer models on CPU. Static quantization also fixes activation scales from a calibration set and writes QuantizeLinear and DequantizeLinear (QDQ) pairs into the graph, the format accelerator EPs such as TensorRT understand.
from onnxruntime.quantization import (
CalibrationDataReader, QuantFormat, QuantType, quantize_dynamic, quantize_static)
quantize_dynamic("model.onnx", "model.int8.onnx", weight_type=QuantType.QInt8)
class Calib(CalibrationDataReader):
def __init__(self, batches):
self._it = iter([{"input": b} for b in batches])
def get_next(self):
return next(self._it, None)
quantize_static("model.onnx", "model.qdq.onnx", Calib(sample_batches),
quant_format=QuantFormat.QDQ,
activation_type=QuantType.QInt8, weight_type=QuantType.QInt8)Use a few hundred representative inputs for calibration, not random tensors, and gate the quantized model on your own accuracy metric before shipping. If one layer degrades badly, exclude it with the nodes_to_exclude argument rather than abandoning quantization.
Worked example: the hidden CPU hop
A team exports an image classifier and sees 9 ms on an older engine but 30 ms in ORT with the CUDA EP. get_providers() shows CUDA first, so the GPU appears to be in use. They turn on enable_profiling and open the trace: one custom post-processing operator near the middle of the graph has no CUDA kernel, so it runs on the CPU EP. The graph is now GPU, CPU, GPU, with a copy at each boundary and a stream synchronisation before each device-to-host copy.
Do the arithmetic. The tensor at the boundary is 8 x 3 x 224 x 224 FP32 values, 4.8 MB. At an effective 25 GB/s over PCIe 4.0 x16 each copy takes about 190 µs, and there are two per call, plus the CPU kernel itself on a single thread, plus the lost overlap from the synchronisations. Moving that operator after the model output, or rewriting it with standard ONNX operators that CUDA supports, collapses the graph into one GPU region. The copies themselves are under half a millisecond; the profile shows that the single-threaded CPU kernel and the stalls around the synchronisations account for most of the 21 ms gap. The general lesson is that with ORT, placement explains most latency surprises, so look at node placement before tuning kernels.
Failure modes
- Silent provider fallback. Missing CUDA or cuDNN libraries, or the CPU wheel imported instead of the GPU wheel, leaves you on the CPU. Assert on
get_providers()at startup. - Partial placement. One unsupported operator splits the graph and inserts copies. Check verbose logs or the profile for nodes on the CPU EP.
- First-request latency. Exhaustive cuDNN search, TensorRT engine builds and CUDA graph capture all happen on first use. Warm up every expected shape before taking traffic.
- Arena growth. The arena keeps memory it has grabbed. Several sessions per GPU need
gpu_mem_limitand a smaller extension strategy, or they exhaust memory under bursty shapes. - Thread oversubscription. Each CPU session starts its own thread pools. Many sessions in one process with default settings fight for cores; set
intra_op_num_threadsdeliberately. - Stale caches and optimised models. TensorRT engines and offline-optimised graphs are tied to versions and hardware. Rebuild them on every upgrade.
- Export mismatches. An opset the runtime does not support, or dynamic axes that were not declared at export, fail at load or force fixed shapes. Pin the opset and test the exported graph against framework outputs with a tolerance.
Trade-offs between providers
| Execution provider | Best at | Watch out for |
|---|---|---|
| CPU | Universal coverage, small models | Thread tuning, AVX features in optimised models |
| CUDA | Broad GPU coverage, simple setup | Launch overhead on small ops, CUDA and cuDNN versions |
| TensorRT | Lowest latency on NVIDIA GPUs | Engine build time, cache invalidation, shape profiles |
| DirectML | Any DirectX 12 GPU on Windows | Windows only, operator coverage varies |
| OpenVINO | Intel CPUs, GPUs and NPUs | Separate build, Intel hardware focus |
| CoreML | Apple silicon and Neural Engine | Partial operator coverage causes splits |
| MIGraphX | AMD GPUs since the ROCm EP removal | Check support for your model's operators |
What to do next
- Pin
onnxruntime-gpu, CUDA, cuDNN and TensorRT versions together and record them with every model artifact. - Add a startup assertion on
get_providers()and fail the deployment if the accelerator is missing. - Profile one request with
enable_profilingand list every node that runs on the CPU EP; remove or move each one. - Use IOBinding when inputs or outputs live on the GPU; try CUDA graphs for fixed-shape, latency-bound models.
- Evaluate the TensorRT EP with FP16, with the engine cache built in CI and keyed on model, ORT, TensorRT and GPU.
- Try dynamic INT8 on CPU and static QDQ on accelerators, gated on your accuracy metric.
- Warm up every expected input shape before a replica receives traffic.
Keep learning: TensorRT-LLM for LLM-specific serving, kernel fusion for why fused graphs are faster, inference latency for measuring tail latency, and OpenVINO and edge inference for non-NVIDIA targets.