Hugging Face Optimum is the part of the Hugging Face stack that takes a model you trained or downloaded with Transformers and gets it running on a runtime other than eager PyTorch: ONNX Runtime on a CPU or NVIDIA GPU, OpenVINO on Intel hardware, the Neuron SDK on AWS accelerators, and several more. It does three jobs. It exports a model to a portable graph without you writing tracing code, it optimizes and quantizes that graph through the target runtime's own tools, and it gives you model classes that keep the Transformers API, so a call to generate() or a pipeline keeps working after the swap.

Optimum is not an accelerator in its own right. Every millisecond you save comes from the runtime underneath, and every failure is an exporter or runtime limit showing through. This article covers the layers, what the exporter does, the optimization and quantization paths, a worked example and the failure modes. Flags were checked against the current optimum-onnx documentation. The packages were reorganised recently, so pin versions and re-check --help for your release.

Optimum is a family of packages

Optimum is a family of packages rather than one library. The core package, optimum, holds the shared machinery: the optimum-cli entry point, the task registry called TasksManager that maps a checkpoint to a task such as text classification or text generation, and the export configuration classes. Hardware-specific code lives in separate packages. The ONNX integration moved out of the core into optimum-onnx, and the project README now tells you to follow that package's install instructions for ONNX work. Intel support lives in optimum-intel (OpenVINO and Intel extensions), AWS accelerators in optimum-neuron, Gaudi in optimum-habana, and NVIDIA's TensorRT-LLM integration ships as a Docker image, huggingface/optimum-nvidia.

Pick a package by where the model will run. The table maps targets to packages and runtimes.

TargetInstallRuntime doing the workTypical use
x86 or ARM CPUpip install "optimum-onnx[onnxruntime]"ONNX Runtime CPU EPencoders, rerankers, small classifiers
NVIDIA GPU, generalpip install "optimum-onnx[onnxruntime-gpu]"ONNX Runtime CUDA or TensorRT EPencoders, vision, Whisper-class models
Intel CPU, GPU, NPUpip install optimum[openvino]OpenVINOclient and edge inference
AWS Inferentia / Trainiumpip install optimum[neuronx]Neuron SDK compilerfixed-shape compiled serving
NVIDIA GPU, LLM servingDocker imageTensorRT-LLMhigh-throughput decoder serving

Older tutorials install the ONNX Runtime path through extras on the core package, such as optimum[onnxruntime-gpu]. Some current documentation pages still show that form. Prefer the optimum-onnx package names above and install either onnxruntime or onnxruntime-gpu, never both. The two wheels conflict, and the usual symptom is CUDAExecutionProvider missing from the available providers.

Where Optimum sits between Transformers and a runtimeTransformers / Diffusers / timmPyTorch model + config + tokenizerOptimum core: TasksManager, export configs, CLIoptimum-cli, task inference, dummy inputs, validationoptimum-onnxexport + ORTModeloptimum-intelOpenVINO, IPEXoptimum-neuronAWS Trainium / Inferentiaoptimum-nvidia, -habanaand other backendsONNX RuntimeCPU, CUDA EP, TensorRT EP, other EPsVendor runtimes and compilersOpenVINO, Neuron SDK, TensorRT-LLMOptimum is glue and recipes: the speed comes from the runtime underneath, not from Optimum itself.
Optimum's core supplies task inference, export configs and the CLI; backend packages hand the graph to the runtime that actually executes it.

What the exporter actually does

Exporting to ONNX means running the model on example inputs and recording its operators as a static graph. A raw torch.onnx.export call leaves you to choose input names, dynamic dimensions, dummy tensors and KV-cache handling, and to check the result. Optimum answers these per architecture through an export configuration (an OnnxConfig subclass) registered for each supported model type.

The configuration declares the inputs and outputs with named dynamic axes such as batch_size and sequence_length, supplies dummy-input generators that build correctly shaped tensors (including past key and value tensors for decoders), sets a default opset, and sets an absolute tolerance for validation. TasksManager chooses the task from Hub metadata. For a local directory you must pass --task, otherwise you get the bare backbone without a task head, which is a quiet way to ship a model that returns hidden states instead of logits.

Generative tasks default to a -with-past variant with KV-cache inputs and outputs; without it every decoding step recomputes attention over the whole prefix. Encoder-decoder models are split into encoder and decoder files, because the encoder runs once and the decoder once per token. By default the decoder graphs with and without past are merged into one file so shared weights are stored once. --no-post-process keeps them separate and --monolith forces a single file.

What optimum-cli export onnx does, step by stepLoad checkpointconfig.json, weightsInfer taskTasksManagerPick OnnxConfiginputs, outputs, axesDummy inputsbatch, seq, past KVtorch.onnx.exportper submodel, opset NPost-processmerge decoders, dedupeValidate vs PyTorchshapes + values within atolOutput directorymodel.onnx (or several files) + config + tokenizerFails here most oftenunsupported op, custom code, atol mismatchValidation runs the same dummy batch through PyTorch and ONNX Runtime and compares every output.
The exporter's pipeline. Most failures surface at the export call (unsupported operator or custom code) or at validation (outputs differ beyond atol).
# CLI: task inferred from the Hub, past-KV export by default for generative models
optimum-cli export onnx --model distilbert-base-uncased-finetuned-sst-2-english sst2_onnx/

# Local checkpoint: name the task explicitly or you lose the classification head
optimum-cli export onnx --model ./my_finetuned --task text-classification --opset 17 sst2_onnx/

# Python equivalent, useful inside a build pipeline
from optimum.exporters.onnx import main_export
main_export("./my_finetuned", output="sst2_onnx", task="text-classification")

# What tasks does this architecture support for ONNX export?
from optimum.exporters.tasks import TasksManager
print(list(TasksManager.get_supported_tasks_for_model_type("distilbert", "onnx")))

Other flags matter in production. --dtype fp16 exports half-precision weights. --device cuda runs the export on a GPU, which the O4 preset requires. --atol overrides the validation tolerance. --dynamo selects PyTorch's newer exporter, which the CLI help recommends for opset 18 and above. --trust-remote-code executes Python from the model repository, so treat it like running an unknown script. For architectures without a registered config you can subclass an existing OnnxConfig and pass it through main_export(..., custom_onnx_configs=...). The documentation shows this for models whose key cache has an unusual layout.

Running the export with ORTModel

The ORTModelForXxx classes, such as ORTModelForSequenceClassification and ORTModelForCausalLM, wrap an ONNX Runtime InferenceSession behind the Transformers interface. They accept the same tokenizer outputs, return the same output dataclasses, and support generate() for decoders. Passing export=True to from_pretrained exports on the fly, which is convenient in a notebook and wrong in a server, where every cold start pays for it. Export once in CI.

The provider argument selects the execution provider. On the CUDA provider, Optimum turns on IOBinding by default (use_io_binding=True). IOBinding pre-allocates outputs on the device and avoids the copy to host after every call. The docs warn that without it, decoding can be slower than plain PyTorch. Always check model.providers after loading. If the CUDA provider failed to initialise, ONNX Runtime falls back to the CPU and your GPU benchmark is measuring the CPU.

from transformers import AutoTokenizer
from optimum.onnxruntime import ORTModelForSequenceClassification, pipeline

tok = AutoTokenizer.from_pretrained("sst2_onnx")
model = ORTModelForSequenceClassification.from_pretrained(
    "sst2_onnx", provider="CUDAExecutionProvider")
assert model.providers[0] == "CUDAExecutionProvider", model.providers

clf = pipeline("text-classification", model=model, tokenizer=tok, device="cuda:0")
print(clf("The latency budget was tight and the export still fit."))

Graph optimization, O1 to O4

After export, ORTOptimizer runs ONNX Runtime's transformer graph optimizer. It folds constants, removes redundant nodes, and fuses common subgraphs: GELU, LayerNorm, skip connections plus LayerNorm, bias plus GELU, and whole attention blocks into single fused kernels. AutoOptimizationConfig packages four presets, and the same presets are available as --optimize at export time.

PresetWhat it addsPortability
O1basic general optimizations (constant folding, redundant node removal)general ONNX
O2O1 plus extended optimizations and transformer-specific fusionsONNX Runtime contrib ops
O3O2 plus GELU approximationONNX Runtime only, tiny numeric change
O4O3 plus fp16 mixed precision; GPU only, needs --device cudaONNX Runtime GPU only

Fusions above O1 emit ONNX Runtime-specific operators. The CLI help says plainly that an --optimize graph may not be usable with OpenVINO or TensorRT. The GPU guide also recommends feeding the unoptimized export to the TensorRT execution provider, which does its own fusion. One subtle trap is attention fusion. It assumes right-side padding for BERT-like encoders and left-side padding for GPT-like decoders. If your tokenizer pads the other way, set use_raw_attention_mask=True in the config and accept a slower graph rather than silently wrong outputs on padded batches.

from optimum.onnxruntime import ORTOptimizer, AutoOptimizationConfig, ORTModelForSequenceClassification

model = ORTModelForSequenceClassification.from_pretrained("sst2_onnx")
optimizer = ORTOptimizer.from_pretrained(model)
optimizer.optimize(save_dir="sst2_o2", optimization_config=AutoOptimizationConfig.O2())

# CLI equivalent on an existing export
# optimum-cli onnxruntime optimize --onnx_model sst2_onnx/ -O2 -o sst2_o2/

Quantization and its hardware limits

ORTQuantizer applies ONNX Runtime's int8 quantization. Dynamic quantization stores weights as int8 and computes each activation's scale at run time. It needs no data and is the right first try for transformer encoders on a CPU. Static quantization fixes the activation ranges ahead of time from a calibration set, which removes the run-time range computation but makes accuracy depend on how representative those samples are. Configs are chosen per instruction set: AutoQuantizationConfig.avx512_vnni, .avx2, .arm64, and .tensorrt for the TensorRT-compatible form. per_channel=True uses one scale per output channel, which usually recovers accuracy at a small cost.

from functools import partial
from transformers import AutoTokenizer
from optimum.onnxruntime import ORTQuantizer, ORTModelForSequenceClassification
from optimum.onnxruntime.configuration import AutoQuantizationConfig, AutoCalibrationConfig

model = ORTModelForSequenceClassification.from_pretrained("sst2_onnx")
quantizer = ORTQuantizer.from_pretrained(model)

# Dynamic: no data needed
dq = AutoQuantizationConfig.avx512_vnni(is_static=False, per_channel=True)
quantizer.quantize(save_dir="sst2_int8_dyn", quantization_config=dq)

# Static: calibrate activation ranges on representative text
tok = AutoTokenizer.from_pretrained("sst2_onnx")
sq = AutoQuantizationConfig.avx512_vnni(is_static=True, per_channel=False)
calib = quantizer.get_calibration_dataset(
    "glue", dataset_config_name="sst2", num_samples=300, dataset_split="train",
    preprocess_function=partial(lambda ex, t: t(ex["sentence"]), t=tok))
ranges = quantizer.fit(dataset=calib,
                       calibration_config=AutoCalibrationConfig.minmax(calib),
                       operators_to_quantize=sq.operators_to_quantize)
quantizer.quantize(save_dir="sst2_int8_static", calibration_tensors_range=ranges,
                   quantization_config=sq)

Know the hardware limits before you quantize for a GPU. Optimum's GPU guide states that current ONNX Runtime limits rule out these int8 models on the CUDA execution provider. Dynamic quantization inserts operators the CUDA provider does not consume. Statically quantized graphs would run their matrix multiplies in floating point with quantize and dequantize nodes around them, so nothing is gained. The TensorRT provider accepts only static, symmetric quantization with fp32 weights in the file. So int8 on this path is a CPU tool; on NVIDIA GPUs, fp16 via O4 or the TensorRT provider is the usual win.

Worked example: a CPU sentiment classifier

Take a DistilBERT sentiment classifier on a CPU-only fleet with a p95 budget of 20 ms per sentence. (1) Export with an explicit task and check that every output validates. (2) Build an O2 graph and a dynamic per-channel int8 graph. (3) Score all three graphs and the PyTorch model on a few thousand labelled held-out sentences, recording accuracy and label agreement. (4) Measure latency with the harness below on the production instance type and thread count. (5) Ship the fastest graph whose accuracy drop is within tolerance, say 0.5 points, with the versions that produced it.

Trust your measurements, not published speedups. Gains depend on sequence length, batch size, threads and the CPU's int8 instructions; a graph that wins at batch 32 can lose at batch 1.

import time, numpy as np, torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from optimum.onnxruntime import ORTModelForSequenceClassification

def bench(model, enc, n=300, warmup=30):
    for _ in range(warmup):
        model(**enc)
    times = []
    for _ in range(n):
        t0 = time.perf_counter(); model(**enc); times.append(time.perf_counter() - t0)
    t = np.array(times) * 1000
    return {"p50_ms": round(float(np.percentile(t, 50)), 2), "p95_ms": round(float(np.percentile(t, 95)), 2)}

tok = AutoTokenizer.from_pretrained("sst2_onnx")
enc = tok(["A representative production sentence of typical length."], return_tensors="pt")
ref = AutoModelForSequenceClassification.from_pretrained("./my_finetuned").eval()
with torch.inference_mode():
    ref_logits = ref(**enc).logits
for path in ["sst2_onnx", "sst2_o2", "sst2_int8_dyn"]:
    m = ORTModelForSequenceClassification.from_pretrained(path)
    drift = (m(**enc).logits - ref_logits).abs().max().item()
    print(path, bench(m, enc), "max |dlogit| =", round(drift, 4))

Failure modes

SymptomLikely causeFix
Export fails with an unsupported operatorarchitecture or op not covered by the opset or exporterraise --opset, try --dynamo, or write a custom OnnxConfig
Validation reports values not closefp16 export, nondeterministic op, or an atol set too tightcompare in fp32 first; loosen --atol only after checking task accuracy
Model returns hidden states, not logitslocal export without --taskexport again with the task named
GPU run is as slow as CPUCUDA provider failed and ONNX Runtime fell back to CPUassert on model.providers; install only onnxruntime-gpu
Generation is slower than PyTorchIOBinding off, or exported without past KVkeep use_io_binding on; use a -with-past task
Wrong outputs on padded batchesattention fusion with an unexpected padding sidematch padding side or set use_raw_attention_mask=True
TensorRT provider errors on an optimized graphONNX Runtime-specific fused opsgive TensorRT the unoptimized export
Accuracy drops after static int8unrepresentative calibration data or outlier activationscalibrate on production-like text; try per-channel or dynamic

Trade-offs and when not to use it

Optimum earns its place for encoders, rerankers, embedding, speech and vision models on CPUs or mixed fleets, and when one export must feed several runtimes. It is a poor fit for high-throughput LLM serving, which needs continuous batching, paged KV caches and weight-only quantization from a dedicated server. If you stay in PyTorch on NVIDIA hardware, torch.compile is often a simpler first step with one artifact and one numerical path.

The costs: a second numerical path to validate on every model update, exports that break when Transformers changes a forward signature, and fast settings (O4, fp16, TensorRT engines) that tie the artifact to one runtime. Treat the export directory as a build output with a recorded toolchain.

Go deeper on the runtime underneath in ONNX Runtime, in depth, on the Intel path in OpenVINO for LLMs, on staying inside PyTorch with torch.compile, on fused attention kernels in FlashAttention, and on dedicated LLM serving in TensorRT-LLM.

What to do next

  1. Install optimum-onnx with exactly one of the CPU or GPU ONNX Runtime extras, in a fresh environment, and record versions.
  2. Export your model with an explicit --task and read every line of the validation log.
  3. Build O2 and dynamic int8 variants. Score each on a labelled held-out set against the PyTorch model.
  4. Benchmark p50 and p95 at your real batch size and thread count, with warmup, on the production instance type.
  5. On GPU, assert model.providers at startup and keep IOBinding on.
  6. Move export into CI, store the directory as a versioned artifact, and fail the build if accuracy or drift exceeds your thresholds.
  7. Re-run the whole pipeline whenever Transformers, Optimum or ONNX Runtime is upgraded.
Key takeaway: Optimum turns a Transformers checkpoint into a validated graph for another runtime and keeps the familiar API on top. The speed comes from that runtime. Export once in CI with an explicit task, keep O2 and higher fusions for ONNX Runtime only, use int8 as a CPU tool, check the active execution provider, and ship only what you have measured against the PyTorch reference on your own data and hardware.