OpenVINO is Intel's open-source inference toolkit. For language models it is the path that turns a Hugging Face checkpoint into something that runs well on hardware most people already own: Intel Xeon and Core CPUs, integrated and discrete Arc GPUs, and the NPUs in recent Core Ultra laptops. It is not a training framework and not a CUDA replacement; it is a compiler and runtime for a frozen graph, with a generation layer on top.

This article explains what each layer does, what the export step really produces, how to choose weight compression, how to budget memory for a concrete 8B model, how to run and tune it on each device, and where deployments fail. Command flags and properties below were checked against the current Optimum Intel and OpenVINO GenAI documentation; options move between releases, so confirm them against the version you install. For the hardware side of laptop accelerators, read Edge NPUs, in depth alongside this.

The four layers of the stack

There are four layers, and most confusion comes from mixing them up.

  • Optimum Intel is the Hugging Face integration. Its optimum-cli export openvino command loads the PyTorch model, traces it, converts it to OpenVINO's intermediate representation and, through the NNCF library, compresses the weights.
  • OpenVINO IR is the on-disk model: an .xml file describing the graph and a .bin file holding the weights. The export also writes tokenizer and detokenizer models, so tokenisation runs inside OpenVINO rather than in Python.
  • OpenVINO Runtime compiles the IR for one device through a plugin. Compilation applies graph transformations and selects kernels for that hardware; it is the step that takes time on first load.
  • OpenVINO GenAI provides LLMPipeline, which owns the generation loop: sampling, chat history, streaming, speculative decoding and, on CPU and GPU, a continuous-batching scheduler with a paged KV cache. The OpenVINO Model Server wraps the same machinery behind an OpenAI-compatible HTTP API.

From a Hugging Face checkpoint to tokens on an Intel CPU, GPU or NPUHF checkpointsafetensors + tokenizeroptimum-cli exporttrace, convert, NNCFOpenVINO IR.xml graph + .binTokenizer IRtokenize, detokenizeOpenVINO GenAI: LLMPipelinesampling, chat history, streaming, schedulerModel Server (OVMS)OpenAI-style HTTP APIOpenVINO Runtime: compile_model per devicegraph transforms, kernel selectionCPU pluginAMX / AVX-512 / AVX2, u8 KVGPU pluginArc / iGPU, XMX enginesNPU pluginstatic shapes, int4 symThe export decides weight precision; the plugin decides kernels, KV-cache precision and shapes at compile time.
The OpenVINO path for LLMs. Export is offline and decides weight precision; compilation happens per device at load time.

What export produces

The minimal command is optimum-cli export openvino --model <id> <output_dir>. Three behaviours of that command matter more than any flag.

First, models larger than one billion parameters are exported with 8-bit weights by default; pass --weight-format fp16 or fp32 if you want an uncompressed baseline for accuracy comparison. Second, decoder models are exported as stateful models by default: the KV cache is held inside the compiled model as internal state rather than exposed as inputs and outputs. That is what lets the runtime keep the cache on the device between steps. --disable-stateful exists for old code that expects explicit past-key-value tensors and is documented as possibly slower. Third, with --weight-format int4, the embedding table and the final projection are compressed to 8 bits unless you pass --all-layers, because those two layers are the most sensitive to error.

The compression flags that shape quality and speed are these:

FlagWhat it controlsGuidance
--weight-format int4|int8|nf4|fp16Weight precisionint4 for memory-bound decode; int8 when quality matters more
--symSymmetric quantization, no zero pointsRequired for the NPU; slightly less accurate than asymmetric
--group-size 128|-1Scale granularity; -1 means per output channel128 is the documented recommendation; channel-wise trades accuracy for speed
--ratio 0.8Fraction of layers in 4 bits, the rest in the backup precisionLower ratio recovers accuracy at more memory
--awqActivation-aware scaling of salient channelsData-free without a dataset; data-aware with one
--scale-estimationFits scales to minimise layer output errorNeeds --dataset; extra time and memory
--gptqLayer-wise error-compensating quantizationSlowest export; try when AWQ is not enough
--dataset wikitext2Calibration data for data-aware methodsEnglish sets work for many non-English models

A sensible starting point for a GPU or CPU is --weight-format int4 --group-size 128 --ratio 1.0 --awq --scale-estimation --dataset wikitext2, then measure; if accuracy falls short, lower the ratio before reaching for GPTQ.

Worked example: budgeting an 8B model

Take Llama 3.1 8B: 32 layers, hidden size 4096, 8 key-value heads of dimension 128, and a vocabulary of 128,256 with separate input embedding and output head. The embedding and head together hold about 1.05 billion parameters; in int8 that is roughly 0.98 GiB. The remaining 6.98 billion parameters in int4 take about 3.25 GiB, plus around 0.1 GiB of fp16 scales at group size 128. Expect a weight file near 4.3 GiB; check the real .bin size after export, because mixed precision and metadata move it.

The KV cache is separate. Per token it stores keys and values for every layer: 2 times 32 layers times 8 heads times 128 dimensions times 2 bytes in fp16 is 131,072 bytes, 128 KiB. A 4,096-token context costs 0.5 GiB per sequence, and a 32k context costs 4 GiB, which is now as large as the weights. Quantizing the cache to u8 roughly halves it, plus a small overhead for per-group scales. Add the runtime, compiled kernels and activations, and an 8B int4 model with a few thousand tokens of context fits comfortably in a 16 GB laptop sharing memory with the integrated GPU, while a 32k context does not leave much room. The general method for this arithmetic is in KV cache sizing.

Decode speed follows from the same numbers. Each generated token reads every weight once, so at about 4.3 GB per token a system with 100 GB/s of usable memory bandwidth cannot exceed roughly 23 tokens per second for one sequence, whatever the compute. That ceiling is why int4 matters more than raw TOPS on client hardware, and why batching several sequences, which reuses each weight read, is the main lever for throughput on servers.

Running it

Export once, then load with GenAI. The device string is the only change between CPU, GPU and NPU at this level.

# pip install openvino-genai optimum[openvino]
# optimum-cli export openvino --model meta-llama/Llama-3.1-8B-Instruct \
#     --weight-format int4 --group-size 128 --awq --scale-estimation \
#     --dataset wikitext2 llama31-8b-int4-ov
import openvino_genai as ov_genai

pipe = ov_genai.LLMPipeline("llama31-8b-int4-ov", "GPU")   # or "CPU", "NPU"

config = pipe.get_generation_config()
config.max_new_tokens = 256
config.temperature = 0.2
config.top_p = 0.9

def streamer(subword):
    print(subword, end="", flush=True)
    return False            # bool convention: True asks the pipeline to stop

pipe.start_chat()           # keeps history and reuses the KV cache across turns
pipe.generate("Summarise what a stateful OpenVINO model is.", config, streamer)
pipe.generate("Now in one sentence.", config, streamer)
pipe.finish_chat()

Newer GenAI releases also accept a streaming-status return value from the callback; check the API reference of your version. For serving many users, GenAI exposes a SchedulerConfig whose cache_size field sets the KV-cache pool in GB, and the Model Server exposes the same scheduler through its configuration, with options such as maximum concurrent sequences, batched-token limits and prefix caching. The pool is paged in blocks, the same idea as paged KV caches in other engines: memory is committed per block as sequences grow, rather than reserved for the maximum length up front.

Speculative decoding

Because decode is bandwidth-bound, the cheapest speedup is to verify several tokens per weight read. GenAI supports speculative decoding: a small draft model from the same family proposes a few tokens, and the main model checks them all in one forward pass, keeping the longest accepted prefix. With greedy decoding the output matches the main model alone; with sampling, check your release's documentation and compare outputs before assuming the same. The speedup depends on how often the draft agrees.

import openvino_genai as ov_genai

scheduler = ov_genai.SchedulerConfig()
scheduler.cache_size = 2                       # KV-cache pool in GB

draft = ov_genai.draft_model("llama32-1b-int4-ov", "CPU")
pipe = ov_genai.LLMPipeline("llama31-8b-int4-ov", "GPU",
                            scheduler_config=scheduler, draft_model=draft)

config = pipe.get_generation_config()
config.max_new_tokens = 200
config.num_assistant_tokens = 5                # tokens proposed per step
print(pipe.generate("Explain paged attention briefly.", config))

The draft and main model must share a tokenizer. Measure acceptance on your own prompts: code and structured output usually accept long runs, while creative text accepts fewer, and a draft that is rejected most of the time makes generation slower than plain decoding. Placing the draft on the CPU and the main model on the GPU, as above, is one option worth benchmarking against putting both on one device.

Device by device

CPU. The CPU plugin uses AMX on recent Xeons and AVX-512 or AVX2 elsewhere. Intel documents two LLM-specific properties: KV_CACHE_PRECISION, where u8 quantizes the cache group-wise, and DYNAMIC_QUANTIZATION_GROUP_SIZE, which quantizes activations on the fly so int4 or int8 weights can feed integer kernels; the KV group size defaults to 32 when the latter is not set. Both are documented as enabled or available by default on CPU in recent releases; set them explicitly when you benchmark so a version change does not silently alter your baseline. Pin threads with the performance hints rather than by hand, and on dual-socket servers run one instance per socket.

GPU. Integrated and Arc GPUs use their matrix engines for prefill and are bandwidth-bound in decode, like every other GPU. The first load compiles kernels, which can take tens of seconds for a large model; set the runtime's CACHE_DIR property so later loads reuse the compiled blob. Discrete cards have their own memory, so the budget above applies to VRAM; integrated GPUs share system memory with everything else.

NPU. The NPU wants static shapes and a narrow set of weight formats. Intel's guide requires --sym, int4 or nf4 weights, --ratio 1.0, and either --group-size 128 for models up to about 4 to 5 billion parameters or --group-size -1 (channel-wise) for larger ones. The pipeline compiles for a fixed prompt window and response length, set with MAX_PROMPT_LEN (default 1024) and MIN_RESPONSE_LEN (default 128), passed as a config dictionary to the constructor; their sum is the maximum context. Compilation is slow, so use the documented NPU cache-directory option, which the releases we checked call NPUW_CACHE_DIR. The NPU's value is power efficiency for a single user, not throughput.

Failure modes

  • Silent int8 default. A team exports without --weight-format, compares against llama.cpp int4, and concludes OpenVINO is slow. Check the precision you actually exported.
  • Prompt longer than the NPU window. A prompt beyond MAX_PROMPT_LEN fails or is rejected rather than streaming slowly; size the window from real prompts, including chat history.
  • Version skew. The IR, openvino, openvino-genai and openvino-tokenizers must come from matching releases; a mismatched tokenizer extension is a common cause of load errors.
  • Unsupported architectures. New model families need exporter support and sometimes a specific transformers version; Optimum Intel's release notes list the compatible range.
  • Quality regressions hidden by perplexity. int4 with channel-wise scales can pass a perplexity check and still fail at arithmetic or tool-call formatting; evaluate on your task.
  • Cold start. Without a compile cache, every container start pays compilation again, which looks like a hang to an orchestrator health check.

Operating it

Treat the exported directory as a build artifact. Record the exact export command, library versions and calibration dataset next to it, and version it like a container image. Gate releases on a small task evaluation that compares the fp16 export with the compressed one, not just on throughput. On servers, watch time to first token separately from inter-token latency: the first tracks prefill and queueing, the second tracks memory bandwidth and batch size. Warm the compile cache at image build time, or mount it from a volume. And benchmark against the alternatives you could run on the same box, such as llama.cpp or Ollama, with matched precision and context, before committing.

Trade-offs

DecisionOption AOption B
Weight precisionint4: smallest, fastest decodeint8: safer quality, about 1.7x the memory for an 8B model
Group size128: better accuracyChannel-wise: faster, needed for large models on NPU
DeviceCPU and GPU: dynamic shapes, batchingNPU: low power, static window, single user
APIGenAI in-process: lowest overheadModel Server: HTTP, multi-client, scheduler
CalibrationData-free: fast, reproducibleAWQ, scale estimation or GPTQ with data: better int4 quality

What to do next

  1. Export your model twice, as fp16 and as int4 with AWQ and scale estimation, and record both commands.
  2. Compute weight and KV memory for your context length with the arithmetic above, then confirm with the .bin size and a measured peak.
  3. Run the pipeline code on CPU and GPU, measure time to first token and tokens per second, and set CACHE_DIR before measuring load time.
  4. If you target an NPU, re-export with the symmetric recipe and size MAX_PROMPT_LEN from real prompts.
  5. Build a 50-prompt task evaluation and gate every re-export on it.
  6. Compare with the engines in LLM serving stacks if you need multi-GPU serving.
Key takeaway: OpenVINO turns a Hugging Face model into a stateful IR whose weight precision is fixed at export, then compiles it per device. Export explicitly, because large models default to int8; use int4 with data-aware calibration for client decode speed; budget weights and KV cache from the model's shapes; set compile caches; and use the symmetric, static-window recipe only when you target the NPU.