The Qualcomm Cloud AI 100 is a PCIe inference accelerator. It does not train models. It runs them at low power. The top SKU, the Cloud AI 100 Ultra, puts four SoCs, 64 AI cores, 576 MB of on-die SRAM and 128 GB of LPDDR4x on one 150 W card. A typical training-class GPU instead pairs a few tens of megabytes of cache with HBM that is far faster and far smaller. That difference shapes everything you do with the card. The model must be compiled ahead of time into a fixed program. The compiler, not a hardware cache, decides which tensors live in fast memory. For large language models, the decode rate is bounded by how many bytes of LPDDR you read per generated token.
This article treats the card as a software target. It covers what is inside an AI core and how the memory tiers behave. It walks the path from a PyTorch model to a compiled QPC binary. A worked example sizes an 8B-parameter model, and the article closes with the failure modes and trade-offs you will meet in production. Specifications are taken from Qualcomm's SDK documentation and product pages as of October 2026. Where a figure differs between those sources, the article says so.
The SKUs and the two numbers that matter
Qualcomm documents several SKUs built from the same AI core. The Standard and Pro cards each carry one SoC. The Ultra carries four, and a cut-down AI 080 Ultra enables fewer cores. The SDK's own table lists:
| SKU | SoCs | AI cores | DRAM | On-die SRAM | INT8 peak | FP16 peak |
|---|---|---|---|---|---|---|
| Standard | 1 | 14 | 16 GB | 126 MB | 325 TOPS | 110 TFLOPS |
| Pro | 1 | 16 | 32 GB | 144 MB | 375 TOPS | 125 TFLOPS |
| Ultra | 4 | 64 | 128 GB | 576 MB | 870 TOPS | 290 TFLOPS |
| AI 080 Ultra | 4 | 32 | 128 GB | 576 MB | 618 TOPS | 222 TFLOPS |
The Ultra product page quotes 288 FP16 TFLOPS where the SDK table rounds to 290. It also gives a 150 W power budget, a PCIe Gen4 x16 host link and 548 GB/s of LPDDR4x bandwidth for the card. The SRAM figures follow directly from the core design. Each core has 9 MB of local memory, made of 8 MB of VTCM and 1 MB of L2, so 16 cores give 144 MB and 64 cores give 576 MB.
Two numbers matter most for planning. The first is the ratio of peak compute to DRAM bandwidth. At 288 FP16 TFLOPS against 548 GB/s, the Ultra can do over 500 floating-point operations for every byte it reads from LPDDR. Any workload with lower arithmetic intensity than that is memory-bound. Single-stream LLM decode falls into that group. The second is capacity. With 128 GB, a 70B-parameter model in 8-bit weights fits on one card with about 58 GB left for KV cache, where an 80 GB GPU keeps only about 10 GB. The card wins on capacity per watt and loses on bandwidth per byte stored.
Inside an AI core
Each AI core has three compute units. The tensor unit has two 2D MAC arrays, sized for 8,192 INT8 or 4,096 FP16 multiply-accumulates per clock. Matrix multiplies and convolutions go here. The vector unit covers elementwise work, reductions, softmax and activation functions, at 512 INT8 or 256 FP16 MACs per clock. The scalar unit is a 4-way VLIW processor with six hardware threads. It runs control flow and any operators the other units cannot express.
The important word is scratchpad. The 8 MB VTCM is Vector Tightly Coupled Memory. It is addressed explicitly, and nothing evicts a line on a miss. The compiler splits each layer into tiles that fit and schedules DMA transfers from LPDDR through the L2 to bring the next tile in while the current one computes. This is why the card wants a static graph with known shapes. A GPU can absorb a surprising shape at run time because its caches and warp scheduler adapt. The AI 100 needs the shape at compile time to build the tiling plan.
Precision follows the units. The tensor unit natively handles INT8 and FP16. FP32 is available on the vector and scalar units, so a model left in FP32 runs far below peak. In practice you compile in FP16, or quantize weights and activations to INT8. For LLMs, Qualcomm's tooling also offers a 6-bit microscaling format for matmul weights, covered below.
Four SoCs on one card: the data path
The Ultra looks like one card but is four SoCs. Each SoC has its own cores, its own SRAM and its own share of LPDDR. Work that spans SoCs has to be partitioned so that data crosses between them as rarely as possible. For a transformer, tensor-parallel splits put a collective in every layer. Pipeline splits put one hand-off per stage. Which works better depends on batch size and on the compiler release. Treat it as something to measure on your SDK version rather than assume.
The data flow for one inference has five steps. First, the host runtime loads a QPC (the compiled program container) onto the device or devices. Second, it copies input tensors over PCIe. Third, cores pull weights from LPDDR into VTCM tile by tile. Fourth, partial results move between cores, and between SoCs where the graph is split. Fifth, outputs return to the host. For LLM serving, the KV cache stays resident in device DRAM between steps. Only token ids and logits, or sampled tokens, cross PCIe.
From model to QPC
The software stack has two layers. The Platform SDK supplies the kernel driver, firmware and runtime. The upstream Linux kernel carries the driver as the qaic accel driver. The Apps SDK supplies the compiler and tools. The flow is: export the model to ONNX, compile it into a QPC for a fixed set of input shapes and a fixed core count, then load and run that QPC. Qualcomm's documentation shows the compiler invoked with profiling enabled like this:
/opt/qti-aic/exec/qaic-compile -model=<model.onnx> \
-aic-binary-dir=<./binaries> \
-stats-level=70For Hugging Face causal language models, Qualcomm maintains the efficient-transformers library (package QEfficient). It rewrites the model into a compiler-friendly form, exports it and compiles it. The quick start reads:
from QEfficient import QEFFAutoModelForCausalLM as AutoModelForCausalLM
from transformers import AutoTokenizer
model_name = "gpt2"
qeff_model = AutoModelForCausalLM.from_pretrained(model_name)
generated_qpc_path = qeff_model.compile(
num_cores=16,
mxfp6_matmul=True,
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
qeff_model.generate(prompts=["My name is"], tokenizer=tokenizer)The command-line equivalent makes the compile-time shape decisions explicit:
python -m QEfficient.cloud.infer --model_name gpt2 --batch_size 1 \
--prompt_len 32 --ctx_len 128 --mxfp6 --num_cores 16 \
--device_group [0] --prompt "My name is" --mos 1 \
--aic_enable_depth_firstRead the flags as a contract. --batch_size, --prompt_len and --ctx_len fix the tensor shapes the QPC accepts. A request longer than ctx_len cannot be served by that binary. --num_cores fixes how many AI cores the program occupies. 16 is one full Pro SoC, or one quarter of an Ultra. --device_group picks the device or devices. --mxfp6 stores matmul weights in a 6-bit microscaling (MX) format, which cuts the bytes streamed per token. The prompt (prefill) phase and the generation (decode) phase have different shapes. The library compiles a specialization for each, so one QPC holds both.
Worked example: sizing decode by bytes
Because decode is memory-bound, sizing starts from bytes, not FLOPs. Each generated token at batch size B reads every weight once, shared across the batch, plus each sequence's KV cache. Take a Llama-3-style 8B model: 32 layers, 8 KV heads, head dimension 128. Here is the arithmetic for one Ultra SoC, assuming the card's 548 GB/s splits evenly so that each SoC gets about 137 GB/s:
def decode_bound(params_b, bits_per_weight, layers, kv_heads, head_dim,
ctx, batch, bw_gbs, kv_bytes=2):
weight_bytes = params_b * 1e9 * bits_per_weight / 8
kv_per_token = 2 * layers * kv_heads * head_dim * kv_bytes # K and V
kv_bytes_step = batch * ctx * kv_per_token # read per step
step_s = (weight_bytes + kv_bytes_step) / (bw_gbs * 1e9)
return weight_bytes / 1e9, kv_per_token, batch / step_s # GB, B, tok/s
for bits in (16, 8, 6.25): # FP16, INT8, MXFP6 (6-bit + shared scale)
for batch in (1, 8, 32):
gb, kvt, tps = decode_bound(8, bits, 32, 8, 128, ctx=4096,
batch=batch, bw_gbs=137)
print(f"{bits:>5} bits batch {batch:>2}: weights {gb:5.1f} GB, "
f"upper bound {tps:7.1f} tok/s")The KV cache costs 2 x 32 x 8 x 128 x 2 bytes = 128 KiB per token in FP16, so a 4,096-token context holds 512 MiB per sequence. At batch 1 with FP16 weights (16 GB), the bound is roughly 137 / 16.5, about 8 tokens per second. MXFP6 weights (about 6.25 GB) raise it to roughly 20. At batch 32 the weights are amortised, but each step now also reads 16 GiB of KV cache, and that becomes the dominant cost. That is the shape of the trade-off on this card. Quantization buys single-stream speed. Batching buys throughput until KV reads catch up. Context length is the knob that decides when they do.
These are upper bounds. They ignore compute, inter-core traffic and scheduling gaps. Measure the real rate with your QPC and compare. If you are well under the bound, the profile (-stats-level output) will show whether DDR traffic or compute dominates. Spreading the model across four SoCs raises the aggregate bandwidth about fourfold, minus whatever the cross-SoC exchanges cost.
Failure modes
The failure modes are mostly ones a GPU user would not expect:
- Shape outside the compiled set. A prompt longer than the compiled prompt length or context length has nowhere to go. Truncate, reject, or compile more specializations. Pick the limits from your traffic's real length distribution, not from the model's maximum.
- Unsupported or slow operators. An operator the tensor unit cannot run falls to the vector or scalar units and can dominate latency. The compiler can list supported operators per framework (
-operators-supported=onnx). Check custom attention variants before committing. - FP32 left in the graph. A model exported without conversion runs far below peak. Confirm the precision of each matmul in the compiled profile.
- Quantization accuracy drift. INT8 or MXFP6 weights change outputs. Gate every compiled binary on a task-level evaluation, not just perplexity, and keep the FP16 QPC for comparison.
- Long compile times in the deploy path. Compilation is ahead of time, so a new model, context length or core count means a new build. Cache QPCs as versioned artefacts and keyed by SDK version, the same way you treat container images.
- SDK drift. Operator coverage and partitioning quality change between releases. Pin the SDK version and rerun your benchmark before upgrading.
Trade-offs and what comes next
Choose the AI 100 when power and rack density limit you, when models are known in advance and change slowly, and when you value memory capacity per watt. A single 150 W card that holds a quantized 70B model is a strong position. Avoid it when you need training, rapidly changing architectures, custom CUDA kernels, or a very high single-stream decode rate. HBM-based GPUs read memory several times faster, and the larger CUDA ecosystem means more serving engines run unchanged. The comparison with Groq's SRAM-only LPU is instructive. Groq keeps weights entirely on-chip and buys speed with many chips. Qualcomm keeps weights in cheap DRAM and buys capacity with modest bandwidth.
Qualcomm's next parts push further toward capacity. In October 2025 it announced the AI200, with 768 GB of LPDDR per card and targeted for 2026, and the AI250, targeted for 2027. Qualcomm says the AI250 uses near-memory computing for more than ten times the effective memory bandwidth. The first named deployment is a 200 MW plan with HUMAIN in Saudi Arabia. Treat performance claims for these parts as unverified until independent benchmarks exist. For costing a fleet either way, use the methods in LLM TCO, in depth and the batching economics in batching cost impact.
What to do next
- Export your target model to ONNX, or load it through QEfficient, and list any operators the compiler reports as unsupported.
- Pull prompt and output lengths from real traffic. Choose
prompt_lenandctx_lento cover the 99th percentile, and decide what happens to the rest. - Run the bandwidth calculation above for your model, precision and batch sizes. This gives you the ceiling before you benchmark.
- Compile FP16 and MXFP6 or INT8 variants. Compare task accuracy, then throughput and latency at your production batch sizes.
- Profile with
-stats-leveland check DDR traffic per core against the bound. - Version QPCs with the SDK release and model hash, and pin the SDK in deployment.
- Compare cost per million tokens against your GPU option with KV cache sizing and TCO numbers, at equal accuracy.