Nemotron is NVIDIA's family of open models, and the name now covers several quite different architectures released since 2024: a dense 340B model built to generate synthetic training data, pruned and distilled Llama derivatives tuned for reasoning, and a line of hybrid Mamba-Transformer models that, from Nemotron 3 onward, add mixture of experts. What ties them together is that each is designed around how NVIDIA GPUs and software actually execute models: shrink the KV cache, keep active compute small, train in low precision that the tensor cores accelerate, and publish the data and recipes so others can reproduce the pipeline on the same stack.

This article is for engineers deciding whether and how to run or fine-tune a Nemotron model. It sorts out the lineage with verified numbers, explains the hybrid layer stack from its published configuration, works through KV cache and state memory for Nemotron 3 Nano, covers FP8 and NVFP4 training and serving, the prune-and-distill recipe, a serving command, failure modes and trade-offs. Model details change quickly; figures here were checked against NVIDIA's technical reports and model cards as of October 2026.

The lineage, with verified numbers

GenerationReleasedWhat is distinctive
Nemotron-4 340B (Base, Instruct, Reward)June 2024Dense; pretrained on 9T tokens; over 98 percent of alignment data synthetic; reward model published for data filtering
Llama-Nemotron Nano 8B, Super 49B, Ultra 253B2025Derived from Llama 3 models; Super and Ultra reshaped with Puzzle architecture search; reasoning toggled by the system prompt detailed thinking on or off
Nemotron-H 8B, 56B, 47BApril 2025Hybrid Mamba-Transformer; 56B pretrained on 20T tokens in FP8; 47B compressed from 56B with MiniPuzzle
Nemotron Nano 2 (9B)August 202512B hybrid base trained on 20T tokens in FP8, then pruned and distilled with Minitron to fit 128k context on one A10G
Nemotron 3 Nano 30B-A3BDecember 2025Hybrid Mamba-Transformer with MoE; 31.6B total, 3.2B active; 25T pretraining tokens; NVFP4 checkpoint via quantization-aware distillation
Nemotron 3 Super 120B-A12BMarch 2026120.6B total, 12.7B active; LatentMoE, multi-token prediction; pretrained mostly in NVFP4 on over 25T tokens

Two practical points follow. First, Nemotron is not one tokenizer, one chat template or one license, so read the model card of the exact checkpoint. Second, the family has moved from dense Transformers to hybrids; serving code, kernels and capacity plans written for Nemotron-4 or Llama-Nemotron do not transfer to Nemotron-H or Nemotron 3 unchanged. A largest member, Nemotron 3 Ultra, is listed by NVIDIA's Nemotron 3 page as released; its specifications are not covered here, so read its model card before planning around it.

The hybrid layer stack

Nemotron 3 Nano 30B-A3B: 52 layers from the published hybrid_override_patternMEMEM*EMEMEM*EMEMEM*EMEMEM*EMEMEM*EMEMEMEM*EMEMEMEMEM: Mamba-2 (23)fixed-size state per sequenceE: MoE FFN (23)128 experts, 6 routed + 2 shared*: GQA attention (6)32 query heads, 2 KV headsOnly the 6 attention layers grow a KV cache with context length.31.6B total parameters must sit in GPU memory; about 3.2B are active per token.
Each cell is a layer, in order. The pattern string is copied from the model's published config.json.

A Transformer layer's attention keeps every past key and value, so its memory grows linearly with context and each decode step reads all of it. A Mamba-2 layer is a selective state space model: it compresses the past into a fixed-size state and updates it per token, so its memory and per-token compute do not grow with context. The math is covered in Mamba math. Pure SSMs are weaker at exact recall from far back in the context, so the Nemotron hybrids keep a few attention layers spread through the depth to provide that.

Nemotron 3 Nano's configuration makes the balance concrete: 52 layers, of which 23 are Mamba-2, 23 are MoE feed-forward layers and only 6 are grouped-query attention, each with 32 query heads and 2 key-value heads of dimension 128. Each MoE layer has 128 routed experts with 6 active per token plus 2 shared experts, which is how 31.6B parameters turn into roughly 3.2B active per token. The routing mechanics are covered in MoE routing on GPUs.

Memory arithmetic: KV cache and SSM state

The payoff shows up in memory arithmetic, which you can do from the config alone. KV cache bytes per token are layers with attention times 2 (keys and values) times KV heads times head dimension times bytes per element.

def kv_bytes_per_token(attn_layers, kv_heads, head_dim, bytes_per=2):
    return attn_layers * 2 * kv_heads * head_dim * bytes_per

def ssm_state_bytes(mamba_layers, heads, head_dim, state_dim, bytes_per=2):
    # recurrent state only; the small convolution state is ignored here
    return mamba_layers * heads * head_dim * state_dim * bytes_per

nano3 = kv_bytes_per_token(6, 2, 128)            # 6,144 bytes per token
print(nano3 * 131072 / 1e9)                      # ~0.81 GB at 128k tokens
state = ssm_state_bytes(23, 64, 64, 128)         # ~24 MB per sequence, any length
print(state / 1e6, state / nano3)                # crossover near 3,900 tokens

# illustrative dense model: 48 attention layers, 8 KV heads, dim 128
dense = kv_bytes_per_token(48, 8, 128)           # 196,608 bytes per token
print(dense * 131072 / 1e9)                      # ~25.8 GB at 128k tokens

Worked example. One 128k-token sequence on Nemotron 3 Nano needs about 0.81 GB of KV cache in BF16, plus about 24 MB of Mamba state (64 heads of dimension 64 with state size 128 across 23 layers; double it if your engine keeps state in FP32). A hypothetical dense model with 48 attention layers and 8 KV heads would need about 25.8 GB for the same sequence, more than 30 times as much. Below roughly 3,900 tokens the fixed Mamba state is actually the larger term, so the hybrid's advantage is a long-context and high-concurrency advantage; for short chat turns the weights dominate anyway.

Weights are the other half of the budget. All 31.6B parameters must be resident even though only about 3.2B are used per token: roughly 63 GB in BF16, about half that in FP8 and about a quarter plus scale overhead in NVFP4. That is why the MoE model is cheap per token but not small. Plug both numbers into the method in GPU serving capacity estimation to get a request rate per replica. NVIDIA reports that on an 8k input, 16k output workload Nemotron 3 Nano delivers 3.3 times the throughput of Qwen3-30B-A3B-Thinking-2507 and 2.2 times that of GPT-OSS-20B; treat those as vendor figures and measure on your own traffic.

Low-precision training: FP8 and NVFP4

Nemotron is also NVIDIA's showcase for low-precision training. Nemotron-H 56B was pretrained on 20T tokens in FP8 using per-tensor current scaling: each tensor is quantized with a single scale computed from its present values rather than a delayed history, which is coarse but simple and was stable at that scale. Nemotron 3 Super went further and pretrained most linear layers in NVFP4 for weights, activations and gradients, while keeping sensitive pieces, including latent projections, multi-token prediction layers, attention projections and embeddings, in BF16 or MXFP8.

NVFP4 stores each value as a 4-bit E2M1 float, with an FP8 (E4M3) scale shared by every block of 16 values and a per-tensor FP32 scale on top. The small blocks let it track local magnitude far better than a single scale. Native FP4 tensor core support arrived with Blackwell, so NVFP4 throughput claims assume B200 or GB200 class hardware; on Hopper you will run FP8 or BF16. Super is the first Nemotron 3 model pretrained in NVFP4; Nano's NVFP4 checkpoint was produced after training by quantization-aware distillation, in which the quantized student is trained to match the full-precision model's outputs.

Pruning and distillation

Several Nemotron models are not trained at their final size. Minitron-style compression trains a larger model, scores the importance of layers, attention heads, MLP channels and embedding channels from activations on a small calibration set, removes the least important ones, and then recovers accuracy by distilling from the original model. Nemotron Nano 2 went from a 12B base to 9B this way, and Nemotron-H 47B came from 56B with the related MiniPuzzle method. The loop looks like this:

# Sketch of activation-based width pruning followed by logit distillation
importance = {}
for batch in calibration_loader:                 # a few thousand samples
    acts = teacher.forward_with_hooks(batch)     # per-layer activations
    for name, a in acts.items():
        importance[name] = importance.get(name, 0) + a.abs().mean(dim=(0, 1))

student = prune_channels(teacher, importance, keep_ratio=0.75)

for batch in distill_loader:                     # tens to hundreds of billions of tokens
    with torch.no_grad():
        t_logits = teacher(batch.tokens)
    s_logits = student(batch.tokens)
    loss = kl_div(log_softmax(s_logits / T), softmax(t_logits / T)) * T * T
    loss.backward(); optimizer.step(); optimizer.zero_grad()

The economics are the point: distillation from a strong teacher needs far fewer tokens than training the small model from scratch. The same idea, generating training signal from a stronger model, runs through the family, from Nemotron-4's synthetic alignment data to the open Nemotron pretraining and post-training datasets. See distillation data on GPUs for the data side.

Serving Nemotron 3 Nano

NVIDIA's own throughput measurements for Nemotron 3 Nano used vLLM and TensorRT-LLM; for other engines, check release notes for the exact checkpoint. The vLLM project's published recipe sets the FlashInfer attention backend and enables the reasoning and tool-call parsers; the BF16 and FP8 checkpoints use the same flags:

export VLLM_ATTENTION_BACKEND=FLASHINFER
vllm serve nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \
    --trust-remote-code \
    --served-model-name nemotron \
    --port 5000 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --reasoning-parser deepseek_r1
from openai import OpenAI
client = OpenAI(base_url="http://localhost:5000/v1", api_key="unused")
r = client.chat.completions.create(
    model="nemotron",
    messages=[{"role": "user", "content": "Explain grouped-query attention in 3 sentences."}],
    max_tokens=2048,
)
print(r.choices[0].message.content)   # the reasoning parser moves the trace to a separate field

Reasoning control differs by generation. Llama-Nemotron models switch modes with the system prompt detailed thinking on or off. Nemotron 3 Nano supports reasoning on and off and was trained with truncated reasoning traces so that a thinking budget can cap reasoning length; use the chat template options documented on its model card rather than the older system-prompt switch. Budget reasoning explicitly: an 8k-token trace costs the same decode time as an 8k-token answer, and it is invisible to users who only see the final content.

Failure modes

  • Engine lacks hybrid kernels. Older inference stacks without Mamba-2 or hybrid cache support fail to load the model or fall back to slow paths. Pin a release that lists the exact checkpoint as supported.
  • Capacity planned like a dense model. Sizing KV for 52 attention layers wastes most of the GPU; forgetting the fixed per-sequence Mamba state overcommits short requests. Use both terms.
  • NVFP4 on the wrong GPU. FP4 checkpoints either will not run or will be emulated on pre-Blackwell hardware. Choose FP8 on Hopper.
  • Context length mismatch. NVIDIA advertises up to 1M tokens for Nemotron 3 Nano, but the published config's max_position_embeddings is 262,144. Check what your engine actually allows and test accuracy at the length you plan to use.
  • Unbounded reasoning. Reasoning mode left on for simple traffic multiplies latency and cost. Route easy requests to reasoning off or set a budget.
  • trust-remote-code drift. Remote model code changes with repository updates. Pin the model revision as well as the engine version.

Trade-offs

ChoiceGainCost
Hybrid Mamba over dense attentionSmall KV cache, steady long-context decodeNewer kernels, less tooling maturity
MoE (3.2B active of 31.6B)Low compute per tokenAll experts resident; routing imbalance at scale
NVFP4 servingHighest throughput on BlackwellBlackwell only; validate accuracy on your tasks
Pruned and distilled sizesFits smaller GPUs cheaplySome capability lost versus the parent
Reasoning onBetter multi-step accuracyMany more output tokens per request

What to do next

  1. Pick the generation that matches your hardware: FP8 or BF16 Nemotron 3 Nano on Hopper, NVFP4 on Blackwell, Nano 2 for a single small GPU.
  2. Read the exact checkpoint's config.json and recompute KV and state memory with the functions above for your context lengths.
  3. Serve it with the pinned vLLM command, then load-test with your real prompt and output length distribution.
  4. Measure quality with reasoning off, on and budgeted, and route traffic by difficulty.
  5. If you fine-tune, start from NVIDIA's NeMo recipes for the checkpoint and keep the high-precision layers the reports identify.
  6. Re-check the model cards monthly; the family is still adding sizes and checkpoints.
Key takeaway: Nemotron is a family, not a model: recent members combine a few attention layers with many Mamba-2 and MoE layers, so KV cache stays small while total weights stay large. Plan capacity from the config, using both KV and fixed SSM state, match precision to your GPU generation, pin engine and model revisions, and budget reasoning tokens explicitly.