Qwen is not one model but a release train. Since 2023 Alibaba's Qwen team has shipped several generations of open-weight checkpoints, each in many sizes, each with instruct, base and often specialist variants, and the architecture underneath has changed more in the last eighteen months than in the two years before. A team that says it uses Qwen can mean a 0.6B dense model on a phone or a 397B mixture of experts in a data centre, and the operational answers differ completely.
This page is the family map: what each generation changed, which branches exist, how the newest models are built and why that matters for memory and runtime support, how the thinking switches work, which serving flags the model cards specify, and how to choose a checkpoint for a real workload. For a side-by-side with Phi and Gemma, including KV cache arithmetic per architecture, see Phi, Qwen and Gemma compared; for the wider small-model market, see the 2026 SLM landscape. The Qwen3.5-9B, Qwen3.6-35B-A3B, Qwen3.6-27B, Qwen3-8B and Qwen2.5 3B and 72B cards on Hugging Face were checked on 2026-10-02; the other rows summarise release announcements, so confirm them on the card of the exact checkpoint before relying on them.
The generations at a glance
Each generation kept the previous one's chat format and tokenizer family where it could, so most migration pain comes from architecture and template details rather than from the API. The table lists the releases that matter for current decisions.
| Generation | When | Open sizes | What changed | Licence |
|---|---|---|---|---|
| Qwen2.5 | September 2024 | 0.5B, 1.5B, 3B, 7B, 14B, 32B, 72B dense | Mature dense decoder, ChatML-style template, strong multilingual and coding base | Apache 2.0 for most sizes; the 3B card lists qwen-research and the 72B card lists qwen |
| Qwen3 | April 2025 | 0.6B, 1.7B, 4B, 8B, 14B, 32B dense; 30B-A3B and 235B-A22B MoE | Thinking and non-thinking in one checkpoint, switched per request | Apache 2.0 |
| Qwen3 2507 updates | July 2025 onward | Selected sizes re-released as separate Instruct and Thinking checkpoints | Split the hybrid back into two models for better quality in each mode | Apache 2.0 |
| Qwen3-Next | September 2025 | 80B-A3B | First hybrid: Gated DeltaNet linear attention mixed with gated full attention, sparse MoE | Apache 2.0 |
| Qwen3.5 | February to March 2026 | 0.8B, 2B, 4B, 9B, 27B dense; 35B-A3B, 122B-A10B, 397B-A17B MoE | Hybrid layout across the whole range, native vision input, 262,144-token native context | Apache 2.0 |
| Qwen3.6 | April 2026 | 35B-A3B MoE and 27B dense open; flagship Max-Preview API-only | Agentic coding focus, an option to preserve earlier reasoning in multi-turn chats | Apache 2.0 for the open checkpoints |
Two notes on reading the table. The suffix A3B, A10B and so on is the number of parameters active per token in a mixture-of-experts model; the first number is what you must load. And release dates for Qwen3.6 vary by a couple of weeks across secondary sources, so only the month is given. Whichever generation you pick, the licence that binds you is the LICENSE file in the exact repository you download, including community quantizations.
The specialist branches
Alongside the general models, the team publishes branches tuned for one job. They share the tokenizer family but not always the template, context length or licence, so treat each as a separate product with its own card.
- Coder: code completion and agentic coding models, including fill-in-the-middle support in the Qwen2.5-Coder line and large agentic coders in the Qwen3-Coder line. The tool-call parser named on recent cards,
qwen3_coder, comes from this branch and is now used for the general models too. - VL and Omni: vision-language and any-to-any models. From Qwen3.5 the general checkpoints accept images and video natively, so a separate VL model is only worth it when its card shows a capability you need.
- Embedding and Reranker: retrieval models in the Qwen3 generation. They are encoders for RAG pipelines, not chat models, and they need their own instruction prefixes as documented on the card.
- Math and earlier reasoning previews: before Qwen3 merged thinking into the main line, reasoning came from separate models such as QwQ. New projects should start from a current general checkpoint.
The practical rule: start with the general model of the newest generation your runtime supports, and move to a branch only after an evaluation on your own data shows a gap the branch closes.
Inside the newest models: dense, MoE and hybrid layers
Up to Qwen3 every layer was a standard transformer block: grouped-query attention followed by a feed-forward network. Attention keeps a key and value vector per token per layer, so memory grows linearly with context, and at long contexts the KV cache, not the weights, decides how many requests fit on a GPU.
Qwen3-Next and then the whole Qwen3.5 range replaced most attention layers with Gated DeltaNet, a linear-attention layer that keeps a fixed-size recurrent state per sequence instead of a cache that grows. The Qwen3.5-9B card describes 32 layers laid out as eight repeats of three Gated DeltaNet layers followed by one gated attention layer, a hidden size of 4,096, and a padded vocabulary of 248,320 tokens, much larger than the roughly 152K of Qwen2.5 and Qwen3. Only the eight full-attention layers pay per-token KV memory, using 4 KV heads of dimension 256.
The DeltaNet state still costs memory, but a fixed amount: on the order of a 128 by 128 matrix per value head per layer. Its exact size depends on head counts and on the dtype your runtime keeps it in, so measure it rather than trusting a formula. The trade is that long-context serving gets much cheaper, while features built around a growing KV cache, such as prefix caching and paged memory, need engine support written for recurrent state.
The MoE models add a second axis. Qwen3.6-35B-A3B has 40 layers in the same three-to-one pattern, and every layer's feed-forward block is a mixture of 256 experts with 8 routed plus 1 shared expert active per token. Compute per token tracks the 3B active parameters, so decode is fast, but all 35B parameters must be resident, so memory planning starts from the total, not the active count. See mixture of experts for small models for routing and memory trade-offs in general.
Thinking modes and sampling settings
Every recent Qwen generation can write a reasoning block wrapped in think tags before the answer. How you control it changed between generations, and getting it wrong either wastes tokens or breaks parsers.
| Generation | Default | How to turn thinking off |
|---|---|---|
| Qwen3 (original release) | Thinking on | enable_thinking=False in the chat template; the cards also document /think and /no_think soft switches in user turns |
| Qwen3 2507 checkpoints | Fixed per checkpoint | Pick the Instruct or the Thinking checkpoint |
| Qwen3.5 and Qwen3.6 | Thinking on | Pass chat_template_kwargs={"enable_thinking": False} to vLLM or SGLang |
Do not assume the Qwen3 soft switches work on Qwen3.5; the 3.5 card documents the template argument, so use that. Qwen3.6 adds a preserve_thinking option that keeps reasoning from earlier turns in the history, which helps iterative agent work at the cost of a longer prompt.
Sampling matters more than usual. The Qwen3 cards warn against greedy decoding in thinking mode because it can cause endless repetition, and recommend temperature 0.6, top-p 0.95 and top-k 20. The Qwen3.5 card gives separate presets; for non-thinking general tasks it lists temperature 0.7, top-p 0.8, top-k 20 and a presence penalty of 1.5, and for thinking-mode general tasks temperature 1.0 with top-p 0.95. Start from the card's preset for your mode and change one value at a time.
Serving and tool calling
The Qwen3.5 card states that Qwen3.5 needed vLLM and SGLang builds from their main branches at release, and the latest Transformers. In practice that means pinning a runtime version you have tested with the exact checkpoint rather than assuming any recent build works. The Qwen3.5-9B card gives these flags for tool use and reasoning output (its own example uses the full 262,144-token length):
vllm serve Qwen/Qwen3.5-9B \
--max-model-len 32768 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3The reasoning parser separates the think block from the answer so clients get clean content. From an OpenAI-compatible client, turn thinking off per request for structured extraction:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
resp = client.chat.completions.create(
model="Qwen/Qwen3.5-9B",
messages=[
{"role": "system", "content": "Classify the ticket. Reply with JSON only."},
{"role": "user", "content": "Refund not received after 14 days."},
],
temperature=0.7, top_p=0.8,
extra_body={"top_k": 20, "chat_template_kwargs": {"enable_thinking": False}},
)
print(resp.choices[0].message.content)Set --max-model-len to what you need, not the 262,144-token maximum; the engine reserves memory for the configured length. For local and edge runtimes, check that your build implements the hybrid layers before downloading a GGUF of a 3.5 or 3.6 model; the GGUF and llama.cpp guide covers how to verify what a file contains. For prompting tool use on small models generally, see SLMs for tool calling.
Worked example: choosing a Qwen for ticket triage
A support team wants a self-hosted model to classify tickets, extract an order number and call one lookup tool. Constraints: one 24 GB GPU, prompts under 6,000 tokens, JSON output, English and Chinese tickets, a p95 latency target of two seconds.
Step one is the memory shortlist, from parameter count times bytes per parameter. Qwen3.5-9B at 16-bit is about 18 GB of weights, which leaves too little room for batching on 24 GB. At 8-bit it is about 9 GB, and at 4-bit roughly 5 to 6 GB. Qwen3.5-4B is about 8 GB at 16-bit. Qwen3.6-35B-A3B needs about 19 to 20 GB even at 4-bit, which fits only with almost no headroom, so it drops out for this GPU despite its fast decode.
Step two is the mode decision. Classification and extraction do not need a reasoning block, and a think block before the JSON breaks a strict parser, so thinking is off for every candidate. Step three is a bake-off on 300 labelled tickets: 4B at 16-bit, 9B at 8-bit and 9B at 4-bit, each served with the same vLLM version and the card's non-thinking preset, scoring label accuracy, JSON validity, correct tool-call arguments and p95 latency at the expected concurrency.
The decision rule is written before the run: choose the smallest configuration within one point of the best accuracy that meets the latency target and produces valid JSON on at least 99.5% of tickets. If none does, add constrained decoding with a JSON schema before moving up a size. Whatever wins, record the checkpoint revision, runtime version, quantization and sampling settings, because changing any of them invalidates the result.
Failure modes
- Think blocks in structured output: thinking is on by default in Qwen3, 3.5 and 3.6, so a JSON parser sees a think tag first and fails.
- Greedy decoding loops: temperature 0 in thinking mode produces repetition that runs until the token limit.
- Wrong tool parser: tool calls arrive as plain text in the content field because the server was started without the parser the card names.
- Runtime without hybrid support: an older engine refuses the checkpoint or, worse, a community conversion runs with silently wrong outputs.
- Memory planned from active parameters: an A3B model sized as if it were 3B goes out of memory at load.
- Licence drift: a community quantization or an older-generation size carries different terms from the model you evaluated.
- Context budget waste: serving with the maximum model length reserves memory for contexts no request uses, cutting concurrency.
Trade-offs
Newer generations give better quality per parameter and far cheaper long contexts, but demand newer runtimes and have less tooling history. Dense models are predictable in memory and supported everywhere; MoE models trade memory for decode speed. A single hybrid-thinking checkpoint is convenient, while split checkpoints tend to be better in each mode. Choose by the constraint that binds you first: memory, runtime support, latency, or licence review.
What to do next
- Write down your hardware, context length, latency target and output format before looking at any model card.
- Shortlist two or three Qwen checkpoints by total parameters times bytes per parameter, leaving headroom for the KV cache and batching.
- Confirm your runtime version supports the generation you picked, and pin it.
- Decide thinking on or off per endpoint, and set the card's sampling preset for that mode.
- Start the server with the tool-call and reasoning parsers the card names, and test one tool call end to end.
- Run a bake-off on a few hundred of your own examples with a decision rule written in advance.
- Record checkpoint revision, quantization, runtime and sampling settings next to the result, and read the LICENSE in the exact repository you deploy.