Gemma is Google's family of open-weight models, built from the same research as Gemini and released in sizes that run on a laptop, a single GPU or a phone. Since February 2024 it has grown from two text models into a family of general checkpoints and specialised variants for code, vision, safety classification, embeddings and medicine. Gemma 4, released in 2026, moved the main models to the Apache 2.0 licence and added mixture-of-experts and audio-capable edge models.
This page explains the family well enough to deploy and adapt it: what was released when, what the Gemma 4 checkpoints are, how the architecture decides memory use, with numbers computed from the published config files, how prompting and thinking work, how to run and fine-tune the models, and which specialised variant fits which job. Comparing Gemma against Phi and Qwen is a separate question, covered in Phi, Qwen, Gemma: small models compared. Facts here were checked against Google's Gemma documentation and the Hugging Face model repositories on 2026-10-02; the family releases often, so check again before you depend on a detail.
The lineage
Each generation changed something that matters to deployment, so it helps to know the sequence. Dates are from Google's Gemma release notes.
| Date | Release | What changed |
|---|---|---|
| Feb 2024 | Gemma 2B and 7B | First open Gemma text models; Gemma 1.1 followed in April |
| Apr to May 2024 | CodeGemma, RecurrentGemma, PaliGemma | Code completion, a recurrent architecture, and vision-language |
| Jun to Jul 2024 | Gemma 2: 2B, 9B, 27B; ShieldGemma | Alternating local and global attention layers; safety classifiers |
| Mar 2025 | Gemma 3: 1B, 4B, 12B, 27B; ShieldGemma 2 | Image input on 4B and up, 128K context on the larger sizes, five local layers per global layer |
| May to Sep 2025 | MedGemma, Gemma 3n, Gemma 3 270M, EmbeddingGemma 308M | Medical variants, on-device models with per-layer embeddings, a tiny fine-tuning base, an embedding model |
| Dec 2025 | FunctionGemma, T5Gemma 2 | Function-calling specialisation and encoder-decoder models |
| Mar 31, 2026 | Gemma 4: E2B, E4B, 26B A4B, 31B | Apache 2.0, a mixture-of-experts size, 256K context on large sizes, audio on edge sizes |
| Jun 3, 2026 | Gemma 4 12B Unified | A 12B model that projects image patches and audio directly into the language model |
Two consequences follow. Licence review depends on the generation: Gemma 3 and earlier shipped under Google's Gemma terms, Gemma 4 under Apache 2.0, and derivatives inherit their base's terms. And each generation introduced layer types that engines had to implement, so a runtime that served Gemma 3 well is not automatically correct for Gemma 4.
The Gemma 4 checkpoints
| Checkpoint | Layers | Context | Inputs | Notes from the config |
|---|---|---|---|---|
| E2B | 35 | 128K | text, image, audio | 2.3B effective, 5.1B with embeddings; 512-token sliding window; 1 KV head; last 20 layers reuse earlier KV |
| E4B | 42 | 128K | text, image, audio | 4.5B effective, 8B with embeddings; per-layer embeddings like E2B |
| 12B Unified | 48 | 256K | text, image, audio | No separate vision or audio encoder |
| 26B A4B | 30 | 256K | text, image | 128 experts, 8 routed per token, about 3.8B active of about 25B |
| 31B | 60 | 256K | text, image | Dense; 1,024-token sliding window; 262,144-token vocabulary |
Every size ships as a pre-trained base and an instruction-tuned -it checkpoint. Use the base model only as a starting point for your own fine-tuning; for prompting, use -it. The vocabulary of 262,144 tokens is shared across the family, which helps multilingual text and makes the embedding and output layers large relative to the small models' compute.
Hybrid attention sets the memory bill
Gemma 2 alternated sliding-window and global layers; from Gemma 3 on, five of every six layers are sliding. A sliding layer attends only to the last few hundred or thousand tokens; a global layer attends to the whole context. In the Gemma 4 31B config the layer_types list is five sliding_attention layers followed by one full_attention layer, repeated ten times, so the last layer is global. The sliding layers use a 1,024-token window with 16 key-value heads of 256 dimensions; the global layers use 4 key-value heads of 512 dimensions.
This matters because the KV cache, not the weights, is what grows with context. A sliding layer never needs to store more than its window, so its cache is fixed once the conversation passes 1,024 tokens. Only the global layers grow with every token. Working it out at bf16, two bytes per value, for a 131,072-token context:
- Sliding layers: 2 (K and V) x 16 heads x 256 dims x 2 bytes = 16 KiB per token per layer, capped at 1,024 tokens = 16 MiB per layer, 800 MiB for all 50.
- Global layers: 2 x 4 x 512 x 2 bytes = 8 KiB per token per layer, 80 KiB per token for all 10, so 10 GiB at 131,072 tokens and 20 GiB at the full 262,144.
The total, about 11 GiB, is an upper bound. The config also sets attention_k_eq_v for the 31B and 26B models, a flag that suggests global layers can share keys and values; whether your engine stores less because of it is something to measure, not assume. The opposite risk is larger: an engine that ignores the sliding window and allocates a full-length cache for every layer would need about 100 GiB for the 50 sliding layers alone. The script below reads the config itself. KV cache mechanics in general are in SLM KV cache.
import json, urllib.request
def kv_bytes(cfg_url, context, bytes_per_value=2):
cfg = json.load(urllib.request.urlopen(cfg_url))["text_config"]
total = 0
for kind in cfg["layer_types"]:
if kind == "sliding_attention":
heads, dim = cfg["num_key_value_heads"], cfg["head_dim"]
tokens = min(context, cfg["sliding_window"])
else: # full_attention
heads = cfg.get("num_global_key_value_heads") or cfg["num_key_value_heads"]
dim = cfg.get("global_head_dim") or cfg["head_dim"]
tokens = context
total += 2 * heads * dim * tokens * bytes_per_value # K and V
return total # upper bound: ignores KV sharing and k_eq_v savings
url = "https://huggingface.co/google/gemma-4-31B-it/raw/main/config.json"
for ctx in (8_192, 32_768, 131_072):
print(ctx, round(kv_bytes(url, ctx) / 2**30, 2), "GiB")
Edge models: per-layer embeddings and KV sharing
E2B and E4B are designed for phones and laptops, and their configs show three memory-saving ideas. First, per-layer embeddings: besides the normal token embedding, each layer receives a small per-token vector looked up from a large table (hidden_size_per_layer_input is 256 in E2B). Those tables account for the gap between 2.3 billion effective and 5.1 billion total parameters. They are lookups, so they add little compute, and a runtime built for it can keep them outside accelerator memory. A generic runtime loads everything, so budget for the larger number until measured.
Second, KV sharing: num_kv_shared_layers is 20 in E2B, meaning the last 20 of its 35 layers reuse key-value tensors computed by earlier layers instead of storing their own. Third, E2B uses a single key-value head and a 512-token sliding window. Together these make its KV cache small even at long context, which is what lets a 128K window be usable on device at all. On-device deployment budgets are covered in SLMs on mobile in 2026.
The 26B A4B mixture of experts
The 26B A4B model replaces most dense feed-forward blocks with a mixture of experts. Its config lists 128 experts per layer with 8 routed per token, plus a shared expert according to the model card, and 30 layers. Roughly 3.8 billion parameters are active for each token, so compute per token is close to a 4B dense model, but all of the roughly 25 billion parameters must be loaded, because different tokens route to different experts.
Plan it accordingly. Memory is sized like a 25B model: about 50 GB at bf16 or roughly 14 to 16 GB at 4-bit, before KV cache. Single-request decoding is fast because each token touches only a fraction of the weights, while heavy batching tends to touch most experts anyway. If memory, not compute, is your limit, a smaller size may fit better.
Prompt format and thinking
Gemma 4 has its own turn markers: a turn opens with <|turn> plus a role of system, user or model, and closes with <turn|>. Gemma 4 accepts a system role. Thinking is switched on by putting <|think|> in the system instruction, and the model then writes its reasoning between <|channel>thought and <channel|> before the answer. Tools use their own pairs: <|tool> for declarations, <|tool_call> for calls and <|tool_response> for results.
<|turn>system
<|think|>You are a careful assistant.<turn|>
<|turn>user
Which layers in this model keep the full context?<turn|>
<|turn>model
<|channel>thought
...model reasoning...<channel|>The global attention layers do.<turn|>Two rules from Google's formatting guide prevent most mistakes. Never build these strings by hand; use the chat template shipped with the checkpoint. And in a normal multi-turn conversation, strip the model's thoughts from previous turns before sending history back, except between tool calls inside one turn, where they must stay. For extraction or classification, leave thinking off: it adds output tokens and latency and can break a parser expecting JSON first.
Running the models
With Hugging Face transformers, Gemma 4 loads through the processor and the image-text-to-text model class even for text-only use, because the checkpoints are multimodal. The config was written by a recent transformers version, so upgrade first and copy the snippet from the model card if class names differ in your version.
from transformers import AutoProcessor, AutoModelForImageTextToText
import torch
model_id = "google/gemma-4-E4B-it"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto")
messages = [
{"role": "system", "content": [{"type": "text", "text": "Answer in one sentence."}]},
{"role": "user", "content": [{"type": "text", "text": "What is a sliding-window attention layer?"}]},
]
inputs = processor.apply_chat_template(messages, add_generation_prompt=True,
tokenize=True, return_dict=True,
return_tensors="pt").to(model.device)
print(processor.decode(inputs["input_ids"][0])) # inspect the real prompt once
out = model.generate(**inputs, max_new_tokens=128)
print(processor.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))For local or edge use, check that your engine names Gemma 4 explicitly and matches the transformers reference on a few prompts; sliding windows, per-layer embeddings and audio towers each need explicit support. Evaluate at the precision you ship; SLM quantization covers the methods.
Fine-tuning with LoRA
Most teams adapt Gemma with LoRA on the instruction-tuned checkpoint. Start from -it unless you have a large dataset and your own instruction-tuning recipe. Apply adapters to the attention and MLP projections of the language model, confirm the module names exist in the checkpoint, and render every example through the chat template so training and serving see the same markers.
from peft import LoraConfig, get_peft_model
# Find the language-model prefix first; the vision and audio towers
# also contain modules named q_proj, so bare names would adapt them too.
for name, _ in model.named_modules():
if name.endswith("q_proj"):
print(name)
LM = r".*language_model.*" # replace with the prefix printed above
lora = LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05,
target_modules=LM + r"\.(q_proj|k_proj|v_proj|o_proj|gate_proj|up_proj|down_proj)",
task_type="CAUSAL_LM",
)
model = get_peft_model(model, lora)
model.print_trainable_parameters()
# Build every training example with processor.apply_chat_template so the
# fine-tuned model sees the same <|turn> markers it will see in production,
# and compute the loss only on the model turns.Three Gemma-specific points. For text-only training, make sure target modules select language-model layers, not the vision tower. On 26B A4B, start with attention projections only, since adapting experts multiplies adapter size. And train with the same thinking setting you serve with. General LoRA practice is in LoRA for SLMs.
Specialised variants
| Variant | Use it for |
|---|---|
| ShieldGemma | Classifying prompts or responses against safety policies, as a guard model in front of or behind a main model |
| EmbeddingGemma (308M) | Text embeddings for retrieval and clustering on small hardware |
| CodeGemma | Code completion and fill-in-the-middle; check whether a newer general model now does your task better |
| PaliGemma | Vision-language tasks such as captioning and visual question answering, usually after fine-tuning |
| MedGemma | Medical text and image starting points; needs your own validation and governance |
| FunctionGemma | Function-calling workloads that need a small specialised model |
| Gemma 3 270M | A tiny base for narrow fine-tuned tasks such as classification or extraction |
Several variants are built on earlier generations, so check each one's licence and runtime support.
Failure modes
- Effective parameters mistaken for memory. Budgeting E2B as a 2B model and finding the loaded size near 5B. Budget total parameters until you have measured.
- Hybrid attention ignored. An engine without sliding-window support allocates a full cache per layer and runs out of memory at modest context. Verify support and measure KV at your context.
- Hand-built prompts. Missing or wrong turn markers degrade answers quietly. Use the processor's chat template and inspect the rendered prompt once.
- Thoughts left in history. Prior-turn reasoning wastes context and departs from the format the model was trained on. Strip it except between tool calls in the same turn.
- Wrong licence assumption. Treating a Gemma 3-based variant as Apache 2.0. Read the licence file in the exact repository you download.
- Mismatched fine-tune format. Training on plain text and serving with the chat template. Render training data with the same template.
What to do next
- Pick the smallest Gemma 4 checkpoint that accepts your input modalities, then confirm your runtime lists that exact checkpoint as supported.
- Run the KV script against its config at your real context length and add the weights; compare with measured peak memory.
- Render one prompt through the chat template and read it, including how thinking and tools appear.
- Decide thinking on or off per task and measure output tokens and latency both ways.
- Evaluate the quantized model you will ship on your own data, not the bf16 original.
- If quality falls short, run a LoRA fine-tune on the -it checkpoint with chat-templated examples and loss on model turns only.
- Check the licence file of every variant, fine-tune and quantization you deploy.