Search for how to quantize a language model and you will find the same three names in thousands of tutorials and model cards: bitsandbytes, AutoAWQ and AutoGPTQ. They are usually presented as interchangeable options on a menu. They are not. One quantizes weights as the model loads, with no data, and is built around fine-tuning; the other two ran calibration algorithms to produce packed checkpoints for fast serving. And in 2026 two of the three are no longer maintained: AutoGPTQ was archived in April 2025 and AutoAWQ in May 2025, with their work continuing in other projects.

This article does not re-derive the algorithms; AWQ architecture and GPTQ architecture do that. Instead it answers the questions an engineer faces when choosing a tool: what each one produces, which runtimes can load the result, what you lose and gain, and how to move a project off an archived library without breaking production.

Advertisement

Two kinds of quantization

The first distinction matters more than any benchmark. On-the-fly quantization takes ordinary BF16 weights and converts them to a low-bit format as they are loaded, using only the weights themselves. There is no calibration data and no separate quantization step, so it works on any model the library supports, immediately. bitsandbytes works this way.

Calibrated quantization runs sample inputs through the model, observes activations, and uses those statistics to choose how to round each weight so the layer's output changes as little as possible. AWQ protects the weight channels that see large activations by scaling them; GPTQ quantizes columns one at a time and pushes each rounding error onto the columns not yet quantized, using second-order information. Both take minutes to hours, need representative data, and write a new checkpoint whose packed layout matches fast inference kernels. AutoAWQ and AutoGPTQ were the popular implementations of these two methods.

The consequence is a rough division of labour. On-the-fly suits experiments and training, where you want to load a model in less memory now. Calibrated suits serving, where you quantize once and run the result millions of times.

BF16 checkpointfrom the hub or your fine-tunebitsandbytesquantize at load, no dataAWQ (llm-compressor)calibrate: activation scalesGPTQ (GPTQModel orllm-compressor): Hessian error feedbackNF4 / FP4 / LLM.int8bnb tensors in memorycompressed-tensorspacked INT4 + scalesGPTQ formatpacked INT4 + scales (+ g_idx)QLoRA trainingfrozen 4-bit base + adaptersHigh-throughput servingvLLM / SGLang with Marlin-class INT4 kernelsSame goal, different contracts: bitsandbytes needs no data and fits training; AWQ and GPTQ need calibration and fit serving
Figure 1. Three routes from one BF16 checkpoint. The route decides the artefact you get and the runtimes that can use it.

Where each project stands in 2026

ProjectStatusWhat to use now
bitsandbytesActively maintained; multi-backend support covers NVIDIA CUDA, AMD ROCm, Intel XPU and CPU backends, with an Apple Silicon backend in recent releasesbitsandbytes itself, through the Transformers BitsAndBytesConfig
AutoAWQArchived in May 2025; its README says the functionality was adopted by the vLLM project's llm-compressorllm-compressor's AWQ modifier; existing AutoAWQ checkpoints still load in vLLM
AutoGPTQArchived on 11 April 2025; Transformers no longer supports it and points to GPTQModelGPTQModel, a maintained fork that has since diverged, or llm-compressor's GPTQ modifier

Archived does not mean broken today. The packages still install, and thousands of existing checkpoints in their formats remain loadable by maintained runtimes. It means no fixes for new model architectures, new CUDA or PyTorch releases, or security issues. Every month an archived dependency stays in your build, the chance that an upgrade elsewhere breaks it grows.

Advertisement

bitsandbytes: what it really does

bitsandbytes provides two inference formats and the optimizers behind memory-efficient training. LLM.int8() stores weights in 8 bits and, during the matrix multiply, finds the few hidden dimensions where activations exceed a threshold (6.0 by default) and computes those in 16-bit, because those outlier features are where naive int8 loses accuracy; LLM.int8() in depth explains why. 4-bit NF4 and FP4 store weights in blocks of 64 with one scale per block; NF4's sixteen levels are placed at quantiles of a normal distribution, which matches how trained weights are distributed, as covered in NF4 explained. Double quantization then quantizes those per-block scales themselves.

At compute time the 4-bit weights are dequantized to the compute dtype, usually BF16, and multiplied. That is why the compute dtype setting matters and why bitsandbytes is primarily a memory saver: it shrinks weights to fit, while the arithmetic stays in 16 bits.

import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training

MODEL_ID = "your-org/base-8b"           # any causal LM in BF16

# 4-bit NF4 with double quantization: the QLoRA recipe. No calibration data.
bnb4 = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",          # or "fp4"
    bnb_4bit_use_double_quant=True,     # quantize the per-block absmax constants too
    bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, quantization_config=bnb4, device_map="auto")
model = prepare_model_for_kbit_training(model)
model = get_peft_model(model, LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
                                         target_modules="all-linear", task_type="CAUSAL_LM"))

# 8-bit LLM.int8(): outlier feature dimensions stay in 16-bit, the rest is int8.
bnb8 = BitsAndBytesConfig(load_in_8bit=True, llm_int8_threshold=6.0)

The snippet is the QLoRA pattern: a frozen 4-bit base with trainable 16-bit low-rank adapters. Gradients flow through the dequantized weights into the adapters, and the paged optimizers in bitsandbytes absorb memory spikes. This is the job bitsandbytes does better than anything else in this comparison. 4-bit models can also be saved and pushed to the hub, and vLLM can load both pre-quantized bitsandbytes checkpoints and quantize in flight with quantization="bitsandbytes"; recent vLLM documentation installs this support through a separate plugin package, so check the docs for your version.

Worked example: memory for an 8B model

Take an 8-billion-parameter model. In BF16 the weights alone need 8e9 times 2 bytes, 16 GB, before the KV cache and activations.

NF4 without double quantization: 4 bits per weight plus one 32-bit scale per block of 64, which is 0.5 bits per weight. That is 4.5 bits, so 8e9 times 4.5 divided by 8 gives 4.5 GB.

NF4 with double quantization: the scales become 8-bit, with a second-level 32-bit constant per 256 scales, so overhead falls to 8/64 plus 32/(64 times 256), about 0.127 bits per weight. Total about 4.13 bits, or 4.13 GB. The 0.37 GB saved can matter on a 24 GB card during fine-tuning.

AWQ or GPTQ INT4, group size 128 with a 16-bit scale and a packed zero point per group: roughly 4.15 to 4.25 bits per weight depending on how zeros are stored, so 4.2 to 4.3 GB. In practice the embedding and output head are often left in 16 bits, which adds noticeably on models with large vocabularies; count them separately.

Weight memory is therefore nearly identical across all three. The differences that decide your choice are elsewhere: whether you need calibration data, what the kernels do with the format, and which runtimes load it.

The calibrated paths today

For new work, the maintained successors are the place to start. llm-compressor applies AWQ or GPTQ as recipe modifiers in a one-shot pass and writes the compressed-tensors format, which vLLM loads natively. GPTQModel keeps a familiar load, quantize and save workflow and writes GPTQ-format checkpoints that Transformers, vLLM and SGLang can use.

# Calibrated INT4 for serving. Both snippets follow each project's README on its
# main branch in 2026; class names and import paths have moved between releases,
# so check the docs for the version you pin.

# --- AWQ, via llm-compressor (successor to AutoAWQ) -----------------------------
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor.modifiers.transform.awq import AWQModifier   # older releases: modifiers.awq

model = AutoModelForCausalLM.from_pretrained("./merged-bf16", torch_dtype="auto")
tok = AutoTokenizer.from_pretrained("./merged-bf16")
recipe = [AWQModifier(),
          QuantizationModifier(ignore=["lm_head"], scheme="W4A16_ASYM", targets=["Linear"])]
oneshot(model=model, dataset=calib_ds, recipe=recipe,          # calib_ds: your templated prompts
        max_seq_length=2048, num_calibration_samples=256)
model.save_pretrained("./model-awq-w4a16", save_compressed=True)
tok.save_pretrained("./model-awq-w4a16")

# --- GPTQ, via GPTQModel (successor to AutoGPTQ) ---------------------------------
from gptqmodel import GPTQModel, GPTQConfig                      # older releases: QuantizeConfig

gptq = GPTQModel.load("./merged-bf16", GPTQConfig(bits=4, group_size=128))
gptq.quantize(calib_texts, batch_size=1)                          # list of strings
gptq.save("./model-gptq-int4")

# Serve either one; vLLM reads the format from the checkpoint's config:
#   vllm serve ./model-awq-w4a16

Notice the input: a merged BF16 model, not a 4-bit one. If you fine-tuned with QLoRA, merge the adapters into a BF16 copy of the base first, then calibrate on prompts that look like production traffic, with your chat template applied. Merging adapters into NF4 weights and serving that loses the accuracy the calibrated methods exist to protect. For the details of calibration data and per-layer audits, see auditing a GPTQ checkpoint.

Checkpoint formats and what loads where

FormatProduced byLoaded byNotes
bitsandbytes 4-bit or 8-bitTransformers with BitsAndBytesConfigTransformers; vLLMCan also be created at load time from BF16
AutoAWQ (GEMM or GEMV packing)AutoAWQ, now archivedvLLM, Transformers and othersLarge legacy population on the hub; no new tooling
compressed-tensorsllm-compressor (AWQ, GPTQ, FP8 and others)vLLM natively; TransformersThe vLLM project's own format
GPTQAutoGPTQ (archived), GPTQModelTransformers via GPTQModel; vLLM; SGLangAct-order adds a group index tensor some kernels handle differently

For serving throughput, the format matters because of the kernels. vLLM runs INT4 weight-only checkpoints with Marlin-class mixed-precision kernels on supported NVIDIA GPUs, which keep the GPU busy when many requests are batched. bitsandbytes kernels were designed for memory savings in training and research and have historically been slower at large batch; recent releases have added a fused 4-bit inference path, so measure on your hardware rather than relying on old comparisons.

A decision guide

  • Fine-tuning a model that does not fit in 16 bits: bitsandbytes NF4 with QLoRA. Nothing else in this comparison is designed for it.
  • Trying a model today on a small GPU: bitsandbytes 4-bit or 8-bit at load time. No calibration, no new files.
  • Serving at high concurrency: a calibrated INT4 checkpoint, AWQ or GPTQ, produced by llm-compressor or GPTQModel and run on vLLM or SGLang.
  • Choosing between AWQ and GPTQ: measure both on your evaluation set. AWQ is quicker to calibrate and less sensitive to calibration data; GPTQ with act-order can be slightly more accurate on some models. The gap is model-dependent; AWQ vs GPTQ compared covers the mechanics.
  • CPU or Apple Silicon deployment: neither archived library is the answer; formats such as GGUF or MLX's own quantization are built for those targets.

Migrating off AutoAWQ and AutoGPTQ

  1. Inventory: list every model, script and container that imports awq or auto_gptq, and every checkpoint produced by them.
  2. Separate artefacts from tools. Existing checkpoints can keep serving on a maintained runtime; it is the quantization scripts that must move.
  3. Pin the current versions in a frozen environment so you can reproduce old checkpoints if needed.
  4. Port each recipe: AutoAWQ to llm-compressor's AWQ modifier, AutoGPTQ to GPTQModel or llm-compressor's GPTQ modifier. Carry over bits, group size, symmetric or asymmetric, act-order, the ignore list and the calibration set.
  5. Re-quantize the most important model with the new tool and compare it with the old checkpoint on the same evaluation gate, below. Expect small differences; investigate large ones.
  6. Switch serving to the new checkpoint behind a canary, then retire the archived packages from your images.

An evaluation gate for any quantized checkpoint

Whichever tool produced it, a quantized model ships only after it passes a gate against its BF16 parent: perplexity on held-out text from your domain, plus accuracy on labelled tasks that represent your product. Perplexity catches broad damage; task accuracy catches the narrow damage that matters, such as broken tool-call formatting or arithmetic.

import math, torch

@torch.no_grad()
def nll_per_token(model, tok, texts, max_len=2048):
    total, count = 0.0, 0
    for t in texts:
        ids = tok(t, return_tensors="pt", truncation=True, max_length=max_len).input_ids.to(model.device)
        out = model(ids, labels=ids)
        n = ids.numel() - 1
        total += out.loss.item() * n
        count += n
    return total / count

def gate(baseline, candidate, tok, held_out, task_eval):
    ppl_b = math.exp(nll_per_token(baseline, tok, held_out))
    ppl_c = math.exp(nll_per_token(candidate, tok, held_out))
    acc_b, acc_c = task_eval(baseline), task_eval(candidate)   # your own labelled tasks
    report = {"ppl_ratio": ppl_c / ppl_b, "task_drop": acc_b - acc_c}
    report["pass"] = report["ppl_ratio"] < 1.05 and report["task_drop"] < 0.01   # set your own bars
    return report

Set the thresholds from your product's tolerance, store the report with the checkpoint along with the tool version and calibration data hash, and rerun the gate whenever the runtime, kernel or tool version changes.

Failure modes

  • Silent dependency rot: an archived library breaks after a PyTorch or CUDA upgrade, discovered only when a re-quantization job fails.
  • Serving bitsandbytes at scale because the fine-tune used it, then discovering throughput far below a calibrated INT4 checkpoint.
  • Calibrating the wrong model: quantizing the base instead of the merged fine-tune, or calibrating on generic web text instead of templated production prompts.
  • Format mismatch: an act-order GPTQ checkpoint routed to a kernel that handles it slowly or not at all.
  • Copied tutorial code whose imports point at archived packages or renamed classes.
  • Memory surprises from unquantized embeddings and output heads, or from KV cache growth that weight quantization does not touch.

Trade-offs at a glance

OptionGainsCosts
bitsandbytes 4-bitNo data, instant, QLoRA trainingCompute in 16 bits; slower serving at large batch
bitsandbytes 8-bitNear-lossless memory halvingOutlier handling costs speed
AWQ (llm-compressor)Fast calibration, robust, fast kernelsNeeds data and a quantization pass
GPTQ (GPTQModel)Strong accuracy with tuning, wide runtime supportLonger calibration; act-order complicates kernels
Keeping archived toolsNo migration work nowNo fixes for new models or toolchains

What to do next

  1. Classify each workload as training, experimentation or serving before choosing a library.
  2. Use bitsandbytes NF4 with double quantization for QLoRA fine-tuning.
  3. For serving, merge adapters into BF16, then produce an AWQ or GPTQ INT4 checkpoint with llm-compressor or GPTQModel.
  4. Build the evaluation gate once and run it for every checkpoint, recording tool versions and calibration hashes.
  5. Inventory and remove imports of AutoAWQ and AutoGPTQ, keeping old checkpoints served from a maintained runtime.
  6. Benchmark throughput on your own GPU and batch sizes rather than trusting older comparisons.
  7. Count embedding, output head and KV cache memory separately when sizing hardware.
Key takeaway: bitsandbytes, AutoAWQ and AutoGPTQ answer different questions. bitsandbytes quantizes at load time with no data and is the tool for QLoRA training and quick experiments. AWQ and GPTQ calibrate on data to produce packed checkpoints for fast serving, and in 2026 they live on in llm-compressor and GPTQModel because the original libraries were archived in 2025. Pick by workload, merge before you calibrate, gate every checkpoint against its parent, and move your scripts off archived packages.