FINE-TUNING · WHERE TRAINING MEMORY GOES

LoRA + QLoRA Memory Calculator

Count the adapter parameters of a published model exactly, then add base weights, optimizer state, activations and logits to see whether training fits your GPU.

Direct tool · updates as you edit

Your experiment

Start with LoRA rank 16 on all linear layers of Llama 3.2 3B, 2,048 tokens, micro-batch 4 and a 16 GiB GPU. Then switch to full fine-tuning, turn off checkpointing, raise the sequence length, and try Qwen2.5-7B with LoRA and with QLoRA.

Every input recomputes the result immediately; there is no animation because nothing here unfolds over time. An input outside its allowed range is rejected with a message and the previous valid result stays on screen.

Computed data

Metrics

Adapter counts are exact: r x (inputs + outputs) for each adapted matrix in every layer. Full fine-tuning uses 16 bytes per parameter for mixed-precision Adam (Rajbhandari et al., ZeRO, 2020); adapters are kept in fp32 with Adam, also 16 bytes per parameter. QLoRA stores linear layers in NF4 with double quantization, 4.127 bits per parameter (Dettmers et al., 2023), and keeps embeddings and the output head in bf16. Activations use the per-layer estimate s x b x h x (34 + 5as/h) bytes of Korthikanti et al. (2022), which assumes a GPT-style block, so treat it as an estimate; framework buffers and fragmentation are not counted.

Counting adapter parameters

A LoRA adapter on a matrix with n inputs and m outputs adds two low-rank factors, r x (n + m) parameters. For Llama 3.2 3B as the Published configuration, with LoRA rank r 16 and Adapted matrices set to all linear layers, that is 868,352 per layer and 24,313,856 in 28 layers, 0.757% of the 3,212,749,824 base parameters. Adapting only q and v gives 4,587,520, all four attention projections 9,175,040, and rank 64 on everything 97,255,424. The trainable state costs 16 bytes per parameter: fp32 weight, gradient and two Adam moments.

The opening run

LoRA on a bf16 base needs 5.98 GiB of base weights, 0.36 GiB of adapter state, 2.11 GiB of activations with checkpointing and 3.91 GiB of logits: 12.37 GiB, which FITS the 16 GiB chosen as GPU memory, the line the total is compared against. The logits term surprises people. The loss needs a score for every vocabulary entry at every position, s x b x V x 4 bytes, and with 128,256 tokens, 2,048 positions and a Micro-batch size of 4 it is almost twice the activations; both the activation and logits terms grow in proportion to the micro-batch.

Three ways to run out

Full fine-tuning keeps 16 bytes for every parameter, 47.87 GiB of state, and needs 53.9 GiB: DOES NOT FIT. Turning off gradient checkpointing stores every layer's activations, 22.31 GiB, for 32.57 GiB in total. Raising the sequence length to 8,192 grows activations and logits together to 8.44 and 15.66 GiB, 30.44 GiB; a micro-batch of 16 does exactly the same. Without a fused attention kernel the score matrices add the 5as/h term, and the opening run needs 14.24 GiB.

QLoRA for the bigger model

Qwen2.5-7B with LoRA needs 21.89 GiB, beyond 16 GiB, because its bf16 base alone is 14.19 GiB. QLoRA stores the linear layers at 4.127 bits, so the base drops to 5.17 GiB and the run needs 12.87 GiB, which fits. On Llama 3.2 3B QLoRA brings the total from 12.37 to 8.47 GiB. The adapters themselves, 40,370,176 parameters for Qwen2.5-7B, are the same in both methods.

Reading the tool and its limits

The stacked bar shows the four terms against the GPU line, the lanes give the trainable count and each term, and the table shows the formulas. The activation formula is an estimate for a GPT-style block; models with gated feed-forward layers and grouped-query attention differ in detail, and frameworks add workspace and fragmentation. Use the result to decide what to try first, then confirm with the peak memory your trainer reports.