SMALL MODELS · FIT BEFORE YOU BENCHMARK
SLM Selector
Filter nine published small models by device memory, weight format, context length and licence, and see which ones actually fit and why the others do not.
Your experiment
Start with 8 GiB, Q4_K weights, 8,192 tokens and any licence. Read why the one rejected model fails, then shrink the memory, switch to F16, ask for 32,768 or 131,072 tokens and restrict the licence.
Every input recomputes the result immediately; there is no animation because nothing here unfolds over time. An input outside its allowed range is rejected with a message and the previous valid result stays on screen.
Computed data
Metrics
Architectures come from each model's config.json and licences from the model cards, fetched on 2026-10-10. Weights are sized as a pure GGUF file of the chosen type (ggml block sizes, with llama.cpp's fallback for rows that do not divide into blocks); the KV cache is f16 for one sequence, with Gemma 2's local layers capped at its 4,096-token window. The tool says nothing about quality: no benchmark scores are used, so shortlist here and evaluate on your own task.
Three gates before any benchmark
A small model is only a candidate if it fits three hard constraints: its weights plus KV cache plus a runtime reserve fit the device, its trained context covers the prompts you need, and its licence allows your use. This tool applies all three to nine published configurations and prints the reason for every rejection. It deliberately uses no quality scores: once you have a shortlist, compare the survivors on your own evaluation set.
The opening shortlist
With 8 GiB of Device memory, Q4_K weights, 8,192 tokens of Context you need and a 1 GiB reserve, 8 of 9 models fit and the largest is Qwen2.5-7B at 5.43 GiB: 3.99 GiB of weights and 0.44 GiB of KV cache. Phi-3-mini-4k is rejected even though it would fit in memory, because its config lists 4,096 positions: context 4,096 below 8,192. SmolLM2-1.7B shows why the cache matters: with 32 key and value heads its cache at 8,192 tokens is 1.5 GiB, more than its 0.9 GiB of weights.
When nothing fits
Weight format multiplies the weight term. At F16 the 8 GiB device holds 7 of 9 and the largest becomes Llama 3.2 3B at 7.86 GiB. Shrink the device to 2 GiB with F16 and the result is NO MODEL FITS: even Qwen2.5-0.5B needs a little over 2 GiB once the reserve is included. Return to Q4_K at 2 GiB and 2 models fit, the largest Llama 3.2 1B at 1.9 GiB. The Reserve for runtime and OS slider is added to every model's total before the comparison with the device; it stands for compute buffers, the runtime and the operating system, and setting it to 0 is optimistic.
Context and licence
Asking for 32,768 tokens leaves 6 models: Gemma 2 2B and SmolLM2-1.7B list 8,192 positions, and Qwen2.5-7B now needs 6.74 GiB. At 131,072 tokens only Llama 3.2 1B remains, at 5.65 GiB, because its config lists 131,072 positions and its 8 KV heads of size 64 keep the cache small. The licence rule Apache-2.0 or MIT only leaves 4 at the opening settings: the Qwen2.5 Research, Gemma and Llama licences carry their own terms, which you must read rather than infer from a label.
Reading the tool and its limits
The bars compare each model with the device line, the lanes list fits and rejections, and the table gives parameters, weight size, KV cache, total, trained context and licence for all nine. Sizes are computed for one sequence; serving several at once multiplies the KV term. Trained context is the max positions value in each config, so a model card that documents a longer context with rope scaling is not reflected here.