Distillation is usually sold as a quality technique: train a small student to imitate a large teacher and keep most of the teacher's skill. Teams actually approve it for a different reason. The student needs a fraction of the GPU time per request, and at enough volume that fraction turns into a large monthly saving. Whether the saving is real depends on numbers most proposals never write down: what the project costs once, what it costs every month after launch, and the traffic volume below which it never pays back.
This article builds that ledger from first principles. It covers the one-time lines, how to measure the serving saving in GPU-seconds per request, the hidden recurring costs (fallback, re-distillation, replica floors), a cost model you can run, and a worked example with a payback date. The mechanics are covered elsewhere on this site: how teacher and student losses work, the teacher scoring pass as a GPU job and sizing the student to a serving budget. Here the subject is money.
Where the saving comes from
Serving cost scales with the work the GPU does per token. Decode is mostly limited by memory bandwidth: every generated token reads the weights, so a model with a tenth of the parameters reads roughly a tenth of the bytes per step. Prefill is compute-bound, and its FLOPs also scale with parameter count, about 2N per token. A smaller model also fits on fewer GPUs, which removes tensor-parallel communication and frees memory for KV cache, so more requests share one batch.
Those effects multiply, which is why an 8B student can cost five to ten times less per request than a 70B teacher on the same hardware. The exact ratio is not something to look up; it is something to measure on your traffic, at your latency target, on your serving stack. The single most common error in distillation business cases is quoting a ratio from parameter counts or a vendor benchmark instead of from a load test.
Distillation is also not the only way to buy that ratio. Quantising the teacher, routing easy requests to an existing small model, prompt caching and better batching all cut cost with less engineering. A distillation case should beat the cheapest of those, not just the status quo.
The unit: GPU-seconds per request at the SLO
The unit that makes the ledger honest is GPU-seconds per request at the SLO. Run a load test that replays a sample of real prompts at a rising request rate, find the highest sustained rate at which p95 latency still meets the target, and divide GPUs by that rate:
gpu_seconds_per_request = gpus_per_replica / max_rps_within_slo
# example: teacher on 2 GPUs sustains 2.2 req/s inside the SLO -> 0.91 GPU-s/request
# student on 1 GPU sustains 8.3 req/s inside the SLO -> 0.12 GPU-s/requestThree corrections turn that into a cost. Multiply by a headroom factor, because you provision for peak and for losing a replica, not for the average; 1.3 to 2.0 is typical. Multiply by the GPU-hour price you actually pay, whether that is a cloud rate, a reserved rate or the amortised cost of owned hardware (see the cost analysis article for the full derivation). And remember the replica floor: a service needs at least two replicas for availability, so below a certain volume the student's cost stops falling with traffic.
Measure both models with the same prompt mix. A student tested on short prompts and a teacher tested on long ones produces a ratio that is wrong in the student's favour, and it will be discovered after launch, when the bill does not move as promised.
The one-time ledger
The one-time side has four lines. Teacher outputs come first. If the teacher is an API model you pay per token for every generated example; if you host it, you pay GPU time for generation, for scoring logits when you use soft targets, and for any judge model that filters bad samples. Self-hosted generation runs at large offline batch sizes, so it is far cheaper per token than interactive serving.
Student training is the second line and is easy to estimate. A dense transformer costs about 6ND FLOPs to train, where N is parameters and D is tokens seen. Divide by the GPU's peak FLOP rate times a realistic utilisation (35 to 50 percent is common for well-tuned fine-tuning) to get GPU-seconds. If you run the teacher online during training to produce soft targets, add 2 x N_teacher x D for its forward passes; with a 70B teacher and an 8B student, that forward pass costs about three times the student's own training. Caching top-k teacher logits once and reusing them across epochs and sweeps is usually the largest compute saving in the whole project.
Evaluation is the third line: human review of a few thousand outputs, judge-model runs and a shadow deployment. The fourth, engineering time, is usually larger than all the others combined. Two engineers for six weeks costs more than every GPU-hour in a typical 8B distillation run. A proposal that lists only GPU costs understates the investment several times over.
Recurring costs that eat the saving
Three recurring lines eat into the saving, and they are why some distillation projects never pay back even though the student works.
| Recurring cost | Why it exists | How to estimate it |
|---|---|---|
| Cascade fallback | Requests the student is unsure about, or fails validation on, are re-sent to the teacher | Escalation rate x teacher GPU-s per request; the student's own work on those requests is spent too |
| Re-distillation | The teacher gets upgraded, the policy changes or the input distribution drifts, and the student goes stale | Cost of a refresh run (data delta, training, evaluation, engineering) x refreshes per year |
| Replica floor | Minimum replicas run whether or not traffic arrives | Floor GPUs x hours x price, compared with what volume alone would need |
| Monitoring | Quality sampling, drift detection, escalation dashboards | Small but non-zero: judge calls on a sample plus on-call time |
The fallback line deserves attention because it drifts. A student that escalates 5 percent of traffic at launch may escalate 15 percent six months later as users find new things to ask. Track the escalation rate as a cost metric, not only as a quality metric, and set a threshold at which a refresh is triggered.
A cost model you can run
The model below turns the ledger into a payback date and a minimum volume. It is deliberately small so that every input is visible and can be argued about. All prices are inputs you supply; the defaults are illustrative, not market quotes.
from dataclasses import dataclass
@dataclass
class Plan:
requests_per_month: float
teacher_gpu_s: float # measured GPU-seconds per request at the SLO
student_gpu_s: float
escalation_rate: float # share of requests also sent to the teacher
headroom: float = 1.5 # peak and failover provisioning
gpu_hour_usd: float = 2.50 # illustrative; use your own rate
one_time_usd: float = 0.0 # data + training + eval + engineering
upkeep_usd_per_month: float = 0.0
def cost_per_request(self, gpu_s):
return gpu_s * self.headroom / 3600 * self.gpu_hour_usd
def saving_per_request(self):
teacher = self.cost_per_request(self.teacher_gpu_s)
student = self.cost_per_request(self.student_gpu_s + self.escalation_rate * self.teacher_gpu_s)
return teacher - student
def net_monthly_saving(self):
return self.requests_per_month * self.saving_per_request() - self.upkeep_usd_per_month
def payback_months(self):
net = self.net_monthly_saving()
return float("inf") if net <= 0 else self.one_time_usd / net
def min_volume_for_upkeep(self):
return self.upkeep_usd_per_month / self.saving_per_request()Two outputs matter. The payback period says whether the project is worth funding. The minimum volume says whether it stays worth running: if traffic falls below the volume at which savings cover upkeep, the student is losing money every month, and serving the teacher (or a quantised teacher) is cheaper.
Worked example: payback in under four months
A support team summarises tickets with a 70B model: 30 million requests a month, about 1,500 input and 250 output tokens each. Load tests show the teacher at 0.9 GPU-seconds per request and an 8B student at 0.12. The student escalates 8 percent of requests to the teacher. Headroom is 1.5 and the illustrative GPU price is $2.50 an hour.
| Line | Arithmetic | Amount |
|---|---|---|
| Teacher serving, per month | 30M x 0.9 s x 1.5 / 3600 x $2.50 | $28,125 |
| Student serving, per month | 30M x 0.12 s x 1.5 / 3600 x $2.50 | $3,750 |
| Fallback, per month | 30M x 0.08 x 0.9 s x 1.5 / 3600 x $2.50 | $2,250 |
| Teacher data via API (illustrative $3 / $15 per M tokens) | 400k examples: 600M in, 120M out | $3,600 |
| Training compute: logits cached once, 4 student runs | about 71 GPU-h scoring once + 4 x 73 GPU-h, about 365 GPU-h, rounded up for failed runs | $1,500 |
| Human evaluation | 2,000 items at $1.50 | $3,000 |
| Engineering | 2 engineers x 6 weeks x $4,000 loaded per week | $48,000 |
| Upkeep: quarterly refresh at $20,000 | per month | $6,667 |
The training estimate checks out from first principles. 400,000 examples at 1,800 tokens is 720M tokens; three epochs make 2.16B. Student training is 6 x 8e9 x 2.16e9, about 1.0e20 FLOPs, and scoring the data once with the teacher is 2 x 70e9 x 0.72e9, about 1.0e20 more. At an assumed 400 TFLOP/s of delivered throughput per GPU that is about 73 GPU-hours per student run and 71 for the one-time scoring pass, a few hundred dollars each. Compute is the smallest line in the project.
The totals: one-time spend of $56,100; net monthly saving of $28,125 - $3,750 - $2,250 - $6,667 = $15,458; payback in 56,100 / 15,458, about 3.6 months. The saving per request is $0.0007375, so upkeep alone needs about 9 million requests a month to cover. At 3 million requests a month the student would serve perfectly well and lose money every month.
Price the alternatives with the same model
Before committing, price the alternatives with the same model. Each changes one input.
| Lever | Input it changes | When it wins over distillation |
|---|---|---|
| Quantise the teacher (FP8, INT4 weights) | teacher_gpu_s falls; no one-time data or training | Quality must match the teacher closely, or volume is too low to amortise a project |
| Route to an existing small model | Escalation is driven by a router, no training | A general small model already handles most of the traffic acceptably |
| Prompt caching and prefix reuse | Prefill cost falls for repeated system prompts | Long shared prefixes dominate input tokens |
| Distil, then quantise the student | student_gpu_s falls further | High volume; stacks with every lever above |
| Fine-tune a small model on labels only | One-time cost falls (no teacher logits) | Task is narrow and labels are plentiful; soft targets add little |
Licensing is a gating input, not a cost line. Many model licences and API terms restrict using outputs to train competing models; check the terms that apply to your teacher before generating a single example. The AI model licences article covers how to build that check into a licence gate.
Failure modes
- Ratio from parameter counts. The case assumes 70B/8B means about 9x cheaper; the load test shows 5x because the student is small enough that batching, not bandwidth, now limits it. Measure first.
- Fallback ignored. Escalated requests pay for the student and the teacher. A 20 percent escalation rate can erase most of the saving.
- Replica floor at low volume. The student needs one GPU for the traffic but runs on two for availability, plus a teacher pool kept warm for fallback. Pooling the fallback with another service is often the fix.
- Stale student. The teacher is upgraded, product compares the two and the student looks worse. Without a funded refresh cadence the student is retired early and the one-time spend is never recovered.
- Online teacher in training. Running the teacher forward every epoch and every sweep multiplies compute several times over. Cache soft targets once.
- Quality cost not priced. If wrong summaries create extra support handling, that cost belongs in the model as a per-error charge, or the saving is overstated.
Trade-offs
Distillation trades a fixed investment and an ongoing maintenance obligation for a lower marginal cost. It is the right trade for high, stable volume on a well-defined task, where the student can be evaluated against a clear standard. It is the wrong trade for low or volatile volume, for tasks that change monthly, and for teams that cannot staff refreshes. The middle ground, quantising and routing first and distilling only the highest-volume task, often gets most of the saving with a fraction of the risk.
What to do next
- Load-test the current model on a replay of real prompts and record GPU-seconds per request at the SLO.
- Price the cheapest alternatives first: quantise the teacher, add caching, route to an existing small model.
- Fill in the cost model with your GPU rate, volume and an honest engineering estimate, and compute payback and minimum volume.
- Check the teacher's licence or API terms for restrictions on training with its outputs before generating data.
- If the case holds, cache teacher logits once, budget sweeps, and measure the student's escalation rate in shadow mode.
- After launch, track escalation rate, volume and upkeep monthly against the minimum-volume figure, and retire the student if it falls below.