Distillation is usually sold as a quality technique: train a small student to imitate a large teacher and keep most of the teacher's skill. Teams actually approve it for a different reason. The student needs a fraction of the GPU time per request, and at enough volume that fraction turns into a large monthly saving. Whether the saving is real depends on numbers most proposals never write down: what the project costs once, what it costs every month after launch, and the traffic volume below which it never pays back.

This article builds that ledger from first principles. It covers the one-time lines, how to measure the serving saving in GPU-seconds per request, the hidden recurring costs (fallback, re-distillation, replica floors), a cost model you can run, and a worked example with a payback date. The mechanics are covered elsewhere on this site: how teacher and student losses work, the teacher scoring pass as a GPU job and sizing the student to a serving budget. Here the subject is money.

Where the saving comes from

Serving cost scales with the work the GPU does per token. Decode is mostly limited by memory bandwidth: every generated token reads the weights, so a model with a tenth of the parameters reads roughly a tenth of the bytes per step. Prefill is compute-bound, and its FLOPs also scale with parameter count, about 2N per token. A smaller model also fits on fewer GPUs, which removes tensor-parallel communication and frees memory for KV cache, so more requests share one batch.

Those effects multiply, which is why an 8B student can cost five to ten times less per request than a 70B teacher on the same hardware. The exact ratio is not something to look up; it is something to measure on your traffic, at your latency target, on your serving stack. The single most common error in distillation business cases is quoting a ratio from parameter counts or a vendor benchmark instead of from a load test.

Distillation is also not the only way to buy that ratio. Quantising the teacher, routing easy requests to an existing small model, prompt caching and better batching all cut cost with less engineering. A distillation case should beat the cheapest of those, not just the status quo.

The unit: GPU-seconds per request at the SLO

The unit that makes the ledger honest is GPU-seconds per request at the SLO. Run a load test that replays a sample of real prompts at a rising request rate, find the highest sustained rate at which p95 latency still meets the target, and divide GPUs by that rate:

gpu_seconds_per_request = gpus_per_replica / max_rps_within_slo

# example: teacher on 2 GPUs sustains 2.2 req/s inside the SLO -> 0.91 GPU-s/request
#          student on 1 GPU  sustains 8.3 req/s inside the SLO -> 0.12 GPU-s/request

Three corrections turn that into a cost. Multiply by a headroom factor, because you provision for peak and for losing a replica, not for the average; 1.3 to 2.0 is typical. Multiply by the GPU-hour price you actually pay, whether that is a cloud rate, a reserved rate or the amortised cost of owned hardware (see the cost analysis article for the full derivation). And remember the replica floor: a service needs at least two replicas for availability, so below a certain volume the student's cost stops falling with traffic.

Measure both models with the same prompt mix. A student tested on short prompts and a teacher tested on long ones produces a ratio that is wrong in the student's favour, and it will be discovered after launch, when the bill does not move as promised.

The one-time ledger

The one-time side has four lines. Teacher outputs come first. If the teacher is an API model you pay per token for every generated example; if you host it, you pay GPU time for generation, for scoring logits when you use soft targets, and for any judge model that filters bad samples. Self-hosted generation runs at large offline batch sizes, so it is far cheaper per token than interactive serving.

Student training is the second line and is easy to estimate. A dense transformer costs about 6ND FLOPs to train, where N is parameters and D is tokens seen. Divide by the GPU's peak FLOP rate times a realistic utilisation (35 to 50 percent is common for well-tuned fine-tuning) to get GPU-seconds. If you run the teacher online during training to produce soft targets, add 2 x N_teacher x D for its forward passes; with a 70B teacher and an 8B student, that forward pass costs about three times the student's own training. Caching top-k teacher logits once and reusing them across epochs and sweeps is usually the largest compute saving in the whole project.

Evaluation is the third line: human review of a few thousand outputs, judge-model runs and a shadow deployment. The fourth, engineering time, is usually larger than all the others combined. Two engineers for six weeks costs more than every GPU-hour in a typical 8B distillation run. A proposal that lists only GPU costs understates the investment several times over.

The distillation ledger: one-time spend up front, a per-request saving every month, upkeep foreverOne-timeTeacher outputsAPI tokens or self-hosted GPU-hStudent training + sweeps6 N D FLOPs per runEvaluationhuman review, judge runsEngineering timeusually the largest linePayback pointone-time / net monthly savingEvery monthTeacher GPU-seconds avoidedper request x volumeCascade fallbackescalations still hit the teacherRe-distillation upkeepteacher, data or policy driftReplica floormin replicas cost even at zero load---Distillation pays only when monthly volume x saving per request outruns upkeep and the replica floor.
One-time costs on the left set the investment; recurring costs on the right are subtracted from the per-request saving before computing payback.

Recurring costs that eat the saving

Three recurring lines eat into the saving, and they are why some distillation projects never pay back even though the student works.

Recurring costWhy it existsHow to estimate it
Cascade fallbackRequests the student is unsure about, or fails validation on, are re-sent to the teacherEscalation rate x teacher GPU-s per request; the student's own work on those requests is spent too
Re-distillationThe teacher gets upgraded, the policy changes or the input distribution drifts, and the student goes staleCost of a refresh run (data delta, training, evaluation, engineering) x refreshes per year
Replica floorMinimum replicas run whether or not traffic arrivesFloor GPUs x hours x price, compared with what volume alone would need
MonitoringQuality sampling, drift detection, escalation dashboardsSmall but non-zero: judge calls on a sample plus on-call time

The fallback line deserves attention because it drifts. A student that escalates 5 percent of traffic at launch may escalate 15 percent six months later as users find new things to ask. Track the escalation rate as a cost metric, not only as a quality metric, and set a threshold at which a refresh is triggered.

A cost model you can run

The model below turns the ledger into a payback date and a minimum volume. It is deliberately small so that every input is visible and can be argued about. All prices are inputs you supply; the defaults are illustrative, not market quotes.

from dataclasses import dataclass

@dataclass
class Plan:
    requests_per_month: float
    teacher_gpu_s: float          # measured GPU-seconds per request at the SLO
    student_gpu_s: float
    escalation_rate: float        # share of requests also sent to the teacher
    headroom: float = 1.5         # peak and failover provisioning
    gpu_hour_usd: float = 2.50    # illustrative; use your own rate
    one_time_usd: float = 0.0     # data + training + eval + engineering
    upkeep_usd_per_month: float = 0.0

    def cost_per_request(self, gpu_s):
        return gpu_s * self.headroom / 3600 * self.gpu_hour_usd

    def saving_per_request(self):
        teacher = self.cost_per_request(self.teacher_gpu_s)
        student = self.cost_per_request(self.student_gpu_s + self.escalation_rate * self.teacher_gpu_s)
        return teacher - student

    def net_monthly_saving(self):
        return self.requests_per_month * self.saving_per_request() - self.upkeep_usd_per_month

    def payback_months(self):
        net = self.net_monthly_saving()
        return float("inf") if net <= 0 else self.one_time_usd / net

    def min_volume_for_upkeep(self):
        return self.upkeep_usd_per_month / self.saving_per_request()

Two outputs matter. The payback period says whether the project is worth funding. The minimum volume says whether it stays worth running: if traffic falls below the volume at which savings cover upkeep, the student is losing money every month, and serving the teacher (or a quantised teacher) is cheaper.

Worked example: payback in under four months

A support team summarises tickets with a 70B model: 30 million requests a month, about 1,500 input and 250 output tokens each. Load tests show the teacher at 0.9 GPU-seconds per request and an 8B student at 0.12. The student escalates 8 percent of requests to the teacher. Headroom is 1.5 and the illustrative GPU price is $2.50 an hour.

LineArithmeticAmount
Teacher serving, per month30M x 0.9 s x 1.5 / 3600 x $2.50$28,125
Student serving, per month30M x 0.12 s x 1.5 / 3600 x $2.50$3,750
Fallback, per month30M x 0.08 x 0.9 s x 1.5 / 3600 x $2.50$2,250
Teacher data via API (illustrative $3 / $15 per M tokens)400k examples: 600M in, 120M out$3,600
Training compute: logits cached once, 4 student runsabout 71 GPU-h scoring once + 4 x 73 GPU-h, about 365 GPU-h, rounded up for failed runs$1,500
Human evaluation2,000 items at $1.50$3,000
Engineering2 engineers x 6 weeks x $4,000 loaded per week$48,000
Upkeep: quarterly refresh at $20,000per month$6,667

The training estimate checks out from first principles. 400,000 examples at 1,800 tokens is 720M tokens; three epochs make 2.16B. Student training is 6 x 8e9 x 2.16e9, about 1.0e20 FLOPs, and scoring the data once with the teacher is 2 x 70e9 x 0.72e9, about 1.0e20 more. At an assumed 400 TFLOP/s of delivered throughput per GPU that is about 73 GPU-hours per student run and 71 for the one-time scoring pass, a few hundred dollars each. Compute is the smallest line in the project.

The totals: one-time spend of $56,100; net monthly saving of $28,125 - $3,750 - $2,250 - $6,667 = $15,458; payback in 56,100 / 15,458, about 3.6 months. The saving per request is $0.0007375, so upkeep alone needs about 9 million requests a month to cover. At 3 million requests a month the student would serve perfectly well and lose money every month.

Cumulative cost in the worked example: crossover at about 3.6 months024681012$0k$100k$200k$300kmonths after launchbreak-evenkeep serving the teacherdistil: one-time + student, fallback, upkeep
Teacher-only cost (red) against the distilled path (green), using the numbers above. A lower volume flattens the gap between the slopes and pushes the crossover right, possibly off the chart.

Price the alternatives with the same model

Before committing, price the alternatives with the same model. Each changes one input.

LeverInput it changesWhen it wins over distillation
Quantise the teacher (FP8, INT4 weights)teacher_gpu_s falls; no one-time data or trainingQuality must match the teacher closely, or volume is too low to amortise a project
Route to an existing small modelEscalation is driven by a router, no trainingA general small model already handles most of the traffic acceptably
Prompt caching and prefix reusePrefill cost falls for repeated system promptsLong shared prefixes dominate input tokens
Distil, then quantise the studentstudent_gpu_s falls furtherHigh volume; stacks with every lever above
Fine-tune a small model on labels onlyOne-time cost falls (no teacher logits)Task is narrow and labels are plentiful; soft targets add little

Licensing is a gating input, not a cost line. Many model licences and API terms restrict using outputs to train competing models; check the terms that apply to your teacher before generating a single example. The AI model licences article covers how to build that check into a licence gate.

Failure modes

  • Ratio from parameter counts. The case assumes 70B/8B means about 9x cheaper; the load test shows 5x because the student is small enough that batching, not bandwidth, now limits it. Measure first.
  • Fallback ignored. Escalated requests pay for the student and the teacher. A 20 percent escalation rate can erase most of the saving.
  • Replica floor at low volume. The student needs one GPU for the traffic but runs on two for availability, plus a teacher pool kept warm for fallback. Pooling the fallback with another service is often the fix.
  • Stale student. The teacher is upgraded, product compares the two and the student looks worse. Without a funded refresh cadence the student is retired early and the one-time spend is never recovered.
  • Online teacher in training. Running the teacher forward every epoch and every sweep multiplies compute several times over. Cache soft targets once.
  • Quality cost not priced. If wrong summaries create extra support handling, that cost belongs in the model as a per-error charge, or the saving is overstated.

Trade-offs

Distillation trades a fixed investment and an ongoing maintenance obligation for a lower marginal cost. It is the right trade for high, stable volume on a well-defined task, where the student can be evaluated against a clear standard. It is the wrong trade for low or volatile volume, for tasks that change monthly, and for teams that cannot staff refreshes. The middle ground, quantising and routing first and distilling only the highest-volume task, often gets most of the saving with a fraction of the risk.

What to do next

  1. Load-test the current model on a replay of real prompts and record GPU-seconds per request at the SLO.
  2. Price the cheapest alternatives first: quantise the teacher, add caching, route to an existing small model.
  3. Fill in the cost model with your GPU rate, volume and an honest engineering estimate, and compute payback and minimum volume.
  4. Check the teacher's licence or API terms for restrictions on training with its outputs before generating data.
  5. If the case holds, cache teacher logits once, budget sweeps, and measure the student's escalation rate in shadow mode.
  6. After launch, track escalation rate, volume and upkeep monthly against the minimum-volume figure, and retire the student if it falls below.
Key takeaway: Distillation saves money only when monthly volume times the per-request saving, measured in GPU-seconds at the SLO, outruns fallback, refresh upkeep and the replica floor. Price the engineering honestly, cache teacher logits, compare against quantising and routing first, and track volume against the minimum that keeps the student profitable.