A data flywheel is the loop in which a deployed model generates the data that trains its next version. Users ask questions, the system logs what happened and how users reacted, a pipeline mines those logs for examples worth learning from, GPUs label and train on them, and a gated release puts the improved model back in front of users. The idea sells itself because each turn should make the product better and the data harder for a competitor to copy. In practice most flywheels stall or, worse, turn the wrong way. They teach the model to flatter users, or to repeat its own mistakes, or to memorise the evaluation set.
This article treats the flywheel as an engineering system with GPU budgets. It covers which signals are worth capturing and how each is biased, an event schema with a consent boundary, code that turns regenerations and edits into preference pairs, a worked sizing of the GPU hours each stage needs, the release gate that stops a bad turn from shipping, and the failure modes that make loops collapse. The training methods themselves are covered in DPO training and the RLHF pipeline; here the subject is the loop around them.
The loop and its GPU workloads
Each arrow in the diagram is a pipeline with its own latency and failure modes. The useful number to track is the turn time: how long from a user behaviour appearing in logs to a model that has learned from it serving traffic. Teams that run a turn per week or per fortnight tend to learn faster than teams that run one large retrain a quarter. Each turn is a small, attributable change that can be rolled back, and a regression shows up within days of the data that caused it.
The flywheel also has a GPU shape that differs from pretraining. Serving is continuous and latency-bound. Judging and labeling are offline batch inference that can soak up idle serving capacity overnight. Training is short and bursty, often a few GPU-hours of LoRA or DPO per turn. Evaluation is batch inference again, run twice, once for the candidate and once for the incumbent. Planning the loop means planning those four workloads on shared hardware.
Signals and their biases
Signals differ enormously in volume, strength and bias. Explicit ratings are clear but rare and come from an unrepresentative slice of users. Implicit signals are plentiful but ambiguous. Outcome signals, such as whether generated code passed its tests or a support ticket was reopened, are the most trustworthy and the slowest to arrive.
| Signal | Typical volume | What it suggests | Main bias or trap |
|---|---|---|---|
| Thumbs up or down | Low, a few percent of sessions at most | Satisfaction with the answer | Angry or delighted users only; rewards flattery |
| Regenerate | Moderate | First answer was unsatisfying | Also used to explore alternatives |
| User edits the output | Moderate in writing and code tools | Edited text is closer to wanted | Edits can be stylistic, not corrections |
| Copy or accept | High | Answer was usable | Copying a wrong answer still counts |
| Task outcome (tests, resolution) | Varies | Answer actually worked | Delayed; attribution across turns is hard |
| Abandonment | High | Something failed | Also happens on success |
Treat every signal as a candidate generator, not a label. A thumbs-down nominates a conversation for review; it does not prove the answer was wrong. The stage that converts candidates into training labels is where judges, rubrics and humans come in, and it is where most of the quality is won.
The event log and the consent boundary
Everything downstream depends on one event schema that is written at serving time and never patched after the fact. It must record the exact model and prompt-template version, because a preference pair is only meaningful against the policy that produced it. It must record the user's consent state at the time of the request, because consent can change later and deletion requests must be honoured in datasets as well as in logs.
from dataclasses import dataclass, field
from typing import Optional
@dataclass(frozen=True)
class TurnEvent:
event_id: str
conversation_id: str
user_id: str # pseudonymous, for deletion requests
ts: float
model_version: str # exact weights + template hash
prompt: str # after PII scrubbing
response: str
consent_training: bool # captured at request time
tenant: str
signals: dict = field(default_factory=dict)
# e.g. {"thumb": -1, "regenerated_by": "evt_123",
# "edited_to": "...", "outcome": "tests_passed"}
def eligible(e: TurnEvent, deleted_users: set, opted_out_tenants: set) -> bool:
return (e.consent_training
and e.tenant not in opted_out_tenants
and e.user_id not in deleted_users)Scrub personal data before the event is written, not in the training pipeline, so that no downstream copy ever holds it. Keep the raw response hash so you can later tell whether a training example came from model output, which matters for the model collapse problem described below.
From logs to preference pairs
Turning events into training data has four steps: filter, deduplicate, judge and pair. Filtering removes ineligible events, very short turns and known abuse. Deduplication matters more than it looks, because popular prompts appear thousands of times and would otherwise dominate a turn; the near-duplicate techniques from GPU curation pipelines apply directly. Judging uses an LLM with a rubric to score candidates and to reject pairs whose preferred side is still wrong. The batch-inference techniques for that, such as prefix caching and constrained labels, are covered in LLM labeling on GPUs.
Pairing is where implicit signals become preference data. A regeneration followed by a thumbs-up or a copy gives a natural pair. A user edit gives a pair of the original against the edited text, but only when the judge agrees that the edit is a correction rather than a style change.
import hashlib
def prompt_hash(prompt):
return hashlib.sha256(" ".join(prompt.lower().split()).encode()).hexdigest()
def build_pairs(events, judge, eval_hashes, min_margin=1.0):
by_id = {e.event_id: e for e in events}
pairs = []
for e in events:
if e.signals.get("thumb") != -1 or "regenerated_by" not in e.signals:
continue
better = by_id.get(e.signals["regenerated_by"])
if better is None or better.signals.get("thumb", 0) < 0:
continue
if prompt_hash(e.prompt) in eval_hashes: # eval firewall
continue
s_bad, s_good = judge(e.prompt, e.response), judge(e.prompt, better.response)
if s_good - s_bad < min_margin or s_good < 4: # 1-5 rubric
continue # weak or wrong pair
pairs.append({"prompt": e.prompt, "chosen": better.response,
"rejected": e.response, "policy": e.model_version})
return pairsThe prompt_hash check is the eval firewall from the diagram. Hash prompts after normalising whitespace and case, and also run a near-duplicate check against the eval set, because a user who pastes a benchmark question with one word changed will slip past an exact hash.
Worked example: GPU hours for one weekly turn
Here is a worked sizing for one weekly turn. Every input is an assumption to replace with your own measurements. The product serves 1,000,000 conversations a day, 2% receive an explicit rating and 3% are regenerated, giving about 350,000 candidate events a week before filtering and deduplication. The judge reads each candidate once at about 2,000 tokens of prompt plus rubric, which is 700 million input tokens and only a few output tokens each. That is prefill-dominated batch work that fits in off-peak serving capacity, and a shared rubric prefix makes it cheaper still.
Suppose 40,000 preference pairs survive, each with an 800-token prompt and 400-token chosen and rejected responses. DPO runs the policy on both sequences, so each pair is about 2,400 tokens and the epoch is 96 million tokens. For an 8-billion-parameter model, full-parameter training costs about 6 x 8e9 x 9.6e7, which is 4.6e18 FLOPs. Precomputing reference log-probabilities adds a forward pass, 2 x 8e9 x 9.6e7, or 1.5e18, for a total of about 6.1e18. Take as inputs a peak of 989 TFLOP/s, the dense BF16 spec-sheet figure for an H100 SXM, and an assumed 35% utilisation. The epoch then needs about 17,700 GPU-seconds, roughly 5 GPU-hours, or under 40 minutes on 8 GPUs.
The FLOPs are not the binding constraint; memory is. Full-parameter training with Adam in mixed precision needs about 16 bytes per parameter before activations, around 128 GB for 8B parameters, so it must be sharded across GPUs. LoRA cuts optimizer state to a small fraction and lets one GPU handle the turn, which is why most flywheels train adapters on every turn and do a full fine-tune less often. Evaluation then costs two batch generations over the frozen suite, often more GPU time than training itself when the suite is large and the outputs are long.
The release gate
The release gate is what makes the loop safe to run every week. It compares the candidate with the incumbent on a frozen evaluation suite that never changes within a quarter. It also uses a rolling suite drawn from recent traffic, held out by the eval firewall, and safety and refusal suites. Report results by slice: tenant, language, task type and prompt length. An aggregate win can hide a large loss on a minority slice. Benchmark evals on GPUs covers how to run these suites efficiently.
def release_gate(cand, inc, slices, max_slice_drop=0.02, min_win=0.0):
report = {}
for name in slices:
d = cand[name] - inc[name]
report[name] = round(d, 4)
if d < -max_slice_drop:
return False, f"slice {name} regressed by {-d:.3f}", report
if cand["safety_refusal_ok"] < inc["safety_refusal_ok"]:
return False, "safety suite regressed", report
if cand["overall"] - inc["overall"] < min_win:
return False, "no overall improvement", report
return True, "promote to shadow", reportPassing the gate earns a shadow deployment, where the candidate answers live prompts without users seeing it, and then a canary at a small traffic share. Watch the same signals that feed the flywheel. A rise in regenerations or thumbs-down on the canary is the fastest regression detector you have, and automatic rollback should key on it.
Failure modes
- Sycophancy from optimising ratings. Users rate agreement and confidence highly. A loop trained on thumbs-up learns to agree. Require a rubric judge on correctness for every chosen response, and keep a suite of prompts where the right answer contradicts the user.
- Selection bias. Ratings come mostly from the most and least happy users, and from some tenants far more than others. Reweight or cap examples per tenant and per prompt cluster.
- Model collapse. Training repeatedly on a model's own outputs narrows the distribution and loses rare cases, as Shumailov and colleagues showed in Nature in 2024. Keep a fixed share of human-written or verified data in every turn, and track output diversity across turns.
- Eval contamination. Popular benchmark prompts arrive through user traffic. Without the firewall the model memorises them and the gate reports progress that does not exist.
- Stale policy pairs. Pairs generated by version N-3 describe mistakes the current model no longer makes. Weight by recency or filter by policy version.
- Consent and deletion drift. A deletion request that removes logs but not derived datasets and trained adapters is a compliance failure. Store dataset lineage per example.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Weekly turns | Fast learning, small attributable changes | Constant eval and release overhead |
| Quarterly retrain | Low operational load | Slow feedback, big risky diffs |
| Implicit signals | Volume | Ambiguity, needs a strong judge |
| Human review of pairs | Highest label quality | Expensive, slow, privacy exposure |
| LoRA per turn | Fits one GPU, quick rollback | Smaller capacity to change behaviour |
| Full fine-tune | Largest gains | Sharded training, heavier evals, harder rollback |
What to do next
- Define one event schema with model version, template hash and consent state, and write it at serving time.
- Scrub personal data before logging, and record lineage from every training example back to its events.
- Build the eval firewall first: hash and near-duplicate check every candidate prompt against all eval suites.
- Start with one high-quality pair source, such as regenerate followed by acceptance, judged on correctness.
- Size judge, training and eval GPU time with your own measured throughput and utilisation, and schedule the batch work into off-peak serving capacity.
- Write the release gate with per-slice regression limits and a safety suite, then add shadow and canary stages.
- Track turn time, pair yield, slice deltas and output diversity on a dashboard reviewed every turn.
- Keep a fixed share of verified non-model data in each turn to guard against collapse.