Human-in-the-loop (HITL) for a large language model service means that some outputs are seen by a person before or after release, and that what the person decides changes the system. The decision can change which outputs are held next time, which examples go into the evaluation set, or which data the next fine-tune learns from. It is not the same as an approval gate on an agent's actions, where a person signs off a payment or a deletion. That pattern is covered in human approval gates for autonomous agents and approval gates as a security control.
This article covers the serving side. It looks at which confidence signals the inference stack can produce and what each costs in GPU time, how to route outputs between release, hold and audit, how to size a review team with plain arithmetic, and how to turn reviewer edits into evaluation and preference data without fooling yourself. It ends with the failure modes that make HITL systems quietly stop working.
The architecture
The pipeline has five parts, shown in the diagram below. The inference server produces a draft and, if asked, per-token log probabilities. Scorers turn the draft and those numbers into one or more risk scores. A router uses the scores to release the answer, hold it for review, or release it and also copy it into a random audit sample. A review queue and console let people approve, edit or reject. A decision store keeps every judgement with enough context to reuse it later.
There are two loops with different speeds. The fast loop runs per request: a held answer waits for a person, which adds minutes of latency, so holds only suit products where that is acceptable, such as email replies, claims or document drafts. Chat products usually release first and review afterwards. The slow loop runs over weeks: decisions recalibrate thresholds and feed training data, and a new model goes through shadow evaluation before it replaces the old one.
Confidence signals and what they cost
A router needs a number. These are the common sources, from cheapest to most expensive in GPU time:
| Signal | How it is computed | Extra GPU cost | Main weakness |
|---|---|---|---|
| Rules | regex, PII detectors, policy keyword lists, schema validation | none | only catches what you predicted |
| Token log probabilities | mean or minimum logprob over the answer, or over key spans | almost none; returned with the draft | often poorly calibrated after RLHF |
| Retrieval coverage | share of answer claims supported by retrieved passages | small embedding or NLI model | only means something for RAG answers |
| Verifier or judge model | a second model scores the draft against a rubric | one extra prefill pass, often on a smaller model | shares blind spots with the generator |
| Self-consistency | sample k answers and measure agreement | about k times the decode cost | expensive; agreement on a wrong answer still happens |
OpenAI-compatible servers, including vLLM, accept a logprobs option on completions that returns per-token log probabilities. These are nearly free because they come from the softmax the server already computes. The catch is calibration. Models tuned with human feedback are often overconfident, so a high mean logprob does not mean the answer is right. Do not choose a threshold by intuition. Use logprobs as a raw feature and map them to error rates with labelled data, as described below.
The best results usually come from combining signals in stages. Run rules and logprob features on every request. Run the verifier only when they are borderline. Run self-consistency, if at all, only on the small slice the verifier cannot settle. That keeps the extra GPU cost a small fraction of serving cost while spending it where the decision is hard.
Routing: release, hold or audit
The router is short code, but each branch has a purpose:
import random, hashlib
AUDIT_RATE = 0.005 # share of auto-released traffic copied to review
def route(req_id, draft, scores, policy):
if scores["rule_block"]:
return "hold", "rule" # never auto-release a rule hit
risk = policy.calibrated_risk(scores) # P(answer is wrong), fitted on labels
if risk >= policy.hold_threshold:
return "hold", f"risk={risk:.2f}"
# Deterministic sampling: the same request id always gets the same decision,
# so retries do not change routing and audits can be replayed.
h = int(hashlib.sha256(req_id.encode()).hexdigest()[:8], 16) / 0xFFFFFFFF
if h < AUDIT_RATE:
return "release_and_audit", f"risk={risk:.2f}"
return "release", f"risk={risk:.2f}"The audit branch is the one teams skip and later regret. If people only ever see held items, you learn the error rate of risky answers and nothing about the answers you released. A small random sample of released traffic is the only unbiased estimate of the error rate users actually experience. It is also the only data that can tell you the threshold has drifted.
calibrated_risk is a small model, often logistic regression or isotonic regression, that maps raw scores to the probability of an error, using reviewed items as labels. Fit it on audit items as well as held ones, weighting each item by the inverse of the probability it had of being reviewed. Otherwise the held items dominate the fit and the curve is skewed toward risky traffic.
Sizing the review team
Review is limited by people, not by GPUs, so size it before you launch. The numbers below are illustrative. Measure your own review times in a pilot week.
| Quantity | Value | Working |
|---|---|---|
| Requests per day | 200,000 | from traffic forecast |
| Hold rate | 3% | chosen threshold |
| Held items per day | 6,000 | 200,000 x 0.03 |
| Audit items per day | 970 | 194,000 released x 0.005 |
| Median review time | 40 s | pilot measurement |
| Reviewer hours per day | about 77 | 6,970 x 40 s / 3,600 |
| Productive hours per reviewer shift | 6 | breaks, calibration sessions |
| Reviewers per day | 13 to 14 | 77 / 6, before peaks |
Traffic is not flat. If 15% of daily volume arrives in the peak hour, then 900 held items need 10 reviewer-hours that hour, so you need about ten people on shift at peak, or a queue that drains later. Little's law gives the wait: average items waiting equals arrival rate times average time in queue. With a 30-minute review SLA, the backlog must stay below about half an hour of arrivals. The threshold is a capacity control: lowering it by a few points can double the hold rate. Change it only after redoing this table.
Decide the overflow policy in advance. When the queue is over its SLA, either fail closed, holding the answer and telling the user it is being checked, or fail open, releasing it with the item still queued for review afterwards. Regulated content usually needs fail-closed. Everything else usually fails open with a temporary rise in audit rate.
Worked example: refund drafts
A support assistant drafts refund decisions. Wrong approvals cost money, so the team holds risky drafts. In a pilot, every draft was reviewed for one week, about 9,000 items, which gave a fully labelled set. The team fitted the risk model on the first five days and scored the last two to choose a threshold:
| Hold threshold on risk | Hold rate | Errors caught by hold | Error rate in released traffic |
|---|---|---|---|
| 0.50 | 1.2% | 25% | 3.0% |
| 0.30 | 3.1% | 45% | 2.3% |
| 0.15 | 7.8% | 78% | 1.0% |
| 0.05 | 21% | 96% | 0.2% |
These figures are this example's own and not a benchmark. The base error rate was about 4%. The business could fund about 3% holds, and the 0.30 threshold cut the released error rate from 4% to about 2.3%. To go further, the team added a verifier model that runs only when risk falls between 0.15 and 0.30, about 5% of traffic. It catches most of the extra errors at a fraction of the reviewer cost of moving the threshold. After launch, the 0.5% audit sample is the check. If its error rate climbs well above 2.3% for a week, the threshold or the model has drifted.
From decisions to data
Every review produces a record. Design it so it is useful for calibration, evaluation and training at once:
@dataclass
class ReviewRecord:
request_id: str
model_version: str # weights + prompt template + decoding config
route: str # "hold" | "release_and_audit"
review_prob: float # chance this item had of being reviewed (for reweighting)
prompt: str
draft: str
decision: str # "approve" | "edit" | "reject"
final: str | None # reviewer text when edited
reason_codes: list[str] # e.g. ["wrong_policy", "tone", "hallucinated_fact"]
reviewer_id: str
seconds_spent: floatThree uses follow from it. Audit records, reweighted by review_prob, calibrate the risk model. A time-based slice, such as the last two weeks, held out from training, becomes the evaluation set for the next model. Edits become preference pairs, with the reviewer's text as the chosen answer and the draft as the rejected one, ready for DPO or reward-model training.
Be careful with those pairs. They come mostly from held items, so they over-represent hard prompts, which helps but is not the real distribution. A reviewer edit is often only one acceptable answer among several, so send about 5% of items to two reviewers and track agreement, using Cohen's kappa for categorical decisions. If two people often disagree, the rubric needs fixing before the data can train anything. Keep prompts that appear in the evaluation set out of the training pairs, and check that your data policy allows training on customer content at all.
Where the GPU time goes
Most GPU cost here is outside the request path. The verifier is the largest per-request cost. A model several times smaller than the generator can share its GPUs, since it only needs a prefill pass. Running it as a separate deployment keeps its load from affecting generation latency. The audit and calibration jobs are cheap batch work that can run on spare capacity.
The training side runs in batches. Collect pairs for a week or two, run a LoRA or full fine-tune, then replay the held-out reviewed set through both models offline and compare error rates by reason code. Promote only if the new model is better on that set and no worse on a general regression suite. After promotion, refit the calibration model, because a new model's logprobs carry a different meaning.
Failure modes
- Automation bias. Reviewers who mostly see correct drafts start approving everything. Mix in known-bad seeded items, about 1 to 2% of the queue, and track each reviewer's catch rate.
- No audit sample. Without one, released-traffic error is unknown and threshold drift goes unnoticed.
- Stale calibration after a model change. The same threshold on a new model can double or halve the hold rate in a day. Refit and redo the capacity table at every promotion.
- Queue overflow with no policy. Holds silently become timeouts. Monitor queue age, not only queue length.
- Training on your own outputs. Approved drafts used as chosen examples teach the model to repeat itself. Prefer edits and rejections, and cap the share of approved-only data.
- Reviewer data leakage. Consoles show customer content. Apply the same access controls, logging and retention to them as to production data.
What to do next
- Write down what a wrong answer costs in your product, and decide whether holds are acceptable or review must happen after release.
- Turn on log probabilities in your serving stack, log them with every response, and add rule-based scorers.
- Run a pilot week reviewing everything, or a large random sample, to get unbiased labels and real review times.
- Fit a calibrated risk model, then choose the threshold from a table like the worked example and check it against reviewer capacity.
- Ship the router with a deterministic audit sample, a queue-age alert and a written overflow policy.
- Store review records with model version and review probability, and set a regular cadence for evaluation and preference-pair export. Use active learning ideas to choose what to label, labelling practice for rubrics, and LLM cost analysis to price the verifier.