Human-in-the-loop (HITL) for a large language model service means that some outputs are seen by a person before or after release, and that what the person decides changes the system. The decision can change which outputs are held next time, which examples go into the evaluation set, or which data the next fine-tune learns from. It is not the same as an approval gate on an agent's actions, where a person signs off a payment or a deletion. That pattern is covered in human approval gates for autonomous agents and approval gates as a security control.

This article covers the serving side. It looks at which confidence signals the inference stack can produce and what each costs in GPU time, how to route outputs between release, hold and audit, how to size a review team with plain arithmetic, and how to turn reviewer edits into evaluation and preference data without fooling yourself. It ends with the failure modes that make HITL systems quietly stop working.

The architecture

The pipeline has five parts, shown in the diagram below. The inference server produces a draft and, if asked, per-token log probabilities. Scorers turn the draft and those numbers into one or more risk scores. A router uses the scores to release the answer, hold it for review, or release it and also copy it into a random audit sample. A review queue and console let people approve, edit or reject. A decision store keeps every judgement with enough context to reuse it later.

Human review on the serving path: score, route, review, and feed decisions backRequestprompt + contextInference serverdraft + token logprobsScorerslogprob, rules, verifierRouterrelease, hold, auditReleasedto user, loggedscore highReview queuepriority + SLAholdrandom auditReviewer consoleapprove, edit, rejectRelease held itemapproved or editedDecision storelabels, edits, reasonsThreshold calibrationfrom audit sampleEval setheld out by timePreference pairsedit over draftFine-tune + shadowGPU, offlineOnly audited and held items are labelled; the audit sample is what keeps the error estimate honest.
Scoring and routing sit between the model and the user. Reviewer decisions feed threshold calibration, the eval set and preference data for the next fine-tune.

There are two loops with different speeds. The fast loop runs per request: a held answer waits for a person, which adds minutes of latency, so holds only suit products where that is acceptable, such as email replies, claims or document drafts. Chat products usually release first and review afterwards. The slow loop runs over weeks: decisions recalibrate thresholds and feed training data, and a new model goes through shadow evaluation before it replaces the old one.

Confidence signals and what they cost

A router needs a number. These are the common sources, from cheapest to most expensive in GPU time:

SignalHow it is computedExtra GPU costMain weakness
Rulesregex, PII detectors, policy keyword lists, schema validationnoneonly catches what you predicted
Token log probabilitiesmean or minimum logprob over the answer, or over key spansalmost none; returned with the draftoften poorly calibrated after RLHF
Retrieval coverageshare of answer claims supported by retrieved passagessmall embedding or NLI modelonly means something for RAG answers
Verifier or judge modela second model scores the draft against a rubricone extra prefill pass, often on a smaller modelshares blind spots with the generator
Self-consistencysample k answers and measure agreementabout k times the decode costexpensive; agreement on a wrong answer still happens

OpenAI-compatible servers, including vLLM, accept a logprobs option on completions that returns per-token log probabilities. These are nearly free because they come from the softmax the server already computes. The catch is calibration. Models tuned with human feedback are often overconfident, so a high mean logprob does not mean the answer is right. Do not choose a threshold by intuition. Use logprobs as a raw feature and map them to error rates with labelled data, as described below.

The best results usually come from combining signals in stages. Run rules and logprob features on every request. Run the verifier only when they are borderline. Run self-consistency, if at all, only on the small slice the verifier cannot settle. That keeps the extra GPU cost a small fraction of serving cost while spending it where the decision is hard.

Routing: release, hold or audit

The router is short code, but each branch has a purpose:

import random, hashlib

AUDIT_RATE = 0.005          # share of auto-released traffic copied to review

def route(req_id, draft, scores, policy):
    if scores["rule_block"]:
        return "hold", "rule"                       # never auto-release a rule hit
    risk = policy.calibrated_risk(scores)          # P(answer is wrong), fitted on labels
    if risk >= policy.hold_threshold:
        return "hold", f"risk={risk:.2f}"
    # Deterministic sampling: the same request id always gets the same decision,
    # so retries do not change routing and audits can be replayed.
    h = int(hashlib.sha256(req_id.encode()).hexdigest()[:8], 16) / 0xFFFFFFFF
    if h < AUDIT_RATE:
        return "release_and_audit", f"risk={risk:.2f}"
    return "release", f"risk={risk:.2f}"

The audit branch is the one teams skip and later regret. If people only ever see held items, you learn the error rate of risky answers and nothing about the answers you released. A small random sample of released traffic is the only unbiased estimate of the error rate users actually experience. It is also the only data that can tell you the threshold has drifted.

calibrated_risk is a small model, often logistic regression or isotonic regression, that maps raw scores to the probability of an error, using reviewed items as labels. Fit it on audit items as well as held ones, weighting each item by the inverse of the probability it had of being reviewed. Otherwise the held items dominate the fit and the curve is skewed toward risky traffic.

Sizing the review team

Review is limited by people, not by GPUs, so size it before you launch. The numbers below are illustrative. Measure your own review times in a pilot week.

QuantityValueWorking
Requests per day200,000from traffic forecast
Hold rate3%chosen threshold
Held items per day6,000200,000 x 0.03
Audit items per day970194,000 released x 0.005
Median review time40 spilot measurement
Reviewer hours per dayabout 776,970 x 40 s / 3,600
Productive hours per reviewer shift6breaks, calibration sessions
Reviewers per day13 to 1477 / 6, before peaks

Traffic is not flat. If 15% of daily volume arrives in the peak hour, then 900 held items need 10 reviewer-hours that hour, so you need about ten people on shift at peak, or a queue that drains later. Little's law gives the wait: average items waiting equals arrival rate times average time in queue. With a 30-minute review SLA, the backlog must stay below about half an hour of arrivals. The threshold is a capacity control: lowering it by a few points can double the hold rate. Change it only after redoing this table.

Decide the overflow policy in advance. When the queue is over its SLA, either fail closed, holding the answer and telling the user it is being checked, or fail open, releasing it with the item still queued for review afterwards. Regulated content usually needs fail-closed. Everything else usually fails open with a temporary rise in audit rate.

Worked example: refund drafts

A support assistant drafts refund decisions. Wrong approvals cost money, so the team holds risky drafts. In a pilot, every draft was reviewed for one week, about 9,000 items, which gave a fully labelled set. The team fitted the risk model on the first five days and scored the last two to choose a threshold:

Hold threshold on riskHold rateErrors caught by holdError rate in released traffic
0.501.2%25%3.0%
0.303.1%45%2.3%
0.157.8%78%1.0%
0.0521%96%0.2%

These figures are this example's own and not a benchmark. The base error rate was about 4%. The business could fund about 3% holds, and the 0.30 threshold cut the released error rate from 4% to about 2.3%. To go further, the team added a verifier model that runs only when risk falls between 0.15 and 0.30, about 5% of traffic. It catches most of the extra errors at a fraction of the reviewer cost of moving the threshold. After launch, the 0.5% audit sample is the check. If its error rate climbs well above 2.3% for a week, the threshold or the model has drifted.

From decisions to data

Every review produces a record. Design it so it is useful for calibration, evaluation and training at once:

@dataclass
class ReviewRecord:
    request_id: str
    model_version: str          # weights + prompt template + decoding config
    route: str                  # "hold" | "release_and_audit"
    review_prob: float          # chance this item had of being reviewed (for reweighting)
    prompt: str
    draft: str
    decision: str               # "approve" | "edit" | "reject"
    final: str | None           # reviewer text when edited
    reason_codes: list[str]     # e.g. ["wrong_policy", "tone", "hallucinated_fact"]
    reviewer_id: str
    seconds_spent: float

Three uses follow from it. Audit records, reweighted by review_prob, calibrate the risk model. A time-based slice, such as the last two weeks, held out from training, becomes the evaluation set for the next model. Edits become preference pairs, with the reviewer's text as the chosen answer and the draft as the rejected one, ready for DPO or reward-model training.

Be careful with those pairs. They come mostly from held items, so they over-represent hard prompts, which helps but is not the real distribution. A reviewer edit is often only one acceptable answer among several, so send about 5% of items to two reviewers and track agreement, using Cohen's kappa for categorical decisions. If two people often disagree, the rubric needs fixing before the data can train anything. Keep prompts that appear in the evaluation set out of the training pairs, and check that your data policy allows training on customer content at all.

Where the GPU time goes

Most GPU cost here is outside the request path. The verifier is the largest per-request cost. A model several times smaller than the generator can share its GPUs, since it only needs a prefill pass. Running it as a separate deployment keeps its load from affecting generation latency. The audit and calibration jobs are cheap batch work that can run on spare capacity.

The training side runs in batches. Collect pairs for a week or two, run a LoRA or full fine-tune, then replay the held-out reviewed set through both models offline and compare error rates by reason code. Promote only if the new model is better on that set and no worse on a general regression suite. After promotion, refit the calibration model, because a new model's logprobs carry a different meaning.

Failure modes

  • Automation bias. Reviewers who mostly see correct drafts start approving everything. Mix in known-bad seeded items, about 1 to 2% of the queue, and track each reviewer's catch rate.
  • No audit sample. Without one, released-traffic error is unknown and threshold drift goes unnoticed.
  • Stale calibration after a model change. The same threshold on a new model can double or halve the hold rate in a day. Refit and redo the capacity table at every promotion.
  • Queue overflow with no policy. Holds silently become timeouts. Monitor queue age, not only queue length.
  • Training on your own outputs. Approved drafts used as chosen examples teach the model to repeat itself. Prefer edits and rejections, and cap the share of approved-only data.
  • Reviewer data leakage. Consoles show customer content. Apply the same access controls, logging and retention to them as to production data.

What to do next

  1. Write down what a wrong answer costs in your product, and decide whether holds are acceptable or review must happen after release.
  2. Turn on log probabilities in your serving stack, log them with every response, and add rule-based scorers.
  3. Run a pilot week reviewing everything, or a large random sample, to get unbiased labels and real review times.
  4. Fit a calibrated risk model, then choose the threshold from a table like the worked example and check it against reviewer capacity.
  5. Ship the router with a deterministic audit sample, a queue-age alert and a written overflow policy.
  6. Store review records with model version and review probability, and set a regular cadence for evaluation and preference-pair export. Use active learning ideas to choose what to label, labelling practice for rubrics, and LLM cost analysis to price the verifier.
Key takeaway: LLM human-in-the-loop on the serving side comes down to scoring, routing, reviewing and learning. Use cheap signals such as rules and logprobs on everything and expensive ones such as a verifier or self-consistency only when the cheap ones are borderline. Calibrate risk on labelled data and always keep a random audit sample. Size the review team from hold rate and review time. Use reviewer edits for evaluation and preference data, keeping their selection bias in mind.