An LLM annotation tool is the software that sits between a pool of raw text and a training run: it shows a person a prompt and model outputs, asks a structured question, records the answer with who gave it and when, and exports the result in a format a fine-tuning job can read. Label Studio, Argilla and Prodigy are the best known tools. The tool decides how many usable examples you get per annotator hour, how much of your model's own bias leaks into the labels, and whether you can later explain where a training row came from.
This article treats the annotation tool as part of the GPU pipeline rather than as a standalone app. GPUs show up on both sides of it: a served model drafts suggestions so that humans edit instead of write, and the accepted labels feed SFT, DPO or reward-model training. We cover the task types that matter for LLM data, the reference architecture, how the common tools differ, a concrete Argilla setup with GPU-generated suggestions, agreement measurement, export to training formats, a worked budget, failure modes and a checklist. Bulk labeling where the LLM itself is the labeler is covered separately in LLM data labeling; here a human stays in the loop.
The task shapes LLM data needs
LLM data work reduces to a handful of task shapes, and the tool must express each one without hacks:
- Demonstration writing or correction (SFT). Show a prompt and a draft answer; the annotator edits the draft until it is what the model should have said. Output: prompt plus final text. Editing a good draft takes a fraction of the time of writing from scratch, which is the main reason pre-annotation exists.
- Pairwise or ranked preference (DPO, reward models). Show a prompt and two to four responses; the annotator ranks them or picks a winner, ideally with a reason code. Output: prompt, chosen, rejected. See preference data collection for interface design in depth.
- Rubric rating. Score a response on separate axes such as correctness, helpfulness, safety and format, usually on a 1-5 scale. Used for evaluation sets and for filtering SFT data.
- Span labeling. Mark the substring that is a hallucinated claim, leaked PII or an unsafe instruction. Output: character offsets plus a label. Useful for token-level reward and safety classifiers.
Reference architecture
Every serious setup has the same six parts, whatever the brand of tool. A source pool holds candidate prompts: production logs that have been scrubbed of personal data, synthetic prompts, or documents. A pre-annotation step runs one or more models over the sample on GPUs and attaches suggestions with confidence scores. The task store holds records (fields to display, questions to answer, suggestions, metadata) and decides who sees what. The annotator UI collects responses. A QA layer computes agreement on overlapping items, scores annotators against hidden gold items and routes disagreements to a reviewer. Finally an export step freezes accepted records into a versioned snapshot that a training job consumes. The loop closes when the new checkpoint becomes the pre-annotation model for the next round.
Choosing a tool
The three common tools make different bets. The table reflects their documented positioning; check current licensing and feature lists before committing, because all three evolve.
| Tool | Model | Strengths for LLM data | Watch out for |
|---|---|---|---|
| Label Studio (HumanSignal) | Open-source community edition plus a commercial enterprise edition; self-host | Very general labeling configs (text, spans, choices, ranking, images, audio); predictions can be imported with tasks or served by an ML backend | LLM-specific views such as chat transcripts and pairwise comparison need custom label configs |
| Argilla | Open-source; Python SDK drives everything; Hugging Face acquired the company in 2024 | Built around LLM feedback: text fields, label, rating, ranking, span and free-text questions, suggestions with scores and agents, task distribution with minimum submissions, export to Hugging Face datasets | Server plus search backend to operate; you write the workflow code yourself |
| Prodigy (Explosion) | Paid licence, runs locally, scriptable recipes | Fast single-user annotation, recipes chain model-in-the-loop steps, good for small expert teams | Built for a few expert annotators rather than a large distributed workforce |
| In-house app | Your code | Exact fit to an unusual task, integrated with internal auth and data stores | You now own agreement metrics, assignment, audit logs and export forever |
Pick the tool whose data model already has your task shapes, then measure seconds per accepted item on a 200-record pilot before scaling.
Setting up a preference task
Here is a pairwise preference task in Argilla's 2.x Python SDK, with two GPU-generated responses per prompt and a model suggestion attached. The class names Settings, TextField, LabelQuestion, RatingQuestion, TextQuestion, TaskDistribution, Record and Suggestion come from the Argilla documentation; the server URL, workspace and question names are placeholders.
import argilla as rg
client = rg.Argilla(api_url="https://argilla.internal", api_key=API_KEY)
settings = rg.Settings(
guidelines="Pick the response a careful expert would send. Correctness beats style.",
fields=[
rg.TextField(name="prompt"),
rg.TextField(name="response_a"),
rg.TextField(name="response_b"),
],
questions=[
rg.LabelQuestion(name="preference", labels=["a", "b", "tie", "both_bad"]),
rg.RatingQuestion(name="confidence", values=[1, 2, 3]),
rg.TextQuestion(name="reason", required=False),
],
distribution=rg.TaskDistribution(min_submitted=2), # two humans per record
)
dataset = rg.Dataset(name="pref_round_07", workspace="alignment",
settings=settings, client=client)
dataset.create()Two design choices in those lines matter more than they look. The both_bad label gives annotators an honest exit; without it they pick the less bad answer and you train the model towards a bad response. And min_submitted=2 buys you an agreement measurement on every record at double the cost; a common compromise is two annotators on a 15-25% overlap slice and one on the rest.
GPU pre-annotation
Pre-annotation is where the GPU earns its keep. You run the current policy (and often a stronger judge) over the sampled prompts in an offline batch, generate the candidate responses, and ask the judge for a preference with a probability. A batch engine such as vLLM keeps the GPUs saturated; see vLLM on GPU for how its scheduler and KV cache behave under this kind of load. The judge's output becomes a suggestion, not a label:
def to_record(item, judge):
# judge = {"winner": "a", "p_winner": 0.81, "model": "judge-v3"}
return rg.Record(
id=item["prompt_id"],
fields={"prompt": item["prompt"],
"response_a": item["resp_a"],
"response_b": item["resp_b"]},
metadata={"source": item["source"], "policy_ckpt": item["ckpt"]},
suggestions=[rg.Suggestion("preference", judge["winner"],
score=judge["p_winner"], agent=judge["model"])],
)
dataset.records.log([to_record(i, j) for i, j in zip(items, judgements)])Label Studio expresses the same idea by importing tasks that already carry a predictions array (each with a model_version, a score and a result in the label config's format), or by connecting an ML backend that returns predictions on demand. Either way, three rules keep pre-annotation honest. Randomize the A/B order before both generation and display so that position bias does not line up with the suggestion. Hide the suggestion on a random control slice (10% is plenty) so you can measure how much it anchors people. And record the agent name and checkpoint, because a suggestion from checkpoint 6 shaping labels that train checkpoint 7 is a feedback loop you will want to audit.
The arithmetic favours spending GPU time: two 400-token responses for 20,000 prompts is about 16 million generated tokens, hours of batch inference for a 7-8B model, while the same items at 45 human seconds each are 250 annotator hours.
Choosing what humans see
Not every prompt deserves a human. Active learning picks the items where a label changes the model most. Two cheap signals do most of the work. Uncertainty: items where the judge's probability is near 0.5, or where the policy's two samples disagree on a verifiable answer. Diversity: embed the prompts on GPU, cluster them, and cap how many items any one cluster contributes so the batch is not 3,000 variations of the same coding question.
import numpy as np
from sklearn.cluster import KMeans
def select_batch(emb, p_judge, budget, n_clusters=200, per_cluster=None):
unc = 1.0 - np.abs(p_judge - 0.5) * 2 # 1 at p=0.5, 0 at p in {0,1}
labels = KMeans(n_clusters=n_clusters, n_init="auto").fit_predict(emb)
per_cluster = per_cluster or max(1, budget // n_clusters * 2)
chosen = []
for k in range(n_clusters):
idx = np.where(labels == k)[0]
chosen.extend(idx[np.argsort(-unc[idx])][:per_cluster])
chosen = np.array(chosen)
return chosen[np.argsort(-unc[chosen])][:budget]Keep 10-20% of each batch uniformly random, or the training set stops looking like real traffic.
Agreement and quality control
Agreement is the measurement that tells you whether the label means anything. For two annotators on a categorical question, Cohen's kappa corrects raw agreement for the agreement you would get by chance: kappa = (p_o - p_e) / (1 - p_e), where p_o is observed agreement and p_e is the agreement expected from each annotator's label frequencies. With more than two annotators or missing responses, use Krippendorff's alpha. A minimal kappa:
from collections import Counter
def cohen_kappa(a, b):
n = len(a)
p_o = sum(x == y for x, y in zip(a, b)) / n
ca, cb = Counter(a), Counter(b)
p_e = sum(ca[k] * cb[k] for k in set(ca) | set(cb)) / (n * n)
return (p_o - p_e) / (1 - p_e) if p_e < 1 else 1.0Interpretation is task-dependent, but for pairwise preference on general chat, kappa in the 0.3-0.5 range is common and above 0.6 is unusually clean. A low number is a signal about the guidelines before it is a signal about people: read 30 disagreements and you will usually find an ambiguous rule. Add hidden gold items (5% of the queue, answers fixed by an expert), track each annotator's gold accuracy over time, and route records where the two humans disagree, or where a human overruled a high-confidence suggestion, to a reviewer.
Exporting training data
The export step turns responses into rows a trainer reads. Two rules prevent most damage: split train and eval by prompt, never by row, so near-duplicate prompts cannot straddle the split; and stamp every snapshot with a content hash, the guideline version and the pre-annotation checkpoint, so any later model regression can be traced to a data round.
import hashlib, json
def to_dpo_rows(records):
rows = []
for r in records:
votes = [resp.value for resp in r.responses["preference"]]
if len(votes) < 2 or votes[0] != votes[1] or votes[0] in ("tie", "both_bad"):
continue # keep only agreed, decisive pairs
win = votes[0]
chosen = r.fields["response_" + win]
rejected = r.fields["response_" + ("b" if win == "a" else "a")]
rows.append({"prompt": r.fields["prompt"], "chosen": chosen,
"rejected": rejected, "prompt_id": r.id})
return rows
def write_snapshot(rows, path, meta):
blob = "\n".join(json.dumps(x, sort_keys=True) for x in rows)
meta["sha256"] = hashlib.sha256(blob.encode()).hexdigest()
open(path, "w").write(blob)
json.dump(meta, open(path + ".meta.json", "w"), indent=2)The resulting prompt / chosen / rejected JSONL is the shape most DPO trainers accept; DPO training on GPU covers what happens to it next, including the memory cost of the reference model. Records with both_bad are not wasted: they are the best source of prompts for the next SFT correction round.
Worked example: a DPO round
Worked example. A team wants 15,000 clean preference pairs for a DPO round on an 8B model. From a scrubbed pool of 400,000 production prompts they embed everything (a few GPU minutes), then generate two responses per prompt for a 60,000-prompt candidate set and have a 70B judge score each pair. Active learning picks 24,000 records: 20,000 by uncertainty with a per-cluster cap and 4,000 uniformly at random.
A pilot of 200 records shows 38 seconds per item with suggestions and 61 seconds without, and agreement with the suggestion is 7 points higher when it is shown than on the hidden control slice, a measurable anchoring effect they decide to accept and monitor. With 20% double annotation, total human time is 24,000 x 1.2 x 38 s, about 304 hours. After review, 18% of records are ties or both_bad and 9% of the doubly annotated slice disagrees; 17,900 pairs survive, comfortably above target. The snapshot meta records guideline v4, judge checkpoint and policy checkpoint, so when the next model gets oddly verbose they can query exactly which pairs preferred the longer answer.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Anchoring on suggestions | Human-suggestion agreement far above control-slice agreement | Hidden control slice, show suggestions after a first click, or only for long edits |
| Length bias | Chosen responses are systematically longer | Length-matched pairs, explicit guideline, monitor chosen/rejected length ratio |
| Position bias | A wins more than 50% on random pairs | Randomize order at display time, log the shown order |
| Guideline drift | Kappa drops after a guideline edit | Version guidelines, re-run gold items, never mix versions in one snapshot unmarked |
| Leaked eval prompts | Eval scores jump after a round | Hash and dedupe prompts against every eval set before annotation |
| PII in source logs | Annotators see customer data | Scrub before sampling, restrict workspaces, log access |
| Silent export bugs | Chosen and rejected swapped for one class | Unit-test the exporter on hand-built records; spot-check 50 rows per snapshot |
Trade-offs
The central trade-off is speed versus independence. Suggestions cut time per item but pull labels towards the model that made them; lean on them for style and format, and prefer blind first passes for factual correctness and safety. Double annotation doubles cost but is the only way to know your noise rate. Human labels cost far more than LLM-judge labels, but they are what you calibrate the judge against, so even a judge-heavy pipeline keeps a human slice. Online DPO shows how that slice is used when pairs are generated every training step.
What to do next
- Write down the task shapes you need (SFT edit, pairwise, rating, span) and pick the tool whose data model already has them.
- Run a 200-record pilot; measure seconds per accepted item with and without suggestions.
- Add a hidden control slice and 5% gold items before scaling past the pilot.
- Generate suggestions on GPU in an offline batch; log agent name and checkpoint on every suggestion.
- Use uncertainty plus cluster caps to pick items, and keep 10-20% uniformly random.
- Double-annotate at least 15% and compute kappa or alpha per question every day.
- Export with prompt-level splits, eval-set dedupe and a hashed, versioned snapshot.
- Feed both_bad records back into the next SFT correction round.