An LLM annotation tool is the software that sits between a pool of raw text and a training run: it shows a person a prompt and model outputs, asks a structured question, records the answer with who gave it and when, and exports the result in a format a fine-tuning job can read. Label Studio, Argilla and Prodigy are the best known tools. The tool decides how many usable examples you get per annotator hour, how much of your model's own bias leaks into the labels, and whether you can later explain where a training row came from.

This article treats the annotation tool as part of the GPU pipeline rather than as a standalone app. GPUs show up on both sides of it: a served model drafts suggestions so that humans edit instead of write, and the accepted labels feed SFT, DPO or reward-model training. We cover the task types that matter for LLM data, the reference architecture, how the common tools differ, a concrete Argilla setup with GPU-generated suggestions, agreement measurement, export to training formats, a worked budget, failure modes and a checklist. Bulk labeling where the LLM itself is the labeler is covered separately in LLM data labeling; here a human stays in the loop.

The task shapes LLM data needs

LLM data work reduces to a handful of task shapes, and the tool must express each one without hacks:

  • Demonstration writing or correction (SFT). Show a prompt and a draft answer; the annotator edits the draft until it is what the model should have said. Output: prompt plus final text. Editing a good draft takes a fraction of the time of writing from scratch, which is the main reason pre-annotation exists.
  • Pairwise or ranked preference (DPO, reward models). Show a prompt and two to four responses; the annotator ranks them or picks a winner, ideally with a reason code. Output: prompt, chosen, rejected. See preference data collection for interface design in depth.
  • Rubric rating. Score a response on separate axes such as correctness, helpfulness, safety and format, usually on a 1-5 scale. Used for evaluation sets and for filtering SFT data.
  • Span labeling. Mark the substring that is a hallucinated claim, leaked PII or an unsafe instruction. Output: character offsets plus a label. Useful for token-level reward and safety classifiers.

Reference architecture

Every serious setup has the same six parts, whatever the brand of tool. A source pool holds candidate prompts: production logs that have been scrubbed of personal data, synthetic prompts, or documents. A pre-annotation step runs one or more models over the sample on GPUs and attaches suggestions with confidence scores. The task store holds records (fields to display, questions to answer, suggestions, metadata) and decides who sees what. The annotator UI collects responses. A QA layer computes agreement on overlapping items, scores annotators against hidden gold items and routes disagreements to a reviewer. Finally an export step freezes accepted records into a versioned snapshot that a training job consumes. The loop closes when the new checkpoint becomes the pre-annotation model for the next round.

Annotation loop for LLM training dataSource poolprompts, logs, docssamplePre-annotationGPU batch inferencesuggestionsTask storefields, questions, queueassignAnnotator UIaccept, edit, rankresponsesQA and agreementoverlap, gold, reviewacceptedExport snapshotSFT / DPO JSONL + hashtrainGPU trainingSFT, DPO, reward modelnew checkpointActive learning: uncertainty and diversity scores decide which pool items are worth a human's minuteThe GPU appears twice: cheaply drafting labels before humans see them, and expensively consuming the labels afterwards.
Pre-annotation and training both run on GPUs; the annotation tool is the bottleneck in between, measured in human minutes.

Choosing a tool

The three common tools make different bets. The table reflects their documented positioning; check current licensing and feature lists before committing, because all three evolve.

ToolModelStrengths for LLM dataWatch out for
Label Studio (HumanSignal)Open-source community edition plus a commercial enterprise edition; self-hostVery general labeling configs (text, spans, choices, ranking, images, audio); predictions can be imported with tasks or served by an ML backendLLM-specific views such as chat transcripts and pairwise comparison need custom label configs
ArgillaOpen-source; Python SDK drives everything; Hugging Face acquired the company in 2024Built around LLM feedback: text fields, label, rating, ranking, span and free-text questions, suggestions with scores and agents, task distribution with minimum submissions, export to Hugging Face datasetsServer plus search backend to operate; you write the workflow code yourself
Prodigy (Explosion)Paid licence, runs locally, scriptable recipesFast single-user annotation, recipes chain model-in-the-loop steps, good for small expert teamsBuilt for a few expert annotators rather than a large distributed workforce
In-house appYour codeExact fit to an unusual task, integrated with internal auth and data storesYou now own agreement metrics, assignment, audit logs and export forever

Pick the tool whose data model already has your task shapes, then measure seconds per accepted item on a 200-record pilot before scaling.

Setting up a preference task

Here is a pairwise preference task in Argilla's 2.x Python SDK, with two GPU-generated responses per prompt and a model suggestion attached. The class names Settings, TextField, LabelQuestion, RatingQuestion, TextQuestion, TaskDistribution, Record and Suggestion come from the Argilla documentation; the server URL, workspace and question names are placeholders.

import argilla as rg

client = rg.Argilla(api_url="https://argilla.internal", api_key=API_KEY)

settings = rg.Settings(
    guidelines="Pick the response a careful expert would send. Correctness beats style.",
    fields=[
        rg.TextField(name="prompt"),
        rg.TextField(name="response_a"),
        rg.TextField(name="response_b"),
    ],
    questions=[
        rg.LabelQuestion(name="preference", labels=["a", "b", "tie", "both_bad"]),
        rg.RatingQuestion(name="confidence", values=[1, 2, 3]),
        rg.TextQuestion(name="reason", required=False),
    ],
    distribution=rg.TaskDistribution(min_submitted=2),   # two humans per record
)
dataset = rg.Dataset(name="pref_round_07", workspace="alignment",
                     settings=settings, client=client)
dataset.create()

Two design choices in those lines matter more than they look. The both_bad label gives annotators an honest exit; without it they pick the less bad answer and you train the model towards a bad response. And min_submitted=2 buys you an agreement measurement on every record at double the cost; a common compromise is two annotators on a 15-25% overlap slice and one on the rest.

GPU pre-annotation

Pre-annotation is where the GPU earns its keep. You run the current policy (and often a stronger judge) over the sampled prompts in an offline batch, generate the candidate responses, and ask the judge for a preference with a probability. A batch engine such as vLLM keeps the GPUs saturated; see vLLM on GPU for how its scheduler and KV cache behave under this kind of load. The judge's output becomes a suggestion, not a label:

def to_record(item, judge):
    # judge = {"winner": "a", "p_winner": 0.81, "model": "judge-v3"}
    return rg.Record(
        id=item["prompt_id"],
        fields={"prompt": item["prompt"],
                "response_a": item["resp_a"],
                "response_b": item["resp_b"]},
        metadata={"source": item["source"], "policy_ckpt": item["ckpt"]},
        suggestions=[rg.Suggestion("preference", judge["winner"],
                                   score=judge["p_winner"], agent=judge["model"])],
    )

dataset.records.log([to_record(i, j) for i, j in zip(items, judgements)])

Label Studio expresses the same idea by importing tasks that already carry a predictions array (each with a model_version, a score and a result in the label config's format), or by connecting an ML backend that returns predictions on demand. Either way, three rules keep pre-annotation honest. Randomize the A/B order before both generation and display so that position bias does not line up with the suggestion. Hide the suggestion on a random control slice (10% is plenty) so you can measure how much it anchors people. And record the agent name and checkpoint, because a suggestion from checkpoint 6 shaping labels that train checkpoint 7 is a feedback loop you will want to audit.

The arithmetic favours spending GPU time: two 400-token responses for 20,000 prompts is about 16 million generated tokens, hours of batch inference for a 7-8B model, while the same items at 45 human seconds each are 250 annotator hours.

Choosing what humans see

Not every prompt deserves a human. Active learning picks the items where a label changes the model most. Two cheap signals do most of the work. Uncertainty: items where the judge's probability is near 0.5, or where the policy's two samples disagree on a verifiable answer. Diversity: embed the prompts on GPU, cluster them, and cap how many items any one cluster contributes so the batch is not 3,000 variations of the same coding question.

import numpy as np
from sklearn.cluster import KMeans

def select_batch(emb, p_judge, budget, n_clusters=200, per_cluster=None):
    unc = 1.0 - np.abs(p_judge - 0.5) * 2          # 1 at p=0.5, 0 at p in {0,1}
    labels = KMeans(n_clusters=n_clusters, n_init="auto").fit_predict(emb)
    per_cluster = per_cluster or max(1, budget // n_clusters * 2)
    chosen = []
    for k in range(n_clusters):
        idx = np.where(labels == k)[0]
        chosen.extend(idx[np.argsort(-unc[idx])][:per_cluster])
    chosen = np.array(chosen)
    return chosen[np.argsort(-unc[chosen])][:budget]

Keep 10-20% of each batch uniformly random, or the training set stops looking like real traffic.

Agreement and quality control

Agreement is the measurement that tells you whether the label means anything. For two annotators on a categorical question, Cohen's kappa corrects raw agreement for the agreement you would get by chance: kappa = (p_o - p_e) / (1 - p_e), where p_o is observed agreement and p_e is the agreement expected from each annotator's label frequencies. With more than two annotators or missing responses, use Krippendorff's alpha. A minimal kappa:

from collections import Counter

def cohen_kappa(a, b):
    n = len(a)
    p_o = sum(x == y for x, y in zip(a, b)) / n
    ca, cb = Counter(a), Counter(b)
    p_e = sum(ca[k] * cb[k] for k in set(ca) | set(cb)) / (n * n)
    return (p_o - p_e) / (1 - p_e) if p_e < 1 else 1.0

Interpretation is task-dependent, but for pairwise preference on general chat, kappa in the 0.3-0.5 range is common and above 0.6 is unusually clean. A low number is a signal about the guidelines before it is a signal about people: read 30 disagreements and you will usually find an ambiguous rule. Add hidden gold items (5% of the queue, answers fixed by an expert), track each annotator's gold accuracy over time, and route records where the two humans disagree, or where a human overruled a high-confidence suggestion, to a reviewer.

Exporting training data

The export step turns responses into rows a trainer reads. Two rules prevent most damage: split train and eval by prompt, never by row, so near-duplicate prompts cannot straddle the split; and stamp every snapshot with a content hash, the guideline version and the pre-annotation checkpoint, so any later model regression can be traced to a data round.

import hashlib, json

def to_dpo_rows(records):
    rows = []
    for r in records:
        votes = [resp.value for resp in r.responses["preference"]]
        if len(votes) < 2 or votes[0] != votes[1] or votes[0] in ("tie", "both_bad"):
            continue                                   # keep only agreed, decisive pairs
        win = votes[0]
        chosen = r.fields["response_" + win]
        rejected = r.fields["response_" + ("b" if win == "a" else "a")]
        rows.append({"prompt": r.fields["prompt"], "chosen": chosen,
                     "rejected": rejected, "prompt_id": r.id})
    return rows

def write_snapshot(rows, path, meta):
    blob = "\n".join(json.dumps(x, sort_keys=True) for x in rows)
    meta["sha256"] = hashlib.sha256(blob.encode()).hexdigest()
    open(path, "w").write(blob)
    json.dump(meta, open(path + ".meta.json", "w"), indent=2)

The resulting prompt / chosen / rejected JSONL is the shape most DPO trainers accept; DPO training on GPU covers what happens to it next, including the memory cost of the reference model. Records with both_bad are not wasted: they are the best source of prompts for the next SFT correction round.

Worked example: a DPO round

Worked example. A team wants 15,000 clean preference pairs for a DPO round on an 8B model. From a scrubbed pool of 400,000 production prompts they embed everything (a few GPU minutes), then generate two responses per prompt for a 60,000-prompt candidate set and have a 70B judge score each pair. Active learning picks 24,000 records: 20,000 by uncertainty with a per-cluster cap and 4,000 uniformly at random.

A pilot of 200 records shows 38 seconds per item with suggestions and 61 seconds without, and agreement with the suggestion is 7 points higher when it is shown than on the hidden control slice, a measurable anchoring effect they decide to accept and monitor. With 20% double annotation, total human time is 24,000 x 1.2 x 38 s, about 304 hours. After review, 18% of records are ties or both_bad and 9% of the doubly annotated slice disagrees; 17,900 pairs survive, comfortably above target. The snapshot meta records guideline v4, judge checkpoint and policy checkpoint, so when the next model gets oddly verbose they can query exactly which pairs preferred the longer answer.

Failure modes

FailureSymptomFix
Anchoring on suggestionsHuman-suggestion agreement far above control-slice agreementHidden control slice, show suggestions after a first click, or only for long edits
Length biasChosen responses are systematically longerLength-matched pairs, explicit guideline, monitor chosen/rejected length ratio
Position biasA wins more than 50% on random pairsRandomize order at display time, log the shown order
Guideline driftKappa drops after a guideline editVersion guidelines, re-run gold items, never mix versions in one snapshot unmarked
Leaked eval promptsEval scores jump after a roundHash and dedupe prompts against every eval set before annotation
PII in source logsAnnotators see customer dataScrub before sampling, restrict workspaces, log access
Silent export bugsChosen and rejected swapped for one classUnit-test the exporter on hand-built records; spot-check 50 rows per snapshot

Trade-offs

The central trade-off is speed versus independence. Suggestions cut time per item but pull labels towards the model that made them; lean on them for style and format, and prefer blind first passes for factual correctness and safety. Double annotation doubles cost but is the only way to know your noise rate. Human labels cost far more than LLM-judge labels, but they are what you calibrate the judge against, so even a judge-heavy pipeline keeps a human slice. Online DPO shows how that slice is used when pairs are generated every training step.

What to do next

  1. Write down the task shapes you need (SFT edit, pairwise, rating, span) and pick the tool whose data model already has them.
  2. Run a 200-record pilot; measure seconds per accepted item with and without suggestions.
  3. Add a hidden control slice and 5% gold items before scaling past the pilot.
  4. Generate suggestions on GPU in an offline batch; log agent name and checkpoint on every suggestion.
  5. Use uncertainty plus cluster caps to pick items, and keep 10-20% uniformly random.
  6. Double-annotate at least 15% and compute kappa or alpha per question every day.
  7. Export with prompt-level splits, eval-set dedupe and a hashed, versioned snapshot.
  8. Feed both_bad records back into the next SFT correction round.
Key takeaway: An annotation tool turns GPU-drafted suggestions into human-verified training rows. Choose it by task shape and measured seconds per item, treat suggestions as anchors to be measured rather than labels, spend humans on uncertain and diverse items, measure agreement continuously, and export versioned, prompt-split snapshots so every training row can be traced back to who labelled it and under which rules.