What makes a dataset golden

Three properties separate a golden dataset from a pile of examples. The labels are adjudicated: each expected outcome was agreed under a written guideline, usually by more than one person, and disagreements were resolved rather than averaged. The dataset is versioned and immutable: a result is always reported against a specific version, and changing a label creates a new version instead of editing history. And it has an owner and a documented scope, so everyone knows which decisions it is fit to support and which it is not.

It is also distinct from training data. Golden cases must never be used for fine-tuning, few-shot examples or prompt tuning, because the moment they influence the system they stop measuring it. Keep them in a separate store with access controls and log who reads them.

Advertisement

Sourcing cases

The best source is real traffic, sampled deliberately. A uniform random sample reproduces the head of the distribution and misses the rare intents, languages and edge conditions where failures concentrate. Stratified sampling fixes that: group logs by intent, language, customer tier, input length or any other dimension that affects behavior, then sample each stratum with a minimum count, and record the sampling weights so that overall scores can be re-weighted to production proportions.

import random
from collections import defaultdict

def stratified_sample(records, key, per_stratum_min=30, total=1500, seed=7):
    rng = random.Random(seed)
    strata = defaultdict(list)
    for r in records:
        strata[key(r)].append(r)
    n = len(records)
    picked = []
    for name, rows in strata.items():
        share = round(total * len(rows) / n)
        take = min(len(rows), max(per_stratum_min, share))
        for r in rng.sample(rows, take):
            # weight restores production proportions when aggregating
            picked.append({**r, "stratum": name, "weight": len(rows) / take})
    return picked

Incidents are the second source, and the most valuable per case. Every escalation, user complaint and postmortem should produce at least one case that reproduces the failure, labeled with the behavior that should have happened. Expert-authored cases fill gaps logs cannot: policy boundaries, adversarial inputs, rare but high-stakes requests and questions whose correct answer is a refusal. Synthetic generation with a language model is useful for coverage, such as paraphrases, other languages or different numeric values, but every synthetic case needs human review, and synthetic cases should be tagged so they can be reported separately; models find model-written text easier than real user text.

Advertisement

The case record

A case needs more than an input and an expected output. It needs the full input context, including conversation history, retrieved documents or tool states if those are fixed for the case, and the expected behavior expressed in the form the scorer needs. That might be an exact value for extraction, a set of acceptable answers, a list of required and forbidden facts, a rubric for a judge, or an expected tool-call sequence for an agent. Provenance, labeling history and data-handling flags belong in the same record so that no one has to reconstruct them later.

{
  "case_id": "refund-eu-0311",
  "dataset_version": "support-golden v7",
  "input": {"messages": [{"role": "user", "content": "..."}], "locale": "de-DE"},
  "expected": {
    "type": "rubric+facts",
    "required_facts": ["14-day withdrawal period applies", "refund to original payment method"],
    "forbidden": ["promise of refund for used digital goods"],
    "rubric": "Answers in German; cites the withdrawal policy; no legal advice beyond policy."
  },
  "slices": ["intent:refund", "region:EU", "lang:de", "source:production"],
  "provenance": {"source": "log-sample", "sampled_at": "2026-08-14", "stratum": "refund/de"},
  "labels": [
    {"annotator": "a17", "verdict_template": "v3", "at": "2026-08-20"},
    {"annotator": "a04", "verdict_template": "v3", "at": "2026-08-21"},
    {"adjudicator": "lead-2", "resolution": "merged facts; added forbidden clause"}
  ],
  "pii": "scrubbed",
  "license": "internal-eval-only"
}

Labeling protocol and agreement

Labels are only as good as the guideline behind them. Write the guideline before labeling starts, with definitions, worked examples of each label and explicit rules for ambiguous cases, then pilot it: have two or three annotators label the same fifty cases, measure agreement, discuss disagreements, and revise the guideline until agreement is acceptable. Only then label at scale, with a fraction of cases, often ten to twenty percent, double-labeled throughout so agreement can be monitored as annotators drift.

Measure agreement with a chance-corrected statistic. Cohen's kappa works for two annotators and categorical labels; Krippendorff's alpha handles more annotators, missing labels and ordinal scales. Raw percentage agreement is misleading when one label dominates: two annotators who both mark 95 percent of cases as correct agree about 90 percent of the time by chance alone.

from collections import Counter

def cohens_kappa(a, b):
    assert len(a) == len(b) and a
    n = len(a)
    observed = sum(x == y for x, y in zip(a, b)) / n
    ca, cb = Counter(a), Counter(b)
    expected = sum(ca[k] * cb[k] for k in set(ca) | set(cb)) / (n * n)
    return (observed - expected) / (1 - expected) if expected < 1 else 1.0

# 92% raw agreement with both annotators at ~90% "pass": expected chance
# agreement is 0.9*0.9 + 0.1*0.1 = 0.82, so kappa = (0.92 - 0.82) / 0.18 = 0.56

Disagreements should go to an adjudicator who records the decision and, where the disagreement exposed a gap, updates the guideline. Persistent low agreement on one slice is a finding in its own right: the task is underspecified there, and product owners, not annotators, need to decide the intended behavior.

How many cases

Dataset size should follow from the smallest difference the dataset must detect. For a pass rate near p measured on n independent cases, the 95 percent confidence interval half-width is roughly 1.96 times the square root of p(1-p)/n. At p = 0.9, 400 cases give about plus or minus 2.9 points and 100 cases about plus or minus 5.9. If the release gate must catch a 3-point drop on a slice, a 100-case slice cannot do it on its own, whatever the headline size of the dataset.

Paired comparisons help considerably. When baseline and candidate run on the same cases, most cases pass or fail on both, and the variance of the difference depends only on the cases that flip. That is why regression gates compare paired results rather than two independent pass rates. Even so, size critical slices first: a few hundred cases per high-stakes slice is a more useful target than a large total dominated by easy, common intents.

import math

def half_width(p, n, z=1.96):
    return z * math.sqrt(p * (1 - p) / n)

def cases_needed(p, target_half_width, z=1.96):
    return math.ceil((z / target_half_width) ** 2 * p * (1 - p))

print(round(half_width(0.9, 400), 3))   # 0.029
print(cases_needed(0.9, 0.02))           # 865 cases for +/- 2 points at 90%

Privacy, licensing and contamination

Cases sampled from production contain personal and sometimes confidential data. Scrub or replace identifiers at intake with consistent placeholders, so that a case referring to the same customer twice still makes sense, and record how each case was scrubbed. Check that the data-processing basis for the original interaction permits use for evaluation, apply the same retention limits, and prefer realistic synthetic replacements for categories of data that should never be copied into an evaluation store at all.

Contamination runs in two directions. Golden cases can leak into the system under test through prompts, few-shot examples, fine-tuning sets or retrieval corpora; scan those artifacts for exact and near-duplicate matches against the dataset before every run, using normalized hashing for exact matches and shingle-based similarity such as MinHash for near-duplicates. Public benchmark items can also leak into third-party models through their pretraining data, which is one reason internal golden sets should stay private and should not be published in documentation or tickets.

Releases and maintenance

Treat the dataset like a software artifact. Each release gets an immutable version tag, a changelog listing added, removed and relabeled cases, and a dataset card describing scope, sources, sampling weights, known gaps, labeling guideline version and measured annotator agreement. Scores are always reported with the dataset version, and a gate result from v6 is never compared directly with one from v7 without re-running the baseline on v7.

Maintenance is continuous. Review cases that every configuration fails; many turn out to be label errors, and fixing them is often the single most valuable data change available. Retire or demote saturated cases that every configuration passes, keeping them in a smoke suite. Watch production for drift, such as new intents, a new market or a changed policy, and add strata when the dataset stops resembling the traffic it is meant to represent. When policy changes, relabel affected cases in a new version rather than letting old expected behavior quietly fail the correct new behavior.

Production logsstratified sampleIncidentsevery escalationExpert-authorededge + policy casesSyntheticreviewed variationsIntakePII scrub, dedup, contamination checkLabelingtwo annotators + adjudicationReleaseimmutable v-tag + cardMaintenance: label fixes, retire saturated, refresh from driftnext versionConsumers: regression gate, model selection, judge calibration, incident verification
Golden dataset lifecycle: cases from four sources pass through privacy scrubbing, deduplication and contamination checks, are double-labeled and adjudicated, and ship as immutable versions with a dataset card; maintenance feeds the next version.

Failure modes

  • Convenience sampling. Cases collected from whatever engineers happened to test are biased toward easy, well-formed inputs and overstate quality.
  • Silent relabeling. Editing labels in place makes historical scores incomparable and can hide a regression behind a label fix.
  • Single-annotator labels. Without agreement measurement, one person's interpretation becomes ground truth and nobody knows its error rate.
  • Synthetic dominance. A dataset mostly written by models rewards systems that resemble the generator, not systems that serve users.
  • Orphaned datasets. Without an owner, cases rot as products change, and the gate slowly stops measuring anything real.

Ownership

Assign each dataset an owner accountable for its scope, labeling quality and release cadence, and each critical slice a domain owner who approves its expected behaviors. Budget annotation as ongoing operating cost, not a one-time project. The return is that every model, prompt and retrieval change can be judged against the same, trusted definition of correct behavior, which is what makes fast iteration safe.

Make the dataset easy to use correctly. Provide a loader that takes a version tag and returns cases with their slices and weights, a standard report that shows per-slice results with intervals, and a request process for proposing new cases or disputing a label. When teams can add a case from an incident in minutes, the dataset keeps pace with the product; when it takes a ticket and a week, people start keeping private test sets, and the shared definition of correct behavior fragments.