Fine-tuning a Claude model means handing a training service a file of example conversations and asking it to make the model behave more like those examples. The file format looks trivial: one JSON object per line, a system prompt and some turns. Almost every failed job and every disappointing result still traces back to the file: a record in the wrong shape, a conversation that ends on the wrong speaker, a label that teaches the wrong habit.

This page explains the Claude training record from first principles, as Amazon Bedrock documents it. Bedrock is where Claude fine-tuning is documented, and its data-preparation guide lists Claude 3 Haiku as the Claude model that supports it. Two dated facts frame everything below. Anthropic retired claude-3-haiku-20240307 on its own Claude API on 20 April 2026, and its deprecations page says partner platforms such as Bedrock set their own schedules. So before you plan a job, check what your account can fine-tune today; the code later on shows how. The format itself is worth learning either way, because a clean set of Claude-shaped conversations is also your evaluation set and your prompt library.

For general line framing and streaming writers see JSONL for fine-tuning; for the OpenAI equivalent see the OpenAI fine-tuning format.

Advertisement

Where the format sits in the pipeline

The dataset passes through four hands. You curate it, a local validator checks it, the Bedrock customization job reads it from S3, and the same rows later score the resulting model. Only the validator runs on your machine, so it is the only place where a mistake costs seconds rather than a failed job and a bill.

Source datatickets, transcriptsCuratededupe, label, redactClaude JSONLsystem + messagesLocal validatorrules + token budgetpassS3 buckettrain.jsonl, val.jsonlCustomization jobFINE_TUNINGmetrics to S3Custom modelprivate copyProvisionedthroughputHeld-out evalsame JSONL rowsfailures become new training rowsError reviewlabel, fix, re-addOne dataset feeds the validator, the job, and the evaluation set; only the validator runs on your machine.
The Claude fine-tuning pipeline on Bedrock. The JSONL file is the contract between every stage.

The record, field by field

Each line of the file is one complete training example. The Claude-native shape has two top-level keys:

{"system": "You triage insurance claims. Reply with JSON only.", "messages": [{"role": "user", "content": "Claim 88213: rear-ended at a light, bumper damage, no injuries."}, {"role": "assistant", "content": "{\"category\": \"auto_collision\", \"injury\": false, \"priority\": \"standard\"}"}]}
  • system is optional and is a plain string. It plays the role of the system prompt you will send at inference time, so write the one you intend to use in production, not a different one per row.
  • messages is a list of turns. Each turn has exactly role (user or assistant) and content as a string.

The rules that Anthropic's Bedrock fine-tuning guide states, and that the validator below enforces: at least two messages; the first from the user; the last from the assistant; roles alternate strictly; and no keys other than these. The guide also warns against the reserved markers Human: and Assistant: on a new line inside prompts, a holdover from Claude's older text-completion format.

Multi-turn records are the same shape with more turns. They teach the model how to behave later in a conversation: asking a clarifying question, then answering once the user replies. Bedrock does not document how the loss is distributed across turns, so assume every assistant turn in a record is something you are teaching. If an early assistant turn is mediocre, it is training data too; fix it or cut the record down to the turns you would sign off on.

Advertisement

Two shapes that look alike: Claude-native and Converse

Bedrock has a second conversational format, built on its Converse API, used for several other models it fine-tunes. It differs in ways a hand-written converter easily gets half right, which is the most common reason a file looks valid and is still rejected.

AspectClaude-native recordConverse record
Version markernone"schemaVersion": "bedrock-conversation-2024"
System promptone string: "system": "..."a list of blocks: "system": [{"text": "..."}]
Turn contenta stringa list of blocks: [{"text": "..."}] (plus image blocks for vision models)
Documented forClaude 3 HaikuLlama 3.2 and 3.3 models, among others

Pick the shape from the documentation for the model you are actually tuning, generate every row through one function, and never mix the two in one file.

Quotas and a worked size estimate

The Bedrock guide lists these limits for Claude 3 Haiku fine-tuning. The total record count is adjustable through Service Quotas; the others are fixed.

LimitValue
Minimum records32
Maximum training records10,000
Maximum validation records1,000
Maximum total records10,000 (adjustable)
Maximum tokens per record32,000
Training file size10 GB
Validation file size1 GB

Work an example. A claims team has 6,000 labelled claims. Each record has a 120-word system prompt, a claim description averaging 250 words and a short JSON answer. At roughly 1.3 tokens per English word that is about 500 tokens a record, far from the 32,000 ceiling, so per-record length is not the constraint. The record budget is: holding out 600 rows for validation leaves 5,400 for training, 6,000 in total, inside the default 10,000. The file is about 15 MB. The binding limits for this team are label quality and record count, not bytes.

Long-document tasks are different. A record that pastes a 40-page contract can pass 32,000 tokens. Truncating it silently teaches the model to answer from text it never saw, so split the task (extract per section, then summarize) or drop the record.

Validate locally before anything is uploaded

The validator below enforces every rule above, reports line numbers, and exits non-zero so it can gate a CI job.

import json, sys

ALLOWED_TOP = {"system", "messages"}
ALLOWED_MSG = {"role", "content"}
RESERVED = ("\nHuman:", "\nAssistant:")

def check_record(line_no, rec, max_tokens=32_000, chars_per_token=3.5):
    errs = []
    extra = set(rec) - ALLOWED_TOP
    if extra:
        errs.append(f"unexpected top-level keys {sorted(extra)}")
    system = rec.get("system", "")
    if not isinstance(system, str):
        errs.append("system must be a string in the Claude-native shape")
    msgs = rec.get("messages")
    if not isinstance(msgs, list) or len(msgs) < 2:
        return errs + ["messages must be a list of at least two turns"]
    for i, m in enumerate(msgs):
        if set(m) - ALLOWED_MSG:
            errs.append(f"turn {i}: unexpected keys {sorted(set(m) - ALLOWED_MSG)}")
        want = "user" if i % 2 == 0 else "assistant"
        if m.get("role") != want:
            errs.append(f"turn {i}: role {m.get('role')!r}, expected {want!r}")
        content = m.get("content")
        if not isinstance(content, str) or not content.strip():
            errs.append(f"turn {i}: content must be a non-empty string")
        elif any(r in content for r in RESERVED):
            errs.append(f"turn {i}: contains a reserved Human:/Assistant: marker")
    if msgs[-1].get("role") != "assistant":
        errs.append("last turn must be the assistant")
    chars = len(system) + sum(len(m.get("content") or "") for m in msgs)
    if chars / chars_per_token > max_tokens * 0.9:     # rough, keep 10% headroom
        errs.append(f"about {int(chars / chars_per_token)} tokens; too close to {max_tokens}")
    return [f"line {line_no}: {e}" for e in errs]

def validate(path, min_records=32, max_records=10_000):
    errors, count = [], 0
    with open(path, encoding="utf-8") as f:
        for n, raw in enumerate(f, 1):
            if not raw.strip():
                errors.append(f"line {n}: blank line"); continue
            try:
                rec = json.loads(raw)
            except json.JSONDecodeError as e:
                errors.append(f"line {n}: not JSON ({e.msg})"); continue
            count += 1
            errors += check_record(n, rec)
    if not min_records <= count <= max_records:
        errors.append(f"{count} records; expected {min_records}-{max_records}")
    return errors

if __name__ == "__main__":
    problems = validate(sys.argv[1])
    print("\n".join(problems) or "ok")
    sys.exit(1 if problems else 0)

Run it on both files, fix every reported line, and add one more check of your own: that no validation record also appears in training.

Converting data you already have

Most teams arrive with data in the OpenAI messages shape, or as chat logs with tool calls. The mapping is mechanical except for three decisions: system and developer messages move to the top-level system string; consecutive turns from the same role are merged, because the Claude record requires strict alternation; and anything that is not plain text, such as tool calls, tool results or images, has no slot in this format and must be rewritten as text or dropped.

def openai_to_claude(rec):
    """Map an OpenAI-style {"messages": [...]} record to the Claude-native shape."""
    system_parts, turns = [], []
    for m in rec["messages"]:
        role, content = m["role"], m.get("content")
        if role in ("system", "developer"):
            system_parts.append(content)
            continue
        if role not in ("user", "assistant") or not isinstance(content, str):
            raise ValueError(f"cannot map role={role!r} or non-text content; drop or rewrite this record")
        if turns and turns[-1]["role"] == role:          # merge consecutive same-role turns
            turns[-1]["content"] += "\n\n" + content
        else:
            turns.append({"role": role, "content": content})
    while turns and turns[0]["role"] != "user":          # must open with the user
        turns.pop(0)
    while turns and turns[-1]["role"] != "assistant":    # must close with the target
        turns.pop()
    out = {"messages": turns}
    if system_parts:
        out["system"] = "\n\n".join(system_parts)
    return out

Log every record the converter rejects and count them by reason. If a large share of your data is tool use, this format is the wrong vehicle for that behaviour; teach it with prompts and tool definitions instead.

Starting the job

The job is one call to the Bedrock control-plane client. Check the base model first: list the foundation models your account can fine-tune in the Region, then pass that exact identifier. Hyperparameter values are passed as strings.

import boto3

bedrock = boto3.client("bedrock", region_name="us-west-2")

# 1. Confirm which base models this account and Region can fine-tune today.
for m in bedrock.list_foundation_models(byCustomizationType="FINE_TUNING")["modelSummaries"]:
    if m["providerName"] == "Anthropic":
        print(m["modelId"], m.get("modelLifecycle", {}).get("status"))

# 2. Start the job with the identifier printed above.
bedrock.create_model_customization_job(
    customizationType="FINE_TUNING",
    jobName="claims-triage-ft-2026-10-01",
    customModelName="claims-triage-v1",
    roleArn="arn:aws:iam::123456789012:role/BedrockCustomizationRole",
    baseModelIdentifier=BASE_MODEL_ID,
    hyperParameters={                       # values are strings
        "epochCount": "2",
        "batchSize": "32",
        "learningRateMultiplier": "1.0",
    },
    trainingDataConfig={"s3Uri": "s3://ml-data/claims/train.jsonl"},
    validationDataConfig={"validators": [{"s3Uri": "s3://ml-data/claims/val.jsonl"}]},
    outputDataConfig={"s3Uri": "s3://ml-data/claims/output/"},
)

Three hyperparameters are exposed: epochCount, batchSize and learningRateMultiplier. AWS material has described their ranges as 1 to 10 epochs (default 2), batch size 4 to 256 (default 32) and a multiplier from 0.1 to 2 (default 1); confirm the current values in the console before relying on them. Start from the defaults. Raise epochs only when the validation loss is still falling at the end of the run; lower the multiplier when the validation loss rises while the training loss falls.

When the job finishes, Bedrock writes training and validation metrics to the output S3 prefix. A fine-tuned Claude 3 Haiku is served through Provisioned Throughput rather than on-demand, so include that capacity in the cost estimate before you start, not after the model looks good.

Worked example: a claims triage model

The claims team wants every claim mapped to one of 14 categories, an injury flag and a priority, returned as strict JSON. A prompted base model gets the categories right 88 percent of the time on their held-out set and sometimes wraps the JSON in prose. Their dataset decisions:

  1. One fixed system prompt for every record, identical to the production prompt.
  2. Assistant turns are minified JSON with keys in a fixed order, so the model learns one exact output shape.
  3. Rare categories are up-sampled to at least 150 examples each; the four commonest are down-sampled so they do not swamp the rest.
  4. Two hundred multi-turn records where the claim is ambiguous: the assistant asks one specific question, the user answers, and the assistant replies with JSON.
  5. Personal data in claim text is replaced with typed placeholders before export, because everything in the file is copied into the training job.

They score both models with the same 600 held-out records, parsing each reply as JSON and comparing fields.

Failure modes

  • Shape mix-up. Converse-style content arrays inside a Claude-native file, or the reverse. Rejected at validation, or worse, accepted and learned as literal text.
  • Ending on the user. Logs exported mid-conversation end with a user turn. The record has no target, so it is rejected; trim it to the last assistant turn.
  • System-prompt drift. Training with one system prompt and serving with another weakens the tuning. Treat the prompt as part of the model version.
  • Teaching the bad habits in your logs. Real transcripts contain apologies, hedges and wrong answers. Fine-tuning reproduces whatever the assistant turns contain.
  • Leakage. Near-duplicate records split across training and validation inflate the validation score and hide overfitting.
  • Base-model lifecycle. A tuned copy is tied to its base. Bedrock sets its own retirement dates, so check its lifecycle terms for custom models built on a retiring base. Keep the dataset, not the model, as the durable asset.

Trade-offs: when to fine-tune at all

Fine-tuning buys consistency of format and tone, and lets a smaller, cheaper model do a narrow job that would otherwise need a larger one. It costs a training job, dedicated serving capacity, and a frozen model that does not improve when the base family does. The alternative, a current model with a carefully written system prompt and a few examples (with prompt caching keeping the repeated prefix cheap), needs no training and moves to every new model for free.

Use the dataset to decide. Run the held-out records against a well-prompted current model first. If it meets the target, stop. If it does not, the same records become training data. Because the Claude-native record is close to a Messages API request (a system string plus alternating turns), the same file serves as few-shot examples, a regression suite and training data, and it can be mapped to open-weight models; see LoRA for small models for that route.

What to do next

  1. Run list_foundation_models(byCustomizationType="FINE_TUNING") in your Region and record which Claude base models, if any, your account can tune today and their lifecycle status.
  2. Export 50 real conversations, convert them with the function above, and run the validator; fix the converter until the error count is zero.
  3. Write the production system prompt first and use it unchanged in every record.
  4. Hold out at least 10 percent as validation, check for near-duplicates across the split, and score a prompted current model on it before training anything.
  5. If you train, start from default hyperparameters, read the validation metrics in the output prefix, and compare per category, not just overall.
  6. Store the dataset, the converter and the validator in version control next to the system prompt, so the work survives the base model's retirement.
Key takeaway: The Claude fine-tuning record is a JSONL line with an optional system string and a strictly alternating list of user and assistant turns that ends on the assistant. On Bedrock it is distinct from the Converse shape, has firm record and token quotas, and should be validated locally before upload. Base models retire on their own schedules, so treat the curated, validated dataset as the asset: it is your evaluation set, your prompt library and your training data at once.