The Alpaca format is the simplest widely used shape for supervised fine-tuning data: every example is a JSON object with three string fields, instruction, input and output. It is so simple that teams often treat it as solved and then lose days to bugs that live around it rather than in it: a prompt template that differs between training and serving, loss computed over the prompt as well as the answer, an end-of-sequence token that never appears, or placeholder inputs that teach the model nonsense.
This article separates the format into its two real parts, the record schema and the prompt template that turns a record into text, and then walks through what a training pipeline does with them: rendering, tokenizing, masking, truncating and validating. It closes with when to keep Alpaca and when to convert your data to chat messages instead. Code is Python with a Hugging Face style tokenizer interface; the ideas carry over to any trainer.
Where the format comes from
The name comes from Stanford's Alpaca project in March 2023, which fine-tuned Meta's LLaMA 7B on about 52,000 instruction-following examples. The examples were produced with the self-instruct method: a small set of human-written seed tasks was expanded by prompting OpenAI's text-davinci-003, which also wrote the outputs. The data was released as alpaca_data.json, a single JSON array (not JSON Lines), under a CC BY-NC 4.0 licence that allows only non-commercial use. According to the project README, around 40% of the examples have a non-empty input.
Two consequences follow for anyone using the original data today. First, the licence and the provenance of the outputs matter: check them before the data goes anywhere near a commercial model. Second, the outputs reflect one 2023 model's style and mistakes, including factual errors and short, generic answers, so treat the dataset as a format reference and a baseline, not as ground truth. The format itself is unencumbered and is now used for many unrelated datasets.
Two parts: the record and the template
The record says what the example is. instruction states the task, input optionally carries the material the task operates on, and output is the desired answer. The split between instruction and input is useful: the same instruction can be paired with many inputs, and you can filter or balance by instruction type.
The template says how the record becomes the token sequence the model actually sees. Stanford's training script defines two templates, chosen by whether the input is empty. These strings are reproduced from the project's train.py:
PROMPT_INPUT = (
"Below is an instruction that describes a task, paired with an input that provides "
"further context. Write a response that appropriately completes the request.\n\n"
"### Instruction:\n{instruction}\n\n### Input:\n{input}\n\n### Response:"
)
PROMPT_NO_INPUT = (
"Below is an instruction that describes a task. "
"Write a response that appropriately completes the request.\n\n"
"### Instruction:\n{instruction}\n\n### Response:"
)
def render_prompt(rec: dict) -> str:
template = PROMPT_INPUT if rec.get("input", "").strip() else PROMPT_NO_INPUT
return template.format(instruction=rec["instruction"], input=rec.get("input", ""))Notice that the prompt ends at ### Response: with no trailing newline, and the target is the output followed by the tokenizer's end-of-sequence token. The model learns that the text after the response marker is the answer and that the answer ends with EOS. Everything about inference depends on reproducing these exact strings, which is why the template deserves to be versioned like code.
The data flow through a trainer
A trainer performs five steps for each record. It picks the template variant, renders the prompt, appends EOS to the output, tokenizes, and builds a label vector the same length as the input ids. Steps four and five are where most silent errors occur, because a wrong label vector still trains without error; it just trains the wrong thing.
Loss masking from first principles
A causal language model is trained to predict token t+1 from tokens 1..t. If you compute the loss over the whole sequence, the model spends capacity learning to reproduce your boilerplate preamble and the instructions themselves. That is wasted at best; at worst, with short outputs, the prompt dominates the gradient and the model learns to continue prompts rather than answer them. The standard fix is to set the label of every prompt position to a value the loss ignores. PyTorch's CrossEntropyLoss ignores -100 by default, and Stanford's script uses IGNORE_INDEX = -100 for exactly this.
Stanford's script tokenizes the prompt alone to learn its length, tokenizes prompt plus response, and sets label[:source_len] = IGNORE_INDEX. That works when the tokenizer splits the joined string at the same place, but tokenizers that merge characters across the boundary can shift it by one token, so a token of the answer is masked or a token of the prompt is trained. The version below avoids the question by tokenizing the halves separately and concatenating ids:
IGNORE_INDEX = -100 # the default ignore_index of torch.nn.CrossEntropyLoss
def build_example(rec, tokenizer, max_len=512):
prompt = render_prompt(rec)
response = rec["output"] + tokenizer.eos_token
# Tokenize the two halves separately and concatenate the ids, so the
# boundary between prompt and response is exact, not inferred.
p_ids = tokenizer(prompt, add_special_tokens=True)["input_ids"]
r_ids = tokenizer(response, add_special_tokens=False)["input_ids"]
if len(p_ids) >= max_len:
return None # the prompt alone does not fit: drop, do not train on nothing
r_ids = r_ids[: max_len - len(p_ids)]
input_ids = p_ids + r_ids
labels = [IGNORE_INDEX] * len(p_ids) + r_ids
return {"input_ids": input_ids, "labels": labels,
"truncated": len(r_ids) < len(tokenizer(response, add_special_tokens=False)["input_ids"])}Two details matter. Special tokens are added once, to the prompt, so the response does not start with a second beginning-of-sequence token. And truncation removes the end of the response, never the prompt: a truncated prompt produces an example whose answer refers to text the model never saw. Track how many responses were truncated; if it is more than a small fraction, raise the maximum length or drop those examples, because a response cut off before EOS teaches the model not to stop.
A worked example
Take one record from a sentiment dataset and follow it through the pipeline:
{"instruction": "Classify the sentiment of the review as positive, negative or mixed.",
"input": "Battery lasts two days, but the screen scratches far too easily.",
"output": "mixed"}
Rendered prompt (with-input variant), then the response:
Below is an instruction that describes a task, paired with an input that provides further
context. Write a response that appropriately completes the request.
### Instruction:
Classify the sentiment of the review as positive, negative or mixed.
### Input:
Battery lasts two days, but the screen scratches far too easily.
### Response:mixed</s>The record has an input, so the with-input template is used. The prompt is about sixty tokens with a typical subword tokenizer, and the response is the word mixed plus EOS, two or three tokens. The label vector is therefore roughly sixty -100 values followed by two or three real ids. The loss for this example is the average negative log-probability of those few tokens, which is exactly the behaviour you want to teach: given this prompt, say mixed and stop.
Now consider the same record trained without masking. Roughly 95% of the loss for this example would come from predicting the preamble and the review text. Across a dataset of short classification answers, the gradient is mostly about copying prompts. This is the cleanest illustration of why masking is not an optional refinement for short-answer data.
Validating and cleaning a dataset
Alpaca-style datasets, including the original, contain predictable defects: empty outputs, placeholder inputs such as <noinput> written as literal text, inputs that repeat the instruction, refusals and AI-assistant boilerplate, duplicated prompts with conflicting answers, and examples too long for the context. Community-cleaned derivatives of the original data exist precisely because of these issues. Run a validator before every training run and keep its report next to the model:
import json, hashlib, collections
BOILERPLATE = ("as an ai language model", "i cannot", "i'm sorry, but")
PLACEHOLDER_INPUTS = {"<noinput>", "n/a", "none", "no input", "noinput"}
def validate(path, tokenizer, max_len=512):
data = json.load(open(path, encoding="utf-8")) # alpaca_data.json is a JSON array, not JSONL
problems, seen, kept = collections.Counter(), set(), []
for i, rec in enumerate(data):
if not isinstance(rec, dict) or set(rec) - {"instruction", "input", "output"}:
problems["bad_keys"] += 1; continue
ins, inp, out = (str(rec.get(k, "")).strip() for k in ("instruction", "input", "output"))
if not ins or not out:
problems["empty_field"] += 1; continue
if inp.lower() in PLACEHOLDER_INPUTS:
inp = ""; problems["placeholder_input_fixed"] += 1
if inp and inp == ins:
problems["input_repeats_instruction"] += 1; continue
if any(b in out.lower()[:80] for b in BOILERPLATE):
problems["boilerplate_output"] += 1; continue
key = hashlib.sha1((ins + "\x00" + inp).lower().encode()).hexdigest()
if key in seen:
problems["duplicate_prompt"] += 1; continue
seen.add(key)
ex = build_example({"instruction": ins, "input": inp, "output": out}, tokenizer, max_len)
if ex is None:
problems["prompt_too_long"] += 1; continue
if ex["truncated"]:
problems["response_truncated"] += 1
kept.append({"instruction": ins, "input": inp, "output": out})
return kept, problemsTreat the counters as a data quality dashboard: a jump in duplicates usually means a generation job looped, and many truncated responses mean your maximum length no longer fits your data. If the data is synthetic, the generation recipe matters more than the format, and the same validation applies; see SLM distillation data recipes for building such data deliberately.
Training and serving must render the same string
A model fine-tuned on Alpaca text has learned a dialect. At inference you must render the same template with the same whitespace and stop at EOS. The most common production bug is serving an Alpaca-tuned model through a chat endpoint that applies the base model's chat template: the model sees markers it was never tuned on, answers in a degraded way, and nobody notices because the output is still fluent. The second most common bug is forgetting EOS during training, which produces a model that answers and then keeps going, inventing a new ### Instruction: section.
Make the template an artifact: store the exact template string, its version and the tokenizer name in the model's metadata, have the serving layer load them from there, and add a test that renders a fixed record in both the training and serving code paths and compares the token ids. That one test catches the majority of format regressions.
Alpaca versus chat formats
| Question | Alpaca record + template | Chat messages + model chat template |
|---|---|---|
| Turns | Single turn only | Any number of turns, plus a system message |
| Who defines the string | You, in your training code | The model's tokenizer chat template |
| Best fit | Base models, single-shot tasks, classification and extraction | Instruct models, assistants, multi-turn and tool use |
| Serving | Raw completion endpoint with your template | Any chat endpoint that applies the same template |
| Main risk | Template drift between train and serve | Template differences between model families |
If you are fine-tuning an already instruction-tuned model, convert to messages and let the model's own template render them, because the model already knows that template and your fine-tune will compose with its prior training. If you are tuning a base model for one narrow task, Alpaca's explicit template is fine and arguably clearer. Conversion is mechanical:
def alpaca_to_messages(rec, system=None):
user = rec["instruction"] if not rec.get("input") else rec["instruction"] + "\n\n" + rec["input"]
msgs = [{"role": "system", "content": system}] if system else []
return msgs + [{"role": "user", "content": user},
{"role": "assistant", "content": rec["output"]}]
# Then let the model's own chat template produce the string, so training and
# serving agree: tokenizer.apply_chat_template(msgs, tokenize=False)Join instruction and input with a blank line in the user turn, and keep the input at the end so long inputs do not separate the task from the answer. For the messages format and its validation rules see the OpenAI fine-tuning data format; for a concrete chat template see the Llama chat template.
Failure modes
- Template mismatch at serving. Fluent but worse answers. Compare token ids from training and serving renders in a test.
- No EOS in targets. The model never stops and starts writing new instructions. Append EOS and confirm it survives tokenization.
- Unmasked prompts. The model learns to copy prompts; short-answer tasks suffer most. Inspect one label vector by hand.
- Boundary drift. Masking by prompt length after joint tokenization is off by one token. Tokenize halves separately.
- Truncated prompts. Answers refer to text the model never saw. Truncate responses or drop the example, never the prompt.
- Placeholder inputs. Literal strings such as
<noinput>select the wrong template. Normalise them to empty. - Licence surprise. The original Stanford data is non-commercial. Check provenance before shipping.
Operational guidance and trade-offs
Store data as JSON Lines even if you start from the Alpaca JSON array; it streams, diffs and appends cleanly, and the record fields are unchanged. Keep a held-out set drawn per instruction type, not at random, so a large category cannot hide regressions in a small one, and evaluate generated answers with the serving template rather than with log-likelihood on the training template. For how much data and which method to use, see LoRA, in depth; for evaluation pitfalls see SLM evaluation.
The trade-off is simplicity against expressiveness: no turns, system prompt or tool calls without inventing a private dialect.
What to do next
- Write your template strings into a versioned module and record the version and tokenizer name with every trained model.
- Build examples by tokenizing prompt and response separately, masking the prompt with -100 and appending EOS; inspect one label vector by eye.
- Run the validator on every dataset build and keep its counters with the training run.
- Add a test that renders a fixed record in both training and serving code and asserts identical token ids.
- If your base is an instruct model, convert records to messages and use its chat template instead.
- Check the licence and provenance of any Alpaca-derived data before it reaches a commercial model.