ShareGPT is the de facto storage format for multi-turn chat fine-tuning data. A record is a JSON object with a conversations list, and each turn is an object with a from field naming the speaker and a value field holding the text. The name comes from ShareGPT, a site where people shared their ChatGPT conversations; Vicuna was trained on conversations collected from it, and the shape stuck long after the site stopped mattering. Today most open fine-tuning tools can read it, and many public chat datasets ship in it.
The format is easy to read and easy to get subtly wrong. It says nothing about how turns become tokens, which roles exist, which turns the model should learn from, or how a conversation longer than the context window is cut. Those decisions live in your pipeline, and a mistake in any of them trains the model on the wrong thing without any error message. This article walks through the record, then the pipeline that turns it into training tensors: normalisation, validation, chat-template rendering, per-turn loss masking, truncation, packing and deduplication. For single-turn instruction data read the Alpaca format; for line framing and loaders read JSONL for fine-tuning.
The record, field by field
A minimal ShareGPT record looks like this:
{"conversations": [
{"from": "system", "value": "You are a terse SQL assistant."},
{"from": "human", "value": "Count orders per day for last week."},
{"from": "gpt", "value": "SELECT order_date, COUNT(*) FROM orders WHERE ..."},
{"from": "human", "value": "Only paid ones."},
{"from": "gpt", "value": "SELECT order_date, COUNT(*) FROM orders WHERE status = 'paid' AND ..."}
]}The common role values are human, gpt and system. Tool-using datasets add more. LLaMA-Factory, for example, documents function_call for the model's tool invocation and observation for the tool's result, accepts a separate top-level system string and a tools string, and lets you rename every field through tags in dataset_info.json. Its documented rule is that human and observation turns sit in odd positions and gpt and function_call turns in even ones, which is a strict alternation between the outside world and the model.
Nothing in the format is enforced by a schema. Records in the wild use user and assistant instead of human and gpt, put the system prompt as a turn or as a top-level field, add per-turn weight flags, or carry metadata such as a source id or language. Treat ShareGPT as a family of dialects rather than a specification, and write a normaliser that accepts the dialects you actually have and rejects everything else.
The pipeline from JSON to training tensors
Each stage has one job. Normalise maps dialects to a single internal shape, usually the OpenAI-style list of role and content that Hugging Face chat templates expect. Validate enforces structural rules. Dedup and split happen per conversation, before any tokenisation, so near-identical conversations cannot straddle train and evaluation. Render uses the chat template that ships with the model you are tuning, never a hand-written one. Tokenise and mask produce input_ids and labels, with labels set to -100 everywhere the model should not learn. Truncate or pack fits the result to your sequence length.
Keeping these as separate, testable functions matters because the failure modes are silent. A broken mask still produces a falling loss curve; it just teaches the model to predict user messages too.
Normalising roles and content
The normaliser below accepts the common dialects and produces role and content messages. It merges consecutive turns from the same speaker, which some scraped datasets contain, and moves a top-level system string to the front.
ROLE_MAP = {
"human": "user", "user": "user",
"gpt": "assistant", "assistant": "assistant", "chatgpt": "assistant",
"system": "system",
"function_call": "tool_call", "observation": "tool",
}
class BadRecord(ValueError):
pass
def normalise(rec):
turns = rec.get("conversations") or rec.get("messages")
if not isinstance(turns, list) or not turns:
raise BadRecord("no conversation list")
msgs = []
if isinstance(rec.get("system"), str) and rec["system"].strip():
msgs.append({"role": "system", "content": rec["system"].strip()})
for i, t in enumerate(turns):
speaker = t.get("from", t.get("role"))
text = t.get("value", t.get("content"))
role = ROLE_MAP.get(str(speaker).lower())
if role is None:
raise BadRecord(f"turn {i}: unknown role {speaker!r}")
if not isinstance(text, str) or not text.strip():
raise BadRecord(f"turn {i}: empty value")
if msgs and msgs[-1]["role"] == role and role in ("user", "assistant"):
msgs[-1]["content"] += "\n\n" + text.strip() # merge runs
else:
msgs.append({"role": role, "content": text.strip()})
return msgsBe deliberate about two choices. Stripping whitespace is usually right, but not for code datasets where trailing newlines or indentation are part of the answer; strip only the ends. Merging consecutive turns is a repair, so count how often it happens per source. If a source needs it on a large fraction of records, the source is probably mis-parsed and deserves inspection rather than silent repair.
Validation rules that catch real problems
After normalisation, check structure. The rules below reject records that would render into nonsense or teach the wrong behaviour:
- At most one system message, and only first. Many chat templates raise an error or silently drop a system message in the middle.
- The first non-system turn is from the user. A conversation that opens with an assistant turn has lost its prompt, usually through a scraping bug.
- Strict alternation. After merging, user and assistant must alternate, with tool calls and tool results paired: every tool_call is followed by a tool turn, and every tool turn is followed by an assistant or another tool_call.
- The last turn is from the assistant. A trailing user turn contributes no learning signal and wastes tokens. Drop it, or reject the record if you want strict sources.
- Length bounds. Reject turns above a character ceiling, and conversations whose total rendered length exceeds several times your sequence length; they will be truncated into fragments.
- Content checks. Reject refusals and boilerplate you do not want imitated, text that mentions the source product by name if that leaks into answers, and encoding damage such as replacement characters.
Log every rejection with source file, line number and reason, and look at the counts per reason. A dataset where 30 percent of records fail alternation is telling you about its own construction.
Rendering with the chat template and masking loss per turn
The model never sees JSON. It sees the string its chat template produces, with special tokens marking role boundaries; the Llama chat template is a worked example of one. Training must use exactly the template used at inference, so render with the tokenizer's own apply_chat_template.
The standard objective for chat SFT trains only on assistant tokens. User text, system prompts and tool results get label -100, so the model learns to answer, not to imitate users. A robust way to find assistant spans for any template is incremental rendering: render the conversation up to each assistant turn with the generation prompt appended, render it again including that turn, and the token difference is the span to label.
IGNORE = -100
def tokens(tok, msgs, gen_prompt=False):
return tok.apply_chat_template(msgs, tokenize=True,
add_generation_prompt=gen_prompt)
def build_example(tok, msgs, max_len):
input_ids = tokens(tok, msgs)
labels = [IGNORE] * len(input_ids)
for i, m in enumerate(msgs):
if m["role"] != "assistant":
continue
before = tokens(tok, msgs[:i], gen_prompt=True)
after = tokens(tok, msgs[:i + 1])
if input_ids[:len(after)] != after or after[:len(before)] != before:
raise BadRecord("template is not prefix-stable; mask needs another method")
labels[len(before):len(after)] = after[len(before):]
return truncate_at_turn(input_ids, labels, max_len)The prefix check is not decoration. Some templates rewrite earlier turns depending on what follows, for example by removing reasoning text from all but the final assistant turn, and then the incremental method mislabels tokens. When the check fails, use the template's own mechanism if it has one: Hugging Face tokenizers can return an assistant-token mask when the template marks assistant spans with a generation block, so check whether yours does. Whatever you use, decode the labelled spans of a few records and read them. That five-minute check catches most masking bugs.
Decide whether the end-of-turn token belongs to the label. It should: a model that never learns to emit it will not stop generating. The incremental method above includes it automatically when the template appends it after the assistant content.
Tool calls and observations
Tool-use conversations add two roles. The model's call (function_call in LLaMA-Factory's naming) is something the model must learn to produce, so it is labelled like an assistant turn. The tool's result (observation) comes from outside and is masked like user text. Templates differ widely in how they render tools: some expect a structured tool_calls field on the assistant message, others a text convention inside the content. Convert your stored call into whatever the target template expects, and store calls as structured JSON in your canonical data so you can re-render for a different model later.
The normaliser above leaves calls under an internal tool_call role that no stock template knows. Convert them just before rendering, so the mask loop's assistant check covers them automatically. The exact tool_calls shape below is a common one; check what your template reads.
import json
def to_template_roles(msgs):
out = []
for m in msgs:
if m["role"] == "tool_call":
call = json.loads(m["content"]) # {"name": ..., "arguments": {...}}
out.append({"role": "assistant", "content": "",
"tool_calls": [{"type": "function", "function": call}]})
else:
out.append(m)
return outValidate tool turns harder than text turns. Parse every call's arguments as JSON, check the function name exists in the record's tool list, and reject observations that are suspiciously long; a pasted web page as a tool result can dominate the token budget of the whole conversation.
Worked example: sizing and cutting a real dataset
Suppose you are tuning a small model with a 4,096-token training length on 120,000 ShareGPT conversations. A token census after rendering shows a median of 900 tokens, a 95th percentile of 5,200 and a long tail to 40,000. Validation rejects 6 percent, mostly for empty turns and trailing user turns. Near-duplicate removal by hashing normalised first user turns plus first assistant turns removes another 9 percent, which is typical for scraped chat data where the same prompt was shared many times.
About 7 percent of what remains exceeds 4,096 tokens. Naively cutting at token 4,096 would leave many records ending mid-answer, teaching the model to stop abruptly. Instead, cut at turn boundaries: keep the system message, then keep whole user and assistant pairs from the start while they fit, and drop the rest. Optionally, emit the remainder as a second example that carries the system message and the last pair before the cut as context, with only the new assistant turns labelled. Short conversations are then packed several to a sequence, with attention masks or position resets so packed conversations do not attend to each other.
def truncate_at_turn(input_ids, labels, max_len):
if len(input_ids) <= max_len:
return input_ids, labels
# last position <= max_len where a labelled span ends (end of an assistant turn)
cut = 0
for j in range(1, max_len + 1):
if labels[j - 1] != IGNORE and (j == len(labels) or labels[j] == IGNORE):
cut = j
if cut == 0:
raise BadRecord("first assistant turn does not fit")
return input_ids[:cut], labels[:cut]
Failure modes
| Symptom | Likely cause | Check |
|---|---|---|
| Model continues by writing the user's next message | Labels cover user turns, or end-of-turn token unlabelled | Decode labelled spans for five records |
| Answers stop mid-sentence | Token-level truncation inside assistant turns | Count examples whose last label is not an end-of-turn token |
| Strong eval scores, weak real use | Duplicates across train and eval | Split by conversation hash after dedup |
| Model ignores the system prompt | System dropped by template or stored in an unread field | Render one record and read the string |
| Training crash on some batches | Unknown role or non-alternating turns reach the template | Validator before rendering, rejections logged |
| Model says it is ChatGPT | Source identity leaked in assistant turns | Content filter on assistant text |
Trade-offs and when to use something else
ShareGPT versus OpenAI-style messages. They carry the same information. Role and content messages are what chat templates consume directly and what the OpenAI fine-tuning format uses, so many teams store that internally and treat ShareGPT purely as an import dialect. Keeping one canonical internal shape is worth more than which one you pick.
Training on all assistant turns versus only the last. Labelling every assistant turn uses more signal per conversation. Labelling only the final turn suits datasets where early answers are low quality. Some tools support a per-turn weight for this; with your own pipeline, it is one condition in the mask loop.
Packing versus padding. Packing raises throughput substantially for short conversations but needs correct cross-example attention handling. Padding is simple and wastes compute. With adapters such as LoRA, the cheaper runs make padding more tolerable while you validate the pipeline.
What to do next
- Pick one internal message shape and write a normaliser with an explicit role map; reject unknown roles rather than guessing.
- Add the validation rules above and log rejections per source and reason.
- Deduplicate per conversation, then split train and evaluation by conversation hash.
- Render with the target model's own chat template and build labels for assistant and tool-call turns only.
- Decode the labelled spans of a few records from every source and read them before the first training run.
- Run a token census, choose a sequence length from it, and truncate at turn boundaries.