Tool calling is the step where a language model stops talking and starts acting: it chooses a function, fills in arguments, and later reads the result. Large hosted models do this well out of the box. Small models, roughly the one-to-eight-billion-parameter range that fits on a single modest GPU or a laptop, are cheaper, faster and private, but they make more mistakes, and in tool calling a mistake is an action rather than a sentence.

Runtime guards such as constrained decoding and validation, covered in function calling with SLMs, stop malformed output. They cannot make a model choose the right tool or know when not to call one. This article is about the model side: what the task really asks of the model, how to pick a base, why the chat template matters more than people expect, how to shrink the problem, how to build training data and fine-tune with LoRA, and how to measure the result honestly. A worked example runs through all of it.

Advertisement

What tool calling asks of a model

It helps to break one tool-calling turn into separate decisions, because a small model can be good at some and poor at others, and each needs different training data.

  1. Decide. Does this request need a tool at all, or should the model answer directly, or ask a clarifying question?
  2. Select. Which tool, out of those offered, matches the intent?
  3. Fill. Extract argument values from the conversation, convert them to the right types and formats, and leave out what the user did not say rather than inventing it.
  4. Sequence. For multi-step tasks, call tools in an order where later calls use earlier results, or issue independent calls in parallel.
  5. Read. Interpret the tool result, including errors, and either answer, retry with corrected arguments or call another tool.

Small models usually handle filling well when the schema is clear, and struggle most with deciding and sequencing: they call a tool when none is needed, or stop after one step. Measuring each decision separately tells you where to spend effort.

Why size matters, and where it does not

A small model has less capacity to hold a long list of tool descriptions in mind, follows long instructions less faithfully, and generalises less well to tools it has never seen. Accuracy typically falls as the number of offered tools grows and as descriptions become similar to each other. The vLLM documentation is candid on this point for one family, noting that smaller Llama models struggle to use tools effectively.

The flip side is that tool calling in a real product is narrow. A support assistant may have forty tools, all known in advance, used in patterns that repeat thousands of times a day. That is exactly the setting where a small model fine-tuned on in-domain examples can match a much larger general model on the cases that matter, at a fraction of the latency and cost.

Where a small model sits in a tool-calling loop, and what you can change around itUser requestplus conversationTool retriever40 tools to top 5Chat templatetools rendered inSmall modelfine-tuned with LoRATool-call parserformat must matchValidatorschema + policyExecutorreal API callvalidTool resultrole tool messageappended to context, model called againOffline looplogs to training data to LoRA to eval gatesingle, parallel, multi-step, no-call and clarify casesnew adapterModel choice, template, catalog size and training data are model-side levers; parser and validator are runtime levers.
The small model is one box in a loop. Retrieval shrinks the catalog, the chat template renders tools in the format the model was trained on, the parser and validator guard the output, and logged traffic feeds an offline fine-tuning loop.
Advertisement

Choosing a base model

Start from an instruction-tuned checkpoint whose maintainers trained it for tool use and ship a chat template that renders tools and tool results. A model without native tool training can be taught, but you then have to invent a format and teach it from scratch, which costs far more data. Beyond that, compare candidates on four things.

  • Format support in your serving stack. Can your server parse its tool-call output? This is a hard requirement, covered in the next section.
  • Context length you can afford. Tool schemas, conversation history and tool results all consume context. Check the context length the model was trained on and what your memory budget allows with the KV cache at your batch size.
  • Licence and deployment target. On-device deployment narrows choices to models that quantise well, which you should test with your own evaluation set, not assume.
  • Baseline on your tasks. Run every candidate, before any fine-tuning, on two or three hundred real examples from your domain. Public leaderboards such as the Berkeley Function Calling Leaderboard are a useful first filter, but rank order on your tools can differ.

Chat templates and tool-call formats

A model learns tool calling in a specific textual format: where tool definitions appear, how a call is written, and how a result is returned. Some families wrap a JSON object in tags, some emit bare JSON, and some emit Python-like function calls. The chat template, usually a Jinja template shipped with the tokenizer, produces this format from structured messages. If you render tools differently from training, or hand-write a prompt that looks similar, accuracy drops sharply, and the drop is easy to misread as the model being weak.

The serving layer must parse the same format back into structured calls. In vLLM, automatic tool calling is enabled with --enable-auto-tool-choice and a --tool-call-parser that matches the model family. Its documentation lists, for example, the hermes parser for Qwen2.5 and Hermes models, llama3_json for Llama 3.1 and 3.2, and a pythonic parser that suits the small Llama 3.2 models, with the caveat that in that mode the model must not produce text and tool calls in the same generation. The tool_choice values required and a named function use vLLM's structured-output backend, so the output is schema-valid by construction; auto relies on the parser.

# Serve a tool-capable model (pick the parser that matches its family)
vllm serve Qwen/Qwen2.5-3B-Instruct \
  --enable-auto-tool-choice --tool-call-parser hermes

# Call it through the OpenAI-compatible API
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")

tools = [{
  "type": "function",
  "function": {
    "name": "get_order_status",
    "description": "Look up the shipping status of one order by its order id.",
    "parameters": {
      "type": "object",
      "properties": {"order_id": {"type": "string", "pattern": "^ORD-[0-9]{6}$"}},
      "required": ["order_id"],
    },
  },
}]

resp = client.chat.completions.create(
    model="Qwen/Qwen2.5-3B-Instruct",
    messages=[{"role": "user", "content": "Where is ORD-482913?"}],
    tools=tools, tool_choice="auto")
print(resp.choices[0].message.tool_calls)

Shrink the problem before training

The cheapest accuracy gain comes from offering fewer, clearer tools. Put a retriever in front of the model: embed each tool's name and description, embed the user request, and pass only the top few candidates, plus any tools the conversation is already using. Going from forty tools to five shortens the prompt, lowers latency and removes most confusable neighbours from view. Measure retrieval recall separately; if the right tool is not in the top five, no model can call it.

Then design schemas for a small reader. Use verb-noun names that differ in their first word, one-sentence descriptions that say when to use the tool and when not to, few required arguments, enums instead of free strings where values are known, and formats such as ISO dates stated in the description. Split a tool that takes a mode flag into two tools. The same principles appear in structured output with SLMs, which covers schema design for extraction.

Building training data

Fine-tuning data for tool calling is a set of complete conversations: system prompt, available tools, user turns, assistant tool calls, tool results and final answers, rendered through the model's own chat template. Good sources are logged production traffic with corrected labels, traces from a larger model run on your tools and then checked, and synthetic requests generated from templates. Whatever the source, cover the whole decision space, not only the happy path.

CaseExampleWhy it matters
Single callWhere is ORD-482913?The bulk of traffic; teaches selection and filling
Parallel callsStatus of these two ordersSmall models often serialise or drop the second
Multi-stepRefund my last orderLook up last order, then refund it with its id
No callWhat are your opening hours?Teaches abstention when the answer is in the prompt
ClarifyCancel my orderAsk which order instead of guessing an id
Irrelevant toolsRequest with only unrelated tools offeredModel must not force a near-miss tool
Error recoveryTool returns not foundRetry with corrected input or explain, never loop

Verify every example mechanically before training: parse the calls, validate arguments against the schema, and where possible execute them against a sandbox to confirm the final answer matches the tool result. Remove near-duplicates and hold out a test set by tool and by phrasing, so evaluation measures generalisation rather than memorisation. A few thousand verified, varied conversations often move a small model more than tens of thousands of noisy ones.

Fine-tuning with LoRA and loss masking

Low-rank adaptation trains small adapter matrices while the base weights stay frozen, so a tool-calling adapter can be trained on one GPU and swapped per product; the LoRA article covers rank and target choices. Two details matter specifically for tool calling. First, render data with the tokenizer's template, passing the tool definitions so they appear exactly as at inference. Second, compute loss only on assistant tokens, the calls and answers, not on the system prompt, the tool definitions or the tool results. Training on tool results teaches the model to hallucinate them.

from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import LoraConfig, get_peft_model

tok = AutoTokenizer.from_pretrained(BASE)
model = AutoModelForCausalLM.from_pretrained(BASE, torch_dtype="bfloat16")
model = get_peft_model(model, LoraConfig(
    r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"]))

def encode(example):
    ids, labels = [], []
    # Render the conversation prefix by prefix so we know which tokens are assistant output.
    for i, msg in enumerate(example["messages"]):
        prefix = tok.apply_chat_template(example["messages"][:i + 1],
                                         tools=example["tools"], tokenize=False)
        new = tok(prefix, add_special_tokens=False)["input_ids"][len(ids):]
        ids += new
        labels += new if msg["role"] == "assistant" else [-100] * len(new)
    return {"input_ids": ids, "labels": labels}

The prefix-by-prefix rendering is simple but assumes the template is append-only, so check it for your model by comparing against a single full render. Keep learning rates modest and train for few epochs; overfitting shows up first as the model calling tools for requests that need none, which is why the no-call cases in the evaluation set are the canary.

Evaluating honestly

Score each decision separately on the held-out set. For calls, compare the predicted call to the reference structurally rather than as strings: same function name, same set of arguments, values equal after normalising types and order. Report executable accuracy as well, whether running the call yields the reference result, because two different argument values can be equally valid. Track abstention as precision and recall on the no-call and clarify cases, and track multi-step task success end to end, since per-call accuracy can look fine while chains fail. The general pitfalls in evaluating SLMs apply: contamination between training and test, and test sets that are easier than traffic.

Gate releases on these numbers. A new adapter ships only if it improves the target metric without regressing abstention or any tool's accuracy beyond a tolerance you set in advance.

Worked example: a forty-tool support assistant

Consider an e-commerce support assistant with forty tools and a 3B-parameter instruction-tuned model. The figures below are illustrative of the process, not measurements to quote. Baseline with all forty tools in the prompt: the model picks the right tool on about two thirds of single-call requests and calls a tool on a large share of requests that need none. Adding a retriever that passes the top five tools, with retrieval recall checked at over 98 percent on the test set, shortens prompts by thousands of tokens and fixes many selection errors. Rewriting schemas to remove a mode flag and two near-duplicate tools fixes more.

Then a LoRA adapter trained on four thousand verified conversations covering all seven cases from the table lifts selection and filling further and, critically, makes abstention reliable. Served by vLLM with the matching parser and a validator, the remaining errors cluster in multi-step refunds, which a confidence rule routes to a larger model that sees only this hard tail.

Failure modes

  • Template mismatch. Hand-built prompts or the wrong parser make a capable model look broken. Always render with the shipped template and confirm the parser by round-tripping known calls.
  • Invented arguments. The model fills a required field the user never gave, such as an order id. Train clarify cases and make policy-sensitive fields come from session state, not the model.
  • Over-calling. Fine-tuning on call-heavy data teaches the model to always call. Balance the data and watch abstention.
  • Loops. After a tool error the model repeats the same call. Cap tool turns per request and include error-recovery examples.
  • Silent drift. Tools change but the adapter does not. Version adapters with the tool catalog and re-run evaluation on every schema change.

Trade-offs

Fine-tuning gives the biggest gains on a fixed catalog but ties the adapter to it; a catalog that changes weekly favours better retrieval and schemas over retraining. Constrained decoding guarantees valid structure but can hide a model's uncertainty, so pair it with abstention metrics. A routing design, small model first and large model for low-confidence or multi-step cases, usually captures most of the savings while keeping quality, at the price of maintaining two paths and a router.

What to do next

  1. Collect two or three hundred real requests with reference tool calls, including no-call and clarify cases, as your evaluation set.
  2. Baseline two or three tool-trained small models on it using their shipped chat templates and the matching server parser.
  3. Add tool retrieval, measure its recall, and rewrite schemas for clarity before any training.
  4. Build a verified training set covering single, parallel, multi-step, no-call, clarify, irrelevant and error cases.
  5. Train a LoRA adapter with loss on assistant tokens only and compare it on every metric, not only call accuracy.
  6. Ship behind a validator, cap tool turns, and route low-confidence or multi-step requests to a larger model.
  7. Re-run the evaluation whenever a tool schema changes, and retrain from logged corrections on a schedule.
Key takeaway: A small model can be a reliable tool caller when the problem is made narrow and the model is trained on that narrow problem. Pick a base trained for tool use, render tools only through its shipped chat template and serve it with the matching parser. Cut the catalog with retrieval and clear schemas, fine-tune a LoRA adapter on verified conversations that include abstention, clarification and error recovery, mask the loss to assistant tokens, and judge every release on per-decision metrics before routing the hard tail to a larger model.