Most prompt engineering is manual search. You change a sentence, run a few examples, squint at the output and change another sentence. It works for a demo, but it does not scale: every model upgrade invalidates the tuning, nobody can say why a phrase is there, and a prompt that improved five examples may have broken fifty others. DSPy, from the Stanford NLP group, turns that search into a program you can rerun. You declare what each language-model step takes and returns, write a metric that says what a good answer is, and let an optimizer pick the instructions and examples that score best on your data.

This article explains DSPy from first principles: the pieces, what the optimizers actually do, a worked example you can adapt, how much an optimization run costs, and the ways it goes wrong. API details were checked against the DSPy 3.4.0 source; DSPy moves quickly, so confirm signatures against the version you install. For background on what demonstrations do to a model, see Few-Shot Prompting, in depth.

Advertisement

The idea: prompts are parameters, the metric is the spec

A DSPy program looks like ordinary Python. It calls one or more language-model steps and combines their results with normal control flow. What differs is that you never write the prompt text for a step. You write a signature, which names the inputs and outputs and gives a short description, and DSPy turns it into a prompt at call time through an adapter that also parses the response back into typed fields.

Because the prompt is generated, it can be changed without touching your code. An optimizer treats each step's instruction and its list of demonstrations as parameters, tries candidate values, runs the program on your examples, and keeps the values your metric scores highest. This is the same shape as training a model, except the parameters are text and the gradient is replaced by search. The consequence matters more than the mechanism: the metric becomes your specification. If it rewards the wrong thing, the optimizer will find prompts that deliver the wrong thing efficiently.

DSPy compile loop: the program is fixed, the prompts are the parametersProgram (dspy.Module)signatures + control flowtrainset / valsetdspy.Example recordsmetric(example, pred, trace)bool or floatOptimizer.compile()bootstrap demos, propose instructions, searchTask LMruns candidate programsPrompt / reflection LMwrites candidate instructionsScores per candidatemetric on valset batchestrialsfeedbackCompiled programinstructions + demos per predictorprogram.save('triage.json')state only, load into same classthe metric is the spec:a weak metric compiles a weak prompt
The compile loop. Your program, data and metric go in; the optimizer runs candidate prompts on the task model, scores them, and returns the same program with better instructions and demonstrations baked into each predictor.

The building blocks

There are five pieces to learn. A dspy.LM wraps a model endpoint; dspy.configure(lm=...) makes it the default. A Signature declares fields. A module such as dspy.Predict or dspy.ChainOfThought executes a signature; ChainOfThought adds a reasoning field before the outputs. A dspy.Module subclass composes modules in forward. A dspy.Example is one data record, and with_inputs marks which fields are inputs so the rest count as labels.

import dspy
from typing import Literal

lm = dspy.LM("provider/model-name", temperature=0.0)   # placeholder model string
dspy.configure(lm=lm)

class TriageTicket(dspy.Signature):
    """Route a customer support ticket to the team that should handle it."""
    ticket: str = dspy.InputField(desc="raw ticket text from the customer")
    team: Literal["billing", "auth", "outage", "feature_request", "other"] = dspy.OutputField()
    urgent: bool = dspy.OutputField(desc="true only if the customer is blocked right now")

class Triage(dspy.Module):
    def __init__(self):
        super().__init__()
        self.classify = dspy.ChainOfThought(TriageTicket)

    def forward(self, ticket):
        return self.classify(ticket=ticket)

program = Triage()
pred = program(ticket="I was charged twice for October and cannot download my invoice")
print(pred.team, pred.urgent, pred.reasoning)

The docstring and field descriptions are the seed instruction. The Literal type constrains the output, and the adapter rejects a response it cannot parse into those types, which is one of the main reasons DSPy programs are more robust than string-templated prompts.

Advertisement

Metrics: the part you cannot delegate

A metric is a function metric(example, pred, trace=None) that returns a number or a bool. It is called in two situations. During evaluation, trace is None and you return a score. During bootstrapping, DSPy passes the execution trace and uses your return value to decide whether a run is good enough to become a demonstration, so the convention is to return a strict bool when trace is not None.

def triage_metric(example, pred, trace=None):
    team_ok = pred.team == example.team
    urgent_ok = pred.urgent == example.urgent
    score = 0.7 * team_ok + 0.3 * urgent_ok
    if trace is not None:          # bootstrapping: only perfect runs become demos
        return team_ok and urgent_ok
    return score

Three rules keep a metric honest. Score what the downstream consumer needs, not what is easy to check. Make partial credit explicit so the optimizer can tell a near miss from a disaster. And when you use a language model as a judge inside the metric, treat that judge as a component with its own error rate: measure its agreement with human labels before trusting the optimizer's numbers. The evaluation design behind this is covered in Prompt Evaluation Architecture in Depth.

GEPA uses a richer metric. Its metric must accept five arguments, (gold, pred, trace, pred_name, pred_trace), and may return dspy.Prediction(score=..., feedback=...). The feedback string is text the reflection model reads, so 'predicted billing, gold was auth; the ticket mentions a password reset' teaches far more than a bare 0.

What the optimizers actually do

DSPy calls optimizers teleprompters in its source, and they differ in what they search and how much they cost. The table lists the ones most teams reach for, in rough order of cost.

OptimizerSearchesNeedsUse when
LabeledFewShotPicks k labeled examples as demosLabeled trainsetBaseline; almost free
BootstrapFewShotRuns a teacher on trainset, keeps traces that pass the metric as demosMetric, roughly 10 to 50 examplesFirst real optimization; demos include reasoning
BootstrapFewShotWithRandomSearchSeveral bootstrapped demo sets, keeps the best on validationMore examples and callsBootstrap helps but varies run to run
MIPROv2Instructions and demos jointly; proposes candidates, then Bayesian search over themMetric, trainset, ideally a valset; auto budgetInstruction wording matters, not just examples
GEPAInstructions, by reflecting on failed traces and textual feedback; keeps a Pareto setFive-argument metric, a reflection_lm, a budgetSmall data, rich failure feedback, multi-step programs

BootstrapFewShot is the one to understand first. It runs the program, often with a stronger teacher, on training examples; each run that passes the metric yields a full input-to-output trace for every predictor, including intermediate reasoning, and those traces become demonstrations. Its defaults are up to four bootstrapped and sixteen labeled demos per predictor and one round. For each round it copies the model at temperature 1.0 with a new rollout id, so it bypasses the cache and gets varied traces.

MIPROv2 adds instruction search. It bootstraps candidate demo sets, asks a prompt model to propose instructions grounded in your data and program, then runs a Bayesian search over combinations, scoring them on minibatches of the validation set. Its auto setting of light, medium or heavy sets the budget; light is the default. GEPA works differently: it reads the traces of failing examples together with your feedback, has a reflection model rewrite the instruction, and keeps a Pareto frontier of candidates that each win on some examples instead of a single best. GEPA asserts that exactly one of auto, max_full_evals or max_metric_calls is set, and that a reflection_lm or custom proposer is supplied.

Worked example: compiling the triage program

Assume 300 labeled historical tickets. Split them once, by time if you can, so the validation and test sets look like future traffic: 60 for training, 120 for validation, 120 held back for a final test that no optimizer ever sees. Then measure, optimize in cheap steps, and measure again.

import json
records = [json.loads(line) for line in open("tickets.jsonl")]
data = [dspy.Example(ticket=r["text"], team=r["team"], urgent=r["urgent"]).with_inputs("ticket")
        for r in records]
train, val, test = data[:60], data[60:180], data[180:]

evaluate = dspy.Evaluate(devset=val, metric=triage_metric, num_threads=8, display_progress=True)
baseline = evaluate(Triage())
print("baseline", baseline.score)            # a percentage, e.g. 71.5, not 0.715

boot = dspy.BootstrapFewShot(metric=triage_metric, max_bootstrapped_demos=4, max_labeled_demos=4)
compiled_boot = boot.compile(Triage(), trainset=train)
print("bootstrap", evaluate(compiled_boot).score)

mipro = dspy.MIPROv2(metric=triage_metric, auto="light", num_threads=8)
compiled_mipro = mipro.compile(Triage(), trainset=train, valset=val)
print("mipro", evaluate(compiled_mipro).score)

final = dspy.Evaluate(devset=test, metric=triage_metric, num_threads=8)
print("test", final(compiled_mipro).score)  # report this number, once
compiled_mipro.save("triage_v3.json")

Read the results in a fixed order. If bootstrap barely moves the baseline, inspect the failing validation examples before spending on MIPROv2: often the labels disagree with each other, or the signature is missing an input the model needs, such as the customer's plan tier. If MIPROv2 improves validation but the test score drops back, the search fitted the validation set; use more validation data or a lighter budget. Then look at what was learned with dspy.inspect_history(n=1) and read the final prompt. A compiled prompt you cannot explain is a prompt you cannot debug.

Notice what you did not do: you never edited a prompt string. When the team adds a new routing category next quarter, you add it to the Literal, relabel, and recompile. When a cheaper model arrives, you change one line and recompile against the same metric, which turns a model migration from a week of prompt archaeology into an afternoon of measurement.

Budgets, caching and what a run costs

An optimization run is many language-model calls, so estimate before you start. Evaluation costs roughly the number of examples times the calls per program run. Bootstrapping costs about one program run per training example per round, more when the teacher is a larger model. MIPROv2 adds instruction proposals from the prompt model plus one minibatch evaluation per trial, with periodic full evaluations. GEPA's budget is explicit in metric calls, which makes it the easiest to cap.

DSPy caches model responses by default, which makes reruns of identical calls nearly free and makes repeated evaluations reproducible. That cache is also a trap: if you change the provider's model behind the same name, cached answers hide the change. Bootstrapping deliberately bypasses it with fresh rollout ids. Set num_threads to what your rate limit allows, not to what your laptop allows; throttled calls that fail are scored as failures, and max_errors decides when a run aborts instead of quietly scoring errors as zero. The per-request arithmetic in Cost Optimization, in depth also applies here: demonstrations add input tokens to every production call, so a compiled program with sixteen demos can cost several times more per request than the baseline.

Saving, versioning and deploying a compiled program

program.save('triage_v3.json') writes the learned state: each predictor's instruction, demonstrations and signature. To load it, construct the same class and call load. Saving with save_program=True pickles the whole program to a directory instead; DSPy requires an explicit opt-in to load pickles because they can run arbitrary code, so prefer the JSON state in production.

program = Triage()
program.load("triage_v3.json")      # same class, same signature, same field names

Treat the JSON file as a build artifact. Store it with the DSPy version, the model string, the metric's version, the data snapshot and the validation and test scores. A compiled state is tied to the model it was optimized for; loading it against a different model usually works syntactically and is unmeasured semantically. Rerun the evaluation in CI whenever any of those inputs change, and recompile rather than hand-editing the saved instructions, because the next compile will overwrite the edit.

Failure modes

SymptomCauseFix
Validation up, test flat or downSearch overfit a small valsetLarger valset, lighter auto budget, report test once
Scores look great, outputs look wrongMetric rewards a proxy, or a judge metric is lenientAudit 30 outputs by hand; calibrate the judge
Bootstrap produces zero demosMetric too strict, or the teacher fails every exampleRelax the trace-time metric, use a stronger teacher
Many parse failuresOutput types the model cannot produce reliablySimplify fields; check with inspect_history
Run aborts or scores collapse midwayRate limits counted as failuresLower num_threads, set max_errors deliberately
Production cost jumps after compileMany long demos per predictorCap max demos; compare cost per correct answer
Quality drops after a provider updateCompiled state is model-specificRe-evaluate on model change; recompile
Leaked test examples as demosOverlapping splits or duplicate ticketsDeduplicate before splitting; split by time

Trade-offs: when DSPy is the wrong tool

DSPy pays off when you have a repeatable task, at least a few dozen labeled examples and a metric you believe. Without those, it optimizes noise. It is overkill for a one-off prompt you will run ten times, and it gives little when the bottleneck is missing context rather than phrasing; no instruction search will make a model know your refund policy if retrieval never supplies it.

It also moves effort rather than removing it. Hand-tuning time becomes data-labeling and metric-design time, plus optimization spend. Teams that already maintain an evaluation set get the most out of it, because the hard part is done. Teams that adopt DSPy to avoid building evaluations discover that the evaluation is the product. If your main lever is choosing examples per request rather than a fixed set, compare with Dynamic Few-Shot, in depth; and if you want to know whether the reasoning field earns its tokens, see Chain-of-Thought Prompting, in depth.

What to do next

  1. Pick one repeatable task and collect 150 to 300 labeled examples; split them once, by time, into train, validation and a test set you touch only at the end.
  2. Write the signature with typed outputs and a one-line docstring, wrap it in a Module, and run dspy.Evaluate for a baseline.
  3. Write the metric with partial credit and a strict bool when trace is set; hand-check 30 scored outputs against it.
  4. Run LabeledFewShot and BootstrapFewShot first; read the failures before paying for MIPROv2 or GEPA.
  5. Run MIPROv2 with auto set to light, or GEPA with an explicit max_metric_calls and a five-argument feedback metric.
  6. Compare cost per correct answer, not only accuracy, then save the JSON state with the model, DSPy version and scores beside it.
  7. Add the evaluation to CI and recompile, never hand-edit, when the model, data or metric changes.
Key takeaway: DSPy replaces hand-tuned prompt strings with signatures, a metric and an optimizer that searches over instructions and demonstrations. The metric is the real specification, so invest there first: give partial credit, return a strict bool while bootstrapping, and calibrate any judge. Start with BootstrapFewShot, move to MIPROv2 or GEPA when instruction wording matters, keep a test set no optimizer sees, and ship the saved JSON state as a versioned artifact tied to its model.