Public SWE-bench leaderboards answer one question: how well does a given model plus a given agent scaffold fix real GitHub issues? Teams that run their own inference ask something different. They have already picked a few candidate models. Now they need to know which serving configuration to put into production: which quantization, which context cap, how many GPUs, whether to turn on speculative decoding, what sampling settings. They need to know how much each choice costs per issue actually fixed.
This page shows how to use SWE-bench as a serving benchmark rather than a model leaderboard. It covers what the benchmark measures, the serving knobs that move its score, a reproducible architecture, the harness commands, the statistics you need before trusting a difference, cost per resolved issue, a worked comparison and the failure modes that make results lie. The numbers in the worked example are illustrative, not measurements of any particular model or GPU.
What SWE-bench actually measures
SWE-bench was introduced by Jimenez and colleagues at Princeton (ICLR 2024). Each task instance is a real pull request from a popular Python repository. The instance has the issue text, the repository at the commit before the fix, a hidden gold patch, and two sets of tests. FAIL_TO_PASS tests fail before the fix and must pass after it. PASS_TO_PASS tests passed already and must keep passing. An instance counts as resolved only if the model's patch applies cleanly and both sets pass inside the instance's container.
The original set has 2,294 instances drawn from 12 repositories. SWE-bench Lite is a 300-instance subset of more self-contained issues. SWE-bench Verified is a 500-instance subset that OpenAI had human annotators screen in 2024 for clear problem statements and fair tests. In February 2026 OpenAI said it had stopped reporting SWE-bench Verified, citing flawed tests and contamination, and pointed to Scale AI's SWE-bench Pro instead. Take two lessons from that. First, an absolute score on an old public split says little about capability, because models may have seen the fixes in training. Second, contamination largely cancels when you compare configs of the same weights, so a relative comparison between serving setups is still informative. That relative use is what this page is about.
For background on code benchmarks in general (HumanEval, pass@k, sandboxing), see coding capability evals.
Serving knobs that move the score
An agentic SWE-bench run makes dozens of model calls per instance: read files, search, edit, run tests, retry. Small serving differences compound over that loop. These are the knobs that matter most, roughly in order of how often they surprise people:
| Knob | How it changes the score | What to log |
|---|---|---|
| Context cap (max model length) | Long trajectories get truncated or the scaffold drops history; the agent forgets what it already tried | Prompt tokens per call, truncation events |
| Quantization (FP8, INT4 weight-only) | Usually small accuracy loss, but errors accumulate over 50+ calls and show up in exact edits | Format errors, failed patch applies |
| Tool-call parser | A parser mismatch turns valid tool calls into plain text; the agent stalls | Calls with no parsed tool call |
| max_tokens per response | A cut-off edit produces a malformed diff | finish_reason == length |
| Temperature and seed | Changes the trajectory; adds run-to-run variance | Sampling params per request |
| Speculative decoding | Lossless for greedy decoding in exact arithmetic; affects speed, not answers, if implemented correctly | Acceptance rate, tokens/s |
| Prefix caching | No effect on outputs; large effect on cost because agent prompts share long prefixes | Cache hit rate |
Two of these are pure speed knobs (speculative decoding, prefix caching). If turning one of them on changes the resolve rate beyond noise, treat that as a bug signal, not a trade-off. For how drafting works, see speculative decoding.
Architecture of a serving benchmark rig
The rig has three parts. The scaffold is the agent loop that turns an issue into a patch. Use one open scaffold and pin it to a commit, with the same system prompt, tools and step limit for every run. If you change the scaffold and the endpoint together, you cannot tell which one moved the score. The endpoint is your production serving stack, for example vLLM with an OpenAI-compatible API, started with exactly the flags you would ship. The harness builds a Docker image per instance, applies each patch and runs the tests. It never sees your endpoint, so cost and latency have to come from your own request logs.
Running the endpoint, scaffold and harness
Start the endpoint with production flags and record them alongside the run. A vLLM example:
# config B: FP8 weights on two GPUs, 64k context, tool calling on
vllm serve ./checkpoints/coder-32b \
--tensor-parallel-size 2 \
--quantization fp8 \
--max-model-len 65536 \
--enable-prefix-caching \
--enable-auto-tool-choice --tool-call-parser hermes \
--port 8000Pick the tool-call parser that matches your model's chat template; the wrong one is the most common cause of a mysteriously low score. Then run the scaffold over a fixed instance list. Its core loop looks like this, whatever framework you use:
def solve(instance, client, cfg, log):
repo = checkout(instance["repo"], instance["base_commit"])
messages = [system_prompt(cfg), user_prompt(instance["problem_statement"])]
for step in range(cfg.max_steps): # e.g. 50
t0 = time.time()
resp = client.chat.completions.create(
model=cfg.model, messages=messages, tools=TOOLS,
temperature=cfg.temperature, max_tokens=cfg.max_tokens, seed=cfg.seed)
log.write(instance["instance_id"], step, resp.usage, time.time() - t0,
resp.choices[0].finish_reason)
msg = resp.choices[0].message
messages.append(msg)
if not msg.tool_calls:
break # agent says it is done
for call in msg.tool_calls:
messages.append(run_tool(repo, call)) # bash, view, edit, ...
return {"instance_id": instance["instance_id"],
"model_name_or_path": cfg.run_name,
"model_patch": git_diff(repo)}Write one JSON object per line to predictions.jsonl with the three keys shown, then run the harness. The long-standing module form is:
python -m swebench.harness.run_evaluation \
--dataset_name princeton-nlp/SWE-bench_Verified \
--predictions_path predictions.jsonl \
--max_workers 8 \
--run_id fp8_tp2_seed0Recent releases add a shorter swebench eval front end; check the README of the version you install. The project recommends an x86_64 machine with plenty of free disk (on the order of 120 GB for images), and ARM support is experimental. Results are cached by run ID, so use a new run ID for every configuration.
Statistics: when is a difference real?
Resolve rate is a proportion over a few hundred instances, so it is noisy. With n = 500 and a true rate near 0.5, the standard error is sqrt(0.25 / 500), about 2.2 percentage points. A 95% interval is roughly plus or minus 4.4 points. Two configs that differ by 3 points on separate samples are not distinguishable.
You can do much better because both configs run on the same instances. Use a paired test. Most instances are solved by both or by neither, and they carry no information about the difference. Only the discordant instances matter: b solved by A only, c solved by B only. McNemar's exact test asks whether b and c could plausibly come from a fair coin:
from math import comb
def mcnemar_exact(a_solved: set, b_solved: set, all_ids: set) -> float:
b = len(a_solved - b_solved) # A only
c = len(b_solved - a_solved) # B only
n, k = b + c, min(b, c)
if n == 0:
return 1.0
tail = sum(comb(n, i) for i in range(k + 1)) / 2 ** n
return min(1.0, 2 * tail) # two-sided p-valueRun-to-run variance matters too. Even at temperature 0, batching and kernel choice can make outputs nondeterministic on GPUs, and agent trajectories amplify tiny differences. Run each config at least twice. If two runs of the same config disagree on 10% of instances, a 5-point gap between configs is within the noise floor, whatever McNemar says about a single pair.
Cost and latency per resolved issue
The metric that decides a serving choice is cost per resolved issue:
cost_per_resolved = (gpu_count * wall_clock_hours * price_per_gpu_hour) / resolved_count
# token view, useful when you share GPUs with other traffic
cost_per_resolved = sum(prompt_tokens * p_in + completion_tokens * p_out) / resolved_countMeasure wall-clock with a fixed concurrency that matches production, for example 16 agents in parallel. Agent workloads re-send long, mostly identical prompts on every step, so prefix caching and prompt-token throughput dominate. A config that looks slower on single-request decode speed can win here. Also report latency per resolved issue at p50 and p90. A developer waiting on an agent cares about the tail, and a config with a lower median but long stalls on large repositories can feel worse. For general serving trade-offs see vLLM on GPUs.
Worked example: FP8 on two GPUs vs BF16 on four
A team compares two configs of the same 32B coder model on SWE-bench Lite (300 instances), with the scaffold pinned, temperature 0 and a 50-step limit. Config A is BF16 on 4 GPUs. Config B is FP8 on 2 GPUs. All figures below are illustrative.
| Config A (BF16, TP=4) | Config B (FP8, TP=2) | |
|---|---|---|
| Resolved | 141 / 300 (47.0%) | 133 / 300 (44.3%) |
| Solved by both | 120 | 120 |
| Solved only by this config | 21 | 13 |
| Wall clock at 16 agents | 3.0 h | 3.4 h |
| GPU-hours | 12.0 | 6.8 |
| Cost at $2.50 per GPU-hour | $30.00 | $17.00 |
| Cost per resolved issue | $0.213 | $0.128 |
McNemar on b = 21, c = 13 gives a two-sided p of about 0.23, so the 2.7-point gap is not significant on this sample. A repeat run of config A disagreed with the first on 24 instances. The team reads the logs before deciding. Config B had 9 responses end with finish_reason == length in the middle of an edit, against 2 for A. Raising max_tokens from 2,048 to 4,096 removes that effect. On a re-run B resolves 139, the discordant counts become 17 and 15, and B costs about 40% less per resolved issue. B ships. The lesson is that the first gap was a serving bug, not a quantization loss, and only the per-request logs showed it.
Failure modes
- Scaffold drift. Updating the agent framework between runs changes prompts and tool schemas. Pin it by commit hash and record the hash in the run ID.
- Silent tool-call failures. The model emits a tool call the server cannot parse. The scaffold sees plain text, decides the agent is done, and submits an empty patch. Count empty patches per config; a spike means a parser or template problem.
- Context truncation. Servers that truncate prompts silently, or scaffolds that drop the oldest messages, hide the real cause. Log prompt tokens against the context cap.
- Harness infrastructure errors. Image build failures, timeouts or disk exhaustion count as unresolved. Separate harness errors from model failures before comparing.
- Contamination. Public splits may be memorized. Use them for relative comparisons only, and add a private set of your own recent issues if you need absolute numbers.
- Benchmark skew. SWE-bench is Python, issue-to-patch, from a handful of repositories. If your users write TypeScript or do refactors, add a custom suite; see custom benchmarks.
Trade-offs
Running full Verified on every candidate config is expensive: hundreds of agent trajectories plus a heavy harness. Lite is cheaper and noisier. A practical split is Lite, or a fixed random 150, for screening, and the full set for the final two candidates. More runs per config buy you variance estimates. Larger samples buy you power for small gaps. If the decision only matters for gaps of 5 points or more, two runs on 300 instances is usually enough.
Greedy decoding makes runs more repeatable but can loop. A small temperature with several seeds gives a more honest estimate of the distribution users will see. Whichever you choose, match production. For how SWE-bench fits with other eval families, see benchmark evals.
What to do next
- Pick one open scaffold, pin it by commit, and freeze its prompts, tools and step limit.
- Write down your production serving flags and start the endpoint with exactly those.
- Fix an instance list and seed; store it in the repository with the run scripts.
- Run every config twice and log usage, latency and finish_reason per request.
- Separate harness errors and empty patches from genuine failures.
- Compare configs with McNemar on discordant instances, against the same-config noise floor.
- Report cost and p90 latency per resolved issue, not just resolve rate.
- Add a small private suite of your own recent issues before trusting absolute numbers.