For a long time the Workbench was the first place developers met Claude: a page in Anthropic's developer Console where you wrote a system prompt, filled in a user message, pressed Run and looked at the answer. It grew into a small prompt-engineering environment, with variables, saved prompts, versions and an Evaluate tab for running a prompt against many test cases. Anthropic has since retired it. As of 1 October 2026 the Claude Help Center describes Workbench (legacy) as retired and the Console offers a playground in its place: a stateless page built on the public Messages API that keeps your draft in the browser and exports it as code.

This page covers what the Workbench did, what the playground does and does not do, how a playground draft maps onto one API request, and then the part that matters most: rebuilding the Workbench's evaluation loop (versioned prompts, test cases, grading, side-by-side comparison) as code you own. The worked example takes a support-ticket triage prompt from first draft to a measured, promotable version.

Advertisement

What the Workbench was, and what replaced it

The Workbench began as a prompt console: choose a model, write a system prompt and messages, adjust settings, run, and copy the result out as SDK code. In 2024 Anthropic added evaluation features. Its announcement described a prompt generator that drafted a prompt from a task description, prompts written with input variables in double braces such as {{ticket_text}}, test cases that Claude could generate or that you could import from a CSV, side-by-side comparison of two or more prompts, a 5-point scale for grading response quality, and the ability to create new prompt versions and re-run the test suite. Those features lived inside the Console, attached to saved prompts.

The current help article, titled How do I use the playground?, says Workbench is now the playground. The playground is stateless: your current draft stays in your browser, and you keep a copy by switching to the code view. Saved prompts, prompt versions, evals and prompt sharing are not part of it. Legacy Workbench data could be exported as JSON from the Console's privacy settings until 1 September 2026. That deadline has passed, so if your team relied on saved Workbench prompts and did not export them, they are no longer reachable from the Console.

The design choice makes sense. A prompt is production configuration, and keeping its only copy in a vendor UI, outside version control and review, was always fragile. The playground keeps the Console to fast exploration and leaves storage, versioning and evaluation to your repository.

What the playground gives you

Open the Claude Console, choose Playground in the navigation, and pick a workspace if your organisation uses workspaces; requests are billed and rate-limited against that workspace. The help article lists the capabilities:

  • Model and settings. A model selector and settings such as maximum output tokens, plus a temperature control.
  • Prompt editing. A system prompt and user messages; Run sends the request and shows the response with token counts and usage.
  • Tools and structured outputs. You can attach tool definitions to test tool use, and define a structured output shape that Claude must return.
  • Raw request and response. The full message structure, the stop reason and usage, exactly as the API sees them.
  • Code export. A code toggle turns the current request into a snippet you can paste into a project.

The raw view is the most educational feature. Every playground draft is one POST /v1/messages request: a model ID, max_tokens, an optional system string, a messages array of alternating user and assistant turns, optional tools, and an optional output_config. The response carries a content array of blocks, a stop_reason such as end_turn, max_tokens, tool_use or refusal, and a usage object with input and output token counts.

Advertisement

Sampling settings on current models

One trap catches people moving from the playground to code. Older guides, and older Workbench habits, set temperature to zero for deterministic classification. On the newest models that no longer works: Claude Opus 5.5 rejects non-default temperature and top_p values with a 400 error, and Claude Sonnet 5.5 rejects non-default sampling values as well. Claude Haiku 4.5 still accepts them. The supported control on the current Opus and Sonnet models is output_config.effort, which trades thoroughness of reasoning against token spend; on Opus 5.5 it defaults to medium, so set it explicitly. Determinism for classification now comes from constraining the output, with a JSON schema whose fields are enums, rather than from sampling settings.

Two other habits also need updating. Assistant-message prefill (starting Claude's reply with an opening brace to force JSON) returns a 400 on the current families; structured outputs replace it. And forcing a specific tool with tool_choice set to any or a named tool is rejected on Opus 5.5 and Sonnet 5.5; use auto with a clear instruction, strict: true on the tool, or structured outputs when the forced call only existed to obtain JSON. The tool use guide walks through the full tool loop.

The loop, rebuilt

From a playground draft to a repeatable prompt evaluation loopConsole playgroundstateless draft, in browserCode exportone Messages API requestPrompt file in gitprompts/triage.v3.jsoncode togglecommitTest casescases.jsonl (+ edge cases)Eval runnerrender, call, record usageMessages APImodel + output_configinputsrequestresponseGraderscode checks + LLM judgeoutputsReportv2 vs v3, per casescoresDecisionpromote or reviseWorkbench (legacy) held prompts, versions and evals inside the Console. The playground does not,so the versioned file, the case set and the report now live in your repository and CI.
The Workbench evaluation loop rebuilt outside the Console: explore in the playground, export, version the prompt in git, and run cases, graders and comparisons in code.

The playground is for exploration, the prompt file is the versioned artefact, the case set is test data, the runner turns a version plus a case into one request, graders score each result, and the report compares versions. Together they replace the Evaluate tab and add what it lacked: diffs, code review and CI.

Worked example: a versioned triage prompt

Suppose an internet provider wants Claude to classify incoming tickets into a category and priority, with a one-sentence reason a human agent can check. After a few rounds in the playground the draft works on the examples you tried. Export it and store it as a file, keeping Workbench's double-brace variable convention so the template stays readable:

{
  "id": "ticket-triage",
  "version": 3,
  "model": "claude-opus-5-5",
  "effort": "low",
  "system": "You triage customer support tickets for an internet service provider. Classify each ticket into exactly one category and one priority. Use only the ticket text; if it is ambiguous, choose the closest category and say why in the reason field.",
  "user_template": "<ticket>\n{{ticket_text}}\n</ticket>\n<customer_tier>{{tier}}</customer_tier>",
  "schema": {
    "type": "object",
    "properties": {
      "category": {"type": "string", "enum": ["billing", "outage", "hardware", "account", "other"]},
      "priority": {"type": "string", "enum": ["p1", "p2", "p3"]},
      "reason":   {"type": "string"}
    },
    "required": ["category", "priority", "reason"],
    "additionalProperties": false
  }
}

The version number is explicit, so every result traces to its prompt. The ticket sits inside XML-style tags, separating instructions from untrusted customer text. The schema uses enums, so a misspelt category is impossible; see structured output.

The case set is a JSON Lines file of objects with an id, vars and expected. Start with twenty to fifty anonymised real tickets and add edge cases: an empty ticket, another language, billing plus an outage, an abusive one, one that tries to instruct the model. Claude can propose edge cases, as the old Generate Test Case button did, but label them yourself.

The runner

import json, re, anthropic

client = anthropic.Anthropic()          # reads ANTHROPIC_API_KEY or an `ant auth login` profile

def render(template: str, variables: dict) -> str:
    """Fill {{name}} placeholders; fail loudly on a missing variable."""
    def sub(m):
        name = m.group(1)
        if name not in variables:
            raise KeyError(f"test case is missing variable {name!r}")
        return str(variables[name])
    return re.sub(r"\{\{(\w+)\}\}", sub, template)

def run_case(prompt: dict, case: dict) -> dict:
    resp = client.messages.create(
        model=prompt["model"],
        max_tokens=16000,                    # thinking tokens count toward this on Opus 5.5
        system=prompt["system"],
        messages=[{"role": "user", "content": render(prompt["user_template"], case["vars"])}],
        output_config={
            "effort": prompt["effort"],
            "format": {"type": "json_schema", "schema": prompt["schema"]},
        },
    )
    record = {
        "case_id": case["id"],
        "prompt_version": prompt["version"],
        "stop_reason": resp.stop_reason,
        "input_tokens": resp.usage.input_tokens,
        "output_tokens": resp.usage.output_tokens,
    }
    if resp.stop_reason != "end_turn":          # refusal, max_tokens, ...: never parse blindly
        record["output"] = None
        return record
    text = next(b.text for b in resp.content if b.type == "text")
    record["output"] = json.loads(text)
    return record

The runner is deliberately dull. It renders the template and refuses to send a request with a missing variable, which is the bug that silently ruins evaluations. It asks for a schema-constrained answer through output_config.format. It records the stop reason and token usage with every result, because a prompt that is two points more accurate but three times the output tokens may not be an improvement. And it never parses output when the stop reason is anything other than end_turn: a max_tokens cut-off or a refusal is a result to count, not an exception to swallow.

For large case sets, the Message Batches API runs the same requests asynchronously at half the price; results arrive in any order, so key them by custom_id. When cases share a long system prompt, prompt caching cuts input cost.

Grading: code first, a judge where you must

JUDGE_MODEL = "claude-haiku-4-5"   # a different, cheaper model grades the reasons

def grade_exact(record: dict, case: dict) -> dict:
    out = record["output"]
    if out is None:
        return {"category_ok": False, "priority_ok": False}
    return {
        "category_ok": out["category"] == case["expected"]["category"],
        "priority_ok": out["priority"] == case["expected"]["priority"],
    }

def grade_reason(record: dict, case: dict) -> int:
    """1-5 rubric score for the free-text reason, in the spirit of Workbench's 5-point grade."""
    if record["output"] is None:
        return 1
    rubric = (
        "Score the REASON from 1 to 5.\n"
        "5: cites the specific ticket facts that justify the category and priority.\n"
        "3: plausible but generic; could apply to many tickets.\n"
        "1: wrong, invented facts, or contradicts the ticket.\n"
        "Reply with the digit only."
    )
    resp = client.messages.create(
        model=JUDGE_MODEL,
        max_tokens=5,
        temperature=0,                 # accepted on Haiku 4.5; current Opus/Sonnet models reject it
        system=rubric,
        messages=[{"role": "user", "content":
            f"<ticket>{case['vars']['ticket_text']}</ticket>\n"
            f"<reason>{record['output']['reason']}</reason>"}],
    )
    digits = re.findall(r"[1-5]", resp.content[0].text)
    return int(digits[0]) if digits else 1

Grade with code wherever the answer is checkable. Category and priority are exact matches against labels, so they need no model at all. The free-text reason is not checkable that way, so a second, cheaper model scores it against a short rubric on a 1 to 5 scale, which recreates the spirit of Workbench's 5-point grading without a human clicking through every row. Anthropic's own evaluation guidance recommends grading with a different model from the one that produced the output, and favouring many automatically graded cases over a few hand-graded ones.

A judge has biases too: hand-grade twenty results, compare with the judge, and tighten the rubric where they disagree. The prompt evaluation guide covers calibration.

Side-by-side comparison and promotion

def compare(old: dict, new: dict, cases: list[dict]) -> None:
    rows = []
    for case in cases:
        a, b = run_case(old, case), run_case(new, case)
        ga, gb = grade_exact(a, case), grade_exact(b, case)
        rows.append((case["id"], ga, gb, a["output_tokens"], b["output_tokens"]))
    def acc(i, key):
        return sum(r[i][key] for r in rows) / len(rows)
    print(f"category accuracy  v{old['version']}={acc(1,'category_ok'):.2%}  v{new['version']}={acc(2,'category_ok'):.2%}")
    print(f"priority accuracy  v{old['version']}={acc(1,'priority_ok'):.2%}  v{new['version']}={acc(2,'priority_ok'):.2%}")
    for case_id, ga, gb, ta, tb in rows:
        if ga != gb:                      # side-by-side view of only the cases that changed
            print(f"  {case_id}: v{old['version']} {ga} -> v{new['version']} {gb}  tokens {ta}->{tb}")

This is the Evaluate tab's comparison view, as text. Run both versions over the same cases, report aggregate accuracy, then list only the cases whose grades changed. That last list is where the real review happens: a new version that fixes eight tickets and breaks two may still be worse if the two it broke are outages marked p3. Promote a version only when it beats the current one on the metrics you defined in advance, and record the comparison output in the pull request that changes the version number.

Wire the runner into CI so any prompt change re-runs the case set and fails the build if accuracy drops below production; eval-driven development describes the practice.

Failure modes

  • The only copy lived in the Console. Teams that kept production prompts as saved Workbench prompts and missed the export deadline now have to reconstruct them from deployed code. Treat the repository as the source of truth from day one.
  • Playground success, production failure. A draft tried on five friendly examples meets real tickets with pasted logs. Only edge-case coverage catches this.
  • Copied settings that the model rejects. Code that sets temperature=0 or prefills JSON fails with a 400 on the newest models. Run exported snippets against the exact model you will deploy.
  • Unparsed stop reasons. Treating a max_tokens cut-off as valid output corrupts accuracy numbers. Record and count every stop reason.

Trade-offs

DecisionOption AOption B
Where prompts liveConsole UI: quick for one person, invisible to review and CIFiles in git: reviewable, testable, needs a little tooling
GradingHuman review of every output: high signal, slow and rarely repeatedCode checks plus a calibrated judge: repeatable on every change, needs calibration
Running casesSynchronous calls: results in seconds, full priceMessage Batches: half price, results arrive later and out of order
Output controlFormat instructions in the prompt: flexible, occasionally malformedJSON schema through output_config: always valid, less room for free-form answers

What to do next

  1. Find every prompt your team still keeps only in a UI or a document, and move it into a versioned file in the repository that deploys it.
  2. Use the playground to explore a new prompt, then export the request with the code toggle and check it runs unchanged against the model you will deploy.
  3. Write a case set of at least twenty real, anonymised inputs plus deliberate edge cases, each with an expected label.
  4. Build the runner and graders above: exact-match checks for anything checkable, a calibrated judge on a different model for the rest, and stop reasons recorded for every case.
  5. Add a CI job that re-runs the case set on any prompt change and blocks the merge if accuracy falls below the production version.
  6. Replace temperature and prefill habits with effort settings and JSON schemas, and review the comparison report in every pull request that bumps a prompt version.
Key takeaway: The Console Workbench is retired, and its replacement playground is a stateless window onto a single Messages API request: excellent for exploration, with no storage, versioning or evaluation. Rebuild what Workbench gave you in your own repository: a versioned prompt file with explicit variables, a labelled case set with edge cases, a runner that records stop reasons and usage, code graders backed by a calibrated judge on a different model, and a side-by-side report that gates every prompt change in CI. Use effort settings and schemas, not temperature or prefill, on current models.