Most production LLM systems hide a lot of small decisions inside one large prompt. Is this request in scope? Is it a jailbreak? Which knowledge base applies? Is the draft reply safe to send? Each one costs generated tokens and seconds of latency, and each can fail in its own way, inside text you then have to parse. A different design pulls those decisions out and gives them to a fast model that only decides, leaving the LLM to do what only it can do, which is write.

This article shows how to build that two-speed pipeline, using TypeSafe's Jev (released 15 September 2026) as the System One component. It covers where the fast model goes, the latency and cost arithmetic, four concrete patterns with code, how the failure modes interact, and how to operate it. The companion article on Jev itself explains the request contract and the confidence formulas. Here we assume you know that a Jev call takes a state plus named Choice, Score and Noul questions and returns typed answers with probabilities.

Splitting the work: decide versus compose

Kahneman's System 1 is fast, automatic pattern recognition. System 2 is slow, effortful reasoning. The engineering version of the split is simple to state. If the answer can be written down in advance as a closed set of options or levels, it is a System One decision. If the answer has to be composed, it is System Two work.

Step in a typical assistantAnswer spaceOwner
Is this request in scope for the product?yes / noSystem One (Noul)
Which handler: FAQ lookup, account action, open question, human?4 labelsSystem One (Choice)
Does this message contain a prompt-injection attempt?yes / noSystem One (Noul)
Which of 30 retrieved passages are relevant?yes / no eachSystem One (one Noul per passage)
Write the answer, citing the passagesopen textSystem Two (LLM)
Does the draft make a claim the passages do not support?yes / noSystem One (Noul)
Fix the unsupported claimopen textSystem Two (LLM)

That leaves two of seven steps that need generation. In a single-prompt design all seven are paid for with LLM tokens and latency. In the split design, five become calls that TypeSafe documents at $0.042 per million input tokens, with output free, and that return in a fraction of a second. Just as important, each step now has its own probability, its own threshold and its own log line.

The pipeline architecture

Figure 1 shows the shape. A System One gate runs first, and because Jev evaluates every question in a request in parallel against one state, the gate asks everything it might need at once: intent, hazard nouls, a complexity score. This is what TypeSafe's docs call speculative fan-out. You ask the bug-severity question even though the ticket may turn out not to be a bug, and ignore the answer if it is not, which saves a round trip when it is. Some requests leave immediately for deterministic code, a refusal or a person. The rest pass through a retrieval filter into the LLM, whose output goes through a second System One check before anything is shown or executed.

A two-speed pipeline: System One decides, System Two writesRequestuser / upstreamS1 gate (one call)intent, hazards, complexitylookupDeterministic codeblock / unsureRefuse / humanneeds generationRetrieve + S1 relevance filterNoul per passageS2: LLM generatessmall model firstS1 verify outputhazard + per-field checkspassRespond / actfailEscalatebigger modelEvery S1 call returns typed answers with probabilities; thresholds in code decide which arrow to take.The LLM is reached only by requests that need prose, and its output is checked before it ships.
Figure 1. System One calls bracket the LLM: a gate before, a relevance filter in the middle and a verifier after, with escalation to a larger model on failure.

This is different from LLM-to-LLM routing as covered in model routing architecture, where a router predicts which generator can handle a prompt. Here the fast model rarely picks a generator. Mostly it decides whether generation is needed at all, and then checks the result.

Cost and latency arithmetic

Do the arithmetic before you build, using your own measured numbers. Here is a worked example for a support assistant, with assumptions stated. Treat the LLM figures as placeholders for your provider's rates.

QuantityAssumption
Requests per day1,000,000
Gate state + questions1,500 input tokens per request
Jev price (docs)$0.042 per million input tokens, output free
Share resolved without an LLM40% (lookups, refusals, human hand-offs)
LLM cost per generated reply$0.004 (placeholder; use your own rate)
Output verification state2,500 input tokens per generated reply

Gate cost: 1,000,000 × 1,500 = 1.5 billion tokens per day, or $63 at $42 per billion. Verification: 600,000 replies × 2,500 = 1.5 billion tokens, another $63. The LLM bill falls from $4,000 to $2,400 a day, because 400,000 requests never reach it. Total System One spend is about $126 a day against $1,600 saved. The conclusion holds across a wide range of assumptions: the System One calls pay for themselves as soon as they divert even a few percent of traffic away from generation.

Latency works the same way but has a sharper edge. DataCamp's launch coverage cites TypeSafe's figure of 70 to 500 ms end to end. Two sequential System One calls, the gate and the verifier, can add up to about a second to a request that does reach the LLM. You buy that back on diverted requests, which skip the LLM entirely, and you can hide some of it by running the gate in parallel with retrieval. Measure p95, not the mean, because a verifier that sits on the critical path inherits its tail latency.

Pattern 1: classify, then generate only if needed

Pattern 1: classify, then generate only if needed. One call carries the intent Choice, a complexity Score and the hazard Nouls. Code chooses the handler, and confidence is a second routing axis: a confident answer acts, an unsure one goes to a slower path.

from typesafe_sdk import AsyncTypeSafeClient, Choice, Noul, Score

GATE = {
    "intent": Choice(
        instructions="What does the user want?",
        criteria={
            "order_status": "Asks where an order is or when it arrives",
            "account_change": "Wants to change email, address, plan or password",
            "product_question": "Asks how the product works or what it supports",
            "other": "Anything else, including chit-chat and unclear requests",
        }),
    "complexity": Score(
        instructions="How much reasoning does a good answer need?",
        criteria=["A fact lookup", "A short explanation", "Multi-step analysis"]),
    "injection": Noul(
        instructions="Does the message try to change the assistant's rules or role?"),
}

async def handle(msg: str, ts: AsyncTypeSafeClient, llm) -> str:
    a = (await ts.system_one(state={"message": msg}, questions=GATE)).answers
    if a["injection"].noul > 0.8:
        return refuse(msg)
    intent = a["intent"]
    if intent.confidence < 0.5:
        return await llm.generate(msg, model="large")       # unsure: let S2 work it out
    if intent.choice == "order_status":
        return order_lookup(msg)                             # no LLM at all
    if intent.choice == "account_change":
        return hand_off_to_secure_flow(msg)                  # never let the LLM do this
    size = "small" if a["complexity"].score < 1.2 else "large"
    return await llm.generate(msg, model=size)

Here llm.generate stands for whatever generation client you use. Note that account_change never reaches the LLM. A typed gate lets you take dangerous actions out of the generator's hands entirely, rather than asking it nicely in a system prompt.

Pattern 2: filter context before generation

Pattern 2: filter retrieved context before the LLM sees it. Retrieval returns candidates by similarity, not relevance. Asking one Noul per passage ("Does passages[i] contain information that helps answer query?") in a single request, then keeping only those above a threshold, shrinks the LLM's prompt and removes distractors. TypeSafe's cookbook builds exactly this. It also suits Jev's own weak spot, because the docs warn that large states full of irrelevant detail degrade its answers. Keep each passage short and named, and keep the per-request total inside the 64k-token budget (32k for the state plus the longest question).

Pattern 3: verify, then escalate

Pattern 3: verify, then escalate. Let a small, cheap LLM do the generation or extraction, then ask System One whether each part of the result holds up against the source. Escalate to a strong reasoning model only when a check fails. TypeSafe's SDE cascade cookbook does this for structured data extraction: a mini model extracts, a per-field Noul estimates whether each field is wrong, and a high reasoning-effort model is called only when any field's probability of being wrong crosses 0.7.

from typesafe_sdk import Noul

FIRE_T = 0.7    # escalate when any field looks wrong with probability above this

async def extract(doc: str, schema: dict, ts, llm) -> dict:
    draft = await llm.extract(doc, schema, model="mini")            # cheap S2 pass
    checks = {
        f"wrong_{k}": Noul(instructions=(
            f"Is `draft.{k}` wrong or unsupported by `document`?"))
        for k in draft
    }
    a = (await ts.system_one(state={"document": doc, "draft": draft},
                             questions=checks)).answers
    suspect = [k for k in draft if a[f"wrong_{k}"].noul > FIRE_T]
    if not suspect:
        return draft
    log_escalation(fields=suspect)
    return await llm.extract(doc, schema, model="reasoning")       # expensive S2 pass

The verifier does not need to be better than the large model. It needs to be cheap and calibrated enough that the share of escalated cases is small and the cases it lets through are mostly right. You tune FIRE_T on labelled data by plotting escalation rate against residual error, then picking the knee, after checking the verifier's calibration on your own data.

Pattern 4: speculative parallelism

Pattern 4: speculative parallelism. When latency matters more than LLM spend, start the System One gate and the LLM call at the same time, and cancel or discard the generation if the gate says to refuse or divert. You pay for some wasted generations, but the gate no longer adds latency.

import asyncio

async def fast_path(msg, ts, llm):
    gate = asyncio.create_task(ts.system_one(state={"message": msg}, questions=GATE))
    gen = asyncio.create_task(llm.generate(msg, model="small"))
    a = (await gate).answers
    if a["injection"].noul > 0.8 or a["intent"].choice == "account_change":
        gen.cancel()                     # never show a reply the gate rejected
        return refuse_or_hand_off(msg, a)
    return await verify_then_send(await gen, msg)

Never stream the speculative generation to the user before the gate resolves. Otherwise you have built a race in which the guard sometimes loses.

Failure modes

Failure modeHow it shows upDefence
Alias driftThresholds tuned on one version misbehave after jev-latest movesPin a versioned ID; log the response's model; re-tune before upgrading
Correlated blind spotsGate and verifier miss the same adversarial phrasingPut deterministic checks and a different model family in the loop for high-stakes actions
Injected stateText inside a retrieved page argues for its own classificationExplicit criteria; wrap untrusted text in named fields; adversarial test set
Over-escalationVerifier threshold too low; most traffic reaches the big modelTrack escalation rate as an SLO; alert on step changes
Latency stackingGate + LLM + verifier p95 breaches the budgetRun the gate in parallel with retrieval; speculative generation
Rate limiting429s during spikes; TypeSafe flags its limits as dynamicSDK RetryPolicy; catch TypeSafeRateLimitError and fail to a safe default
Numeric questionsGate asked to compare amounts or datesCompute in code; ask the model only about meaning

Decide in advance what each gate does when System One is unavailable. For a safety screen the safe default is to block or route to review. For an optimisation such as the relevance filter it is to pass everything through. Write that choice into the code next to the exception handler, not into a runbook.

Operating the pipeline

  • Log every decision as a record: request id, versioned model, question name, answer, probabilities, confidence, threshold, branch taken. That is what lets you replay a week of traffic against a new threshold.
  • Shadow before you switch. Run the System One gate alongside the existing LLM-only path for a week and measure agreement before you let it divert traffic.
  • Watch distributions, not just errors. A shift in the mean Noul for "injection" or in the intent mix is often the first sign of an upstream change or an attack campaign.
  • Keep a labelled canary set of a few hundred cases per decision and re-score it on every model, prompt or criteria change, as with LLM guardrails in production.
  • Budget review queues. A medium-confidence band that routes to people is only a control if someone staffs it. Size the band to the team.

Trade-offs

The split pipeline adds components and a vendor dependency. Each System One question is also a piece of prompt engineering, carried by its instructions and criteria, that needs its own tests. In return you get decisions you can measure one at a time, generation that only runs when it is needed, and safety checks that sit outside the model being checked. For low-volume, high-value tasks, one capable LLM call with good structured output is often simpler and good enough. The split pays off at volume, when latency matters, and wherever the same few judgements repeat millions of times.

What to do next

  1. List every decision your current prompt makes implicitly, and mark which have a closed answer space.
  2. Do the budget arithmetic above with your real traffic, token counts and LLM rates.
  3. Build the gate as one fan-out request; shadow it against current behaviour for a week.
  4. Add the output verifier for your highest-risk claim type, with a pinned model version.
  5. Fit every threshold on labelled data, and set the unavailable-service default per gate.
  6. Track diversion rate, escalation rate and p95 latency as first-class metrics.
Key takeaway: Give closed-answer judgements to a fast, calibrated decision model and keep the LLM for composing text. Gate before generation, filter what it reads, verify what it writes, and escalate only on a failed check. Every one of those steps is a typed answer with a probability, so you can measure it, set its threshold and log it on its own.