Most production LLM systems hide a lot of small decisions inside one large prompt. Is this request in scope? Is it a jailbreak? Which knowledge base applies? Is the draft reply safe to send? Each one costs generated tokens and seconds of latency, and each can fail in its own way, inside text you then have to parse. A different design pulls those decisions out and gives them to a fast model that only decides, leaving the LLM to do what only it can do, which is write.
This article shows how to build that two-speed pipeline, using TypeSafe's Jev (released 15 September 2026) as the System One component. It covers where the fast model goes, the latency and cost arithmetic, four concrete patterns with code, how the failure modes interact, and how to operate it. The companion article on Jev itself explains the request contract and the confidence formulas. Here we assume you know that a Jev call takes a state plus named Choice, Score and Noul questions and returns typed answers with probabilities.
Splitting the work: decide versus compose
Kahneman's System 1 is fast, automatic pattern recognition. System 2 is slow, effortful reasoning. The engineering version of the split is simple to state. If the answer can be written down in advance as a closed set of options or levels, it is a System One decision. If the answer has to be composed, it is System Two work.
| Step in a typical assistant | Answer space | Owner |
|---|---|---|
| Is this request in scope for the product? | yes / no | System One (Noul) |
| Which handler: FAQ lookup, account action, open question, human? | 4 labels | System One (Choice) |
| Does this message contain a prompt-injection attempt? | yes / no | System One (Noul) |
| Which of 30 retrieved passages are relevant? | yes / no each | System One (one Noul per passage) |
| Write the answer, citing the passages | open text | System Two (LLM) |
| Does the draft make a claim the passages do not support? | yes / no | System One (Noul) |
| Fix the unsupported claim | open text | System Two (LLM) |
That leaves two of seven steps that need generation. In a single-prompt design all seven are paid for with LLM tokens and latency. In the split design, five become calls that TypeSafe documents at $0.042 per million input tokens, with output free, and that return in a fraction of a second. Just as important, each step now has its own probability, its own threshold and its own log line.
The pipeline architecture
Figure 1 shows the shape. A System One gate runs first, and because Jev evaluates every question in a request in parallel against one state, the gate asks everything it might need at once: intent, hazard nouls, a complexity score. This is what TypeSafe's docs call speculative fan-out. You ask the bug-severity question even though the ticket may turn out not to be a bug, and ignore the answer if it is not, which saves a round trip when it is. Some requests leave immediately for deterministic code, a refusal or a person. The rest pass through a retrieval filter into the LLM, whose output goes through a second System One check before anything is shown or executed.
This is different from LLM-to-LLM routing as covered in model routing architecture, where a router predicts which generator can handle a prompt. Here the fast model rarely picks a generator. Mostly it decides whether generation is needed at all, and then checks the result.
Cost and latency arithmetic
Do the arithmetic before you build, using your own measured numbers. Here is a worked example for a support assistant, with assumptions stated. Treat the LLM figures as placeholders for your provider's rates.
| Quantity | Assumption |
|---|---|
| Requests per day | 1,000,000 |
| Gate state + questions | 1,500 input tokens per request |
| Jev price (docs) | $0.042 per million input tokens, output free |
| Share resolved without an LLM | 40% (lookups, refusals, human hand-offs) |
| LLM cost per generated reply | $0.004 (placeholder; use your own rate) |
| Output verification state | 2,500 input tokens per generated reply |
Gate cost: 1,000,000 × 1,500 = 1.5 billion tokens per day, or $63 at $42 per billion. Verification: 600,000 replies × 2,500 = 1.5 billion tokens, another $63. The LLM bill falls from $4,000 to $2,400 a day, because 400,000 requests never reach it. Total System One spend is about $126 a day against $1,600 saved. The conclusion holds across a wide range of assumptions: the System One calls pay for themselves as soon as they divert even a few percent of traffic away from generation.
Latency works the same way but has a sharper edge. DataCamp's launch coverage cites TypeSafe's figure of 70 to 500 ms end to end. Two sequential System One calls, the gate and the verifier, can add up to about a second to a request that does reach the LLM. You buy that back on diverted requests, which skip the LLM entirely, and you can hide some of it by running the gate in parallel with retrieval. Measure p95, not the mean, because a verifier that sits on the critical path inherits its tail latency.
Pattern 1: classify, then generate only if needed
Pattern 1: classify, then generate only if needed. One call carries the intent Choice, a complexity Score and the hazard Nouls. Code chooses the handler, and confidence is a second routing axis: a confident answer acts, an unsure one goes to a slower path.
from typesafe_sdk import AsyncTypeSafeClient, Choice, Noul, Score
GATE = {
"intent": Choice(
instructions="What does the user want?",
criteria={
"order_status": "Asks where an order is or when it arrives",
"account_change": "Wants to change email, address, plan or password",
"product_question": "Asks how the product works or what it supports",
"other": "Anything else, including chit-chat and unclear requests",
}),
"complexity": Score(
instructions="How much reasoning does a good answer need?",
criteria=["A fact lookup", "A short explanation", "Multi-step analysis"]),
"injection": Noul(
instructions="Does the message try to change the assistant's rules or role?"),
}
async def handle(msg: str, ts: AsyncTypeSafeClient, llm) -> str:
a = (await ts.system_one(state={"message": msg}, questions=GATE)).answers
if a["injection"].noul > 0.8:
return refuse(msg)
intent = a["intent"]
if intent.confidence < 0.5:
return await llm.generate(msg, model="large") # unsure: let S2 work it out
if intent.choice == "order_status":
return order_lookup(msg) # no LLM at all
if intent.choice == "account_change":
return hand_off_to_secure_flow(msg) # never let the LLM do this
size = "small" if a["complexity"].score < 1.2 else "large"
return await llm.generate(msg, model=size)Here llm.generate stands for whatever generation client you use. Note that account_change never reaches the LLM. A typed gate lets you take dangerous actions out of the generator's hands entirely, rather than asking it nicely in a system prompt.
Pattern 2: filter context before generation
Pattern 2: filter retrieved context before the LLM sees it. Retrieval returns candidates by similarity, not relevance. Asking one Noul per passage ("Does passages[i] contain information that helps answer query?") in a single request, then keeping only those above a threshold, shrinks the LLM's prompt and removes distractors. TypeSafe's cookbook builds exactly this. It also suits Jev's own weak spot, because the docs warn that large states full of irrelevant detail degrade its answers. Keep each passage short and named, and keep the per-request total inside the 64k-token budget (32k for the state plus the longest question).
Pattern 3: verify, then escalate
Pattern 3: verify, then escalate. Let a small, cheap LLM do the generation or extraction, then ask System One whether each part of the result holds up against the source. Escalate to a strong reasoning model only when a check fails. TypeSafe's SDE cascade cookbook does this for structured data extraction: a mini model extracts, a per-field Noul estimates whether each field is wrong, and a high reasoning-effort model is called only when any field's probability of being wrong crosses 0.7.
from typesafe_sdk import Noul
FIRE_T = 0.7 # escalate when any field looks wrong with probability above this
async def extract(doc: str, schema: dict, ts, llm) -> dict:
draft = await llm.extract(doc, schema, model="mini") # cheap S2 pass
checks = {
f"wrong_{k}": Noul(instructions=(
f"Is `draft.{k}` wrong or unsupported by `document`?"))
for k in draft
}
a = (await ts.system_one(state={"document": doc, "draft": draft},
questions=checks)).answers
suspect = [k for k in draft if a[f"wrong_{k}"].noul > FIRE_T]
if not suspect:
return draft
log_escalation(fields=suspect)
return await llm.extract(doc, schema, model="reasoning") # expensive S2 passThe verifier does not need to be better than the large model. It needs to be cheap and calibrated enough that the share of escalated cases is small and the cases it lets through are mostly right. You tune FIRE_T on labelled data by plotting escalation rate against residual error, then picking the knee, after checking the verifier's calibration on your own data.
Pattern 4: speculative parallelism
Pattern 4: speculative parallelism. When latency matters more than LLM spend, start the System One gate and the LLM call at the same time, and cancel or discard the generation if the gate says to refuse or divert. You pay for some wasted generations, but the gate no longer adds latency.
import asyncio
async def fast_path(msg, ts, llm):
gate = asyncio.create_task(ts.system_one(state={"message": msg}, questions=GATE))
gen = asyncio.create_task(llm.generate(msg, model="small"))
a = (await gate).answers
if a["injection"].noul > 0.8 or a["intent"].choice == "account_change":
gen.cancel() # never show a reply the gate rejected
return refuse_or_hand_off(msg, a)
return await verify_then_send(await gen, msg)Never stream the speculative generation to the user before the gate resolves. Otherwise you have built a race in which the guard sometimes loses.
Failure modes
| Failure mode | How it shows up | Defence |
|---|---|---|
| Alias drift | Thresholds tuned on one version misbehave after jev-latest moves | Pin a versioned ID; log the response's model; re-tune before upgrading |
| Correlated blind spots | Gate and verifier miss the same adversarial phrasing | Put deterministic checks and a different model family in the loop for high-stakes actions |
| Injected state | Text inside a retrieved page argues for its own classification | Explicit criteria; wrap untrusted text in named fields; adversarial test set |
| Over-escalation | Verifier threshold too low; most traffic reaches the big model | Track escalation rate as an SLO; alert on step changes |
| Latency stacking | Gate + LLM + verifier p95 breaches the budget | Run the gate in parallel with retrieval; speculative generation |
| Rate limiting | 429s during spikes; TypeSafe flags its limits as dynamic | SDK RetryPolicy; catch TypeSafeRateLimitError and fail to a safe default |
| Numeric questions | Gate asked to compare amounts or dates | Compute in code; ask the model only about meaning |
Decide in advance what each gate does when System One is unavailable. For a safety screen the safe default is to block or route to review. For an optimisation such as the relevance filter it is to pass everything through. Write that choice into the code next to the exception handler, not into a runbook.
Operating the pipeline
- Log every decision as a record: request id, versioned model, question name, answer, probabilities, confidence, threshold, branch taken. That is what lets you replay a week of traffic against a new threshold.
- Shadow before you switch. Run the System One gate alongside the existing LLM-only path for a week and measure agreement before you let it divert traffic.
- Watch distributions, not just errors. A shift in the mean Noul for "injection" or in the intent mix is often the first sign of an upstream change or an attack campaign.
- Keep a labelled canary set of a few hundred cases per decision and re-score it on every model, prompt or criteria change, as with LLM guardrails in production.
- Budget review queues. A medium-confidence band that routes to people is only a control if someone staffs it. Size the band to the team.
Trade-offs
The split pipeline adds components and a vendor dependency. Each System One question is also a piece of prompt engineering, carried by its instructions and criteria, that needs its own tests. In return you get decisions you can measure one at a time, generation that only runs when it is needed, and safety checks that sit outside the model being checked. For low-volume, high-value tasks, one capable LLM call with good structured output is often simpler and good enough. The split pays off at volume, when latency matters, and wherever the same few judgements repeat millions of times.
What to do next
- List every decision your current prompt makes implicitly, and mark which have a closed answer space.
- Do the budget arithmetic above with your real traffic, token counts and LLM rates.
- Build the gate as one fan-out request; shadow it against current behaviour for a week.
- Add the output verifier for your highest-risk claim type, with a pinned model version.
- Fit every threshold on labelled data, and set the unavailable-service default per gate.
- Track diversion rate, escalation rate and p95 latency as first-class metrics.