A calibrated model tells you how likely something is. It does not tell you what to do. The step from "p = 0.83 that this refund claim is legitimate" to "approve it", "deny it" or "send it to a person" is a separate piece of software with its own inputs: what each mistake costs, what a human review costs, how much reviewer capacity exists today, and how much evidence you have that the probabilities are still right. Most teams write it as a single hard-coded if p > 0.5 and never revisit it.
This article is about that decision layer. It assumes you already have calibrated probabilities; if not, start with model calibration, which covers reliability diagrams, ECE and temperature scaling, and is not repeated here. We derive thresholds from costs, add abstain and escalate options, price model cascades, fit thresholds with a finite-sample guarantee and size the human queue, using one refund-approval example throughout.
The decision layer is its own component
Treat the decision policy as a component with a version number, not as a constant inside the model server. Its input is a probability vector over outcomes; its output is one action from a finite set; its parameters are a cost matrix and a few thresholds. Separating it lets you change behaviour without retraining, log which policy version made each decision, and test it in isolation as a pure function. It only works if a probability of 0.9 is still right about 90 percent of the time on today's traffic.
One threshold from a cost matrix
Start with two actions. Let p be the calibrated probability that a claim is legitimate. Approving a fraudulent claim costs C_fp (the money plus handling). Denying a legitimate one costs C_fn (support contact, goodwill, some churn). Correct decisions cost zero, which loses nothing because only cost differences matter.
The expected cost of approving is (1 - p) C_fp, since you are wrong whenever the claim is fraudulent. The expected cost of denying is p C_fn. Approve when the first is smaller:
(1 - p) * C_fp < p * C_fn
C_fp < p * (C_fp + C_fn)
p > C_fp / (C_fp + C_fn) # the Bayes threshold tWith C_fp = 120 dollars and C_fn = 40 dollars, t = 120 / 160 = 0.75. The threshold of 0.5 is correct only when the two errors cost the same, which is almost never true. Notice also that the threshold ignores the fraud base rate: a calibrated p already contains it, so "tuning for imbalance" on top double-counts it.
Adding abstain: a reject band priced in money
Now add a third action, refer to a human, with cost C_r per review. Assume for now that the reviewer is always right (we return to that). Refer when its fixed cost beats both automatic actions. This is Chow's reject rule written with costs:
cost(approve) = (1 - p) * C_fp
cost(deny) = p * C_fn
cost(refer) = C_r
refer when p * C_fn > C_r and (1 - p) * C_fp > C_r
i.e. C_r / C_fn < p < 1 - C_r / C_fp
the band is non-empty only if C_r < C_fp * C_fn / (C_fp + C_fn)In the example C_r = 6. The lower edge is 6 / 40 = 0.15 and the upper edge is 1 - 6 / 120 = 0.95. So the policy denies below 0.15, approves above 0.95 and refers everything between. The band condition holds because 6 is below 120 x 40 / 160 = 30. If review cost 35 dollars there would be no band: no probability would make a review worth paying for, and the right policy is the plain threshold at 0.75.
The band is asymmetric, reaching far toward 1 because false approvals are expensive, and its width is set by money rather than intuition: halve review cost and it widens.
Many outcomes, many actions: the policy object
Real policies have more outcomes and more actions: approve, deny, approve with a hold, request a document, refer. The general rule is one line. With outcome probabilities p(y) and a cost table C[a][y], pick the action with the smallest expected cost. The code below is the whole policy object, including versioning and the reason string you will want in logs.
from dataclasses import dataclass
import numpy as np
@dataclass(frozen=True)
class DecisionPolicy:
version: str
actions: tuple # e.g. ("approve", "deny", "refer")
outcomes: tuple # e.g. ("legit", "fraud")
cost: np.ndarray # shape (n_actions, n_outcomes), in dollars
def expected_costs(self, probs):
probs = np.asarray(probs, dtype=float)
if probs.shape != (len(self.outcomes),) or abs(probs.sum() - 1) > 1e-6:
raise ValueError("probs must be a distribution over outcomes")
return self.cost @ probs
def decide(self, probs):
ec = self.expected_costs(probs)
i = int(np.argmin(ec))
runner_up = float(np.partition(ec, 1)[1]) if len(ec) > 1 else float("inf")
return {
"action": self.actions[i],
"expected_cost": float(ec[i]),
"margin": runner_up - float(ec[i]), # small margin = fragile decision
"policy_version": self.version,
}
refund_v3 = DecisionPolicy(
version="refund-2026-10-07",
actions=("approve", "deny", "refer"),
outcomes=("legit", "fraud"),
cost=np.array([[0.0, 120.0], # approve: free if legit, 120 if fraud
[40.0, 0.0], # deny: 40 if legit, free if fraud
[6.0, 6.0]]), # refer: flat review cost
)
print(refund_v3.decide([0.83, 0.17])) # refer: approve would cost 20.4, deny 33.2Log the margin: small-margin decisions are the ones that flip when costs or calibration move.
Escalation cascades and cost-based routing
Model escalation sends a request to a slower, more accurate model only when the cheap one is unsure. The same arithmetic applies once you count the second model's compute.
Define the value of information of calling the large model as the expected decision cost if you act now on p1, minus the expected decision cost after you see p2. Escalate when that value exceeds the large model's cost. You cannot compute p2 before calling it, but you can estimate the value per bucket of p1 from logs where both models scored the same requests and the label later arrived.
def voi_table(p1, p2, y, policy, bins=20):
"""Per-bucket value of calling model 2, from paired offline scores and labels.
p1, p2: arrays of P(legit); y: 1 if legit. Returns {bucket: dollars saved}."""
out = {}
edges = np.linspace(0, 1, bins + 1)
for b in range(bins):
m = (p1 >= edges[b]) & (p1 < edges[b + 1])
if m.sum() < 200: # too few samples: do not trust the estimate
continue
def realised(ps):
acts = [policy.decide([q, 1 - q])["action"] for q in ps]
idx = [policy.actions.index(a) for a in acts]
return np.mean([policy.cost[i, 0 if yy else 1] for i, yy in zip(idx, y[m])])
out[b] = realised(p1[m]) - realised(p2[m])
return out
# escalate a request whose p1 falls in bucket b iff voi[b] > cost_of_model_2Using realised costs on labelled data protects you if the large model is miscalibrated, and in the confident tails the value is near zero, so spend goes where p1 is ambiguous. For router architectures see LLM routers, and for routing by topic rather than confidence, semantic routing.
Choosing thresholds from data with a guarantee
Often you want a guarantee stated as a rate: "at most 1 percent of auto-approved claims are fraudulent". Set it from held-out data with a margin for sampling error.
The procedure: sort held-out examples by p, scan candidate thresholds from high to low, and for each compute an upper confidence bound on the error rate among examples above it. Keep the lowest threshold, the most coverage, whose bound is under the target.
from scipy.stats import beta
def clopper_pearson_upper(errors, n, delta=0.05):
return 1.0 if n == 0 else beta.ppf(1 - delta, errors + 1, n - errors)
def lowest_safe_threshold(p, is_error, target=0.01, delta=0.05, min_n=300):
"""p: P(legit) on held-out data; is_error: 1 if auto-approving would be wrong.
Returns (threshold, coverage) or None if no threshold meets the target."""
order = np.argsort(-p)
p_sorted, e_sorted = p[order], is_error[order]
best = None
errs = 0
for k in range(len(p_sorted)):
errs += e_sorted[k]
n = k + 1
if n >= min_n and clopper_pearson_upper(errs, n, delta) <= target:
best = (float(p_sorted[k]), n / len(p_sorted))
return bestCoverage against the error bound is the risk-coverage curve a product owner chooses a point on. Keeping the best of many thresholds is a mild multiple-testing problem; the Learn-then-Test framework fixes it by stopping at the first failure, one break in the loop above.
When reviewer capacity is the constraint
The refer band assumes unlimited reviewers. They are not. Suppose 10,000 claims arrive per day and the calibrated distribution of p puts 2,600 of them inside 0.15 to 0.95, but the team can review 1,200. Narrowing the band symmetrically is the wrong fix. The right one is to rank band members by how much a review saves and review the top 1,200.
The saving from reviewing one claim is min(cost(approve), cost(deny)) - C_r. It is largest near the two-action threshold of 0.75 and smallest near the band edges. Unreviewed claims fall back to the cheaper automatic action, and the saving of the last reviewed item is the shadow price of one more reviewer.
import heapq
def assign_reviews(batch, policy, capacity):
"""batch: list of (claim_id, p_legit). Returns {claim_id: action}."""
heap, result = [], {}
for cid, p in batch:
ec = policy.expected_costs([p, 1 - p])
auto = int(np.argmin(ec[:2])) # best of approve / deny
saving = ec[auto] - policy.cost[2, 0] # minus flat review cost
result[cid] = policy.actions[auto]
if saving > 0:
heapq.heappush(heap, (-saving, cid))
for _ in range(min(capacity, len(heap))):
_, cid = heapq.heappop(heap)
result[cid] = "refer"
return result
Worked example: one day of refunds
Putting the refund pieces together for one day. The small model scores all 10,000 claims at 2 cents each. The value-of-information table says calling the large model at 30 cents pays only for p1 between 0.4 and 0.97, which is 3,100 claims. After the large model, 1,900 claims sit inside the refer band and are ranked by saving; the top 1,200 go to reviewers and the remaining 700 take their cheaper automatic action. The risk-coverage check on last month's labels confirms that auto-approvals above 0.95 had a Clopper-Pearson upper bound of 0.8 percent fraud, under the 1 percent target.
Compute costs 200 dollars for the small model plus 930 for the large one; review costs 7,200. The numbers are illustrative; use your own logged costs.
Failure modes
The policy is simple. The ways it goes wrong are mostly about its inputs.
- Prior shift. If fraud doubles, a model calibrated last quarter is now overconfident about legitimacy. When only the base rate changes, correct the odds directly: multiply p / (1 - p) by the ratio of the new prior odds to the old, then convert back. Detect it by watching the mean predicted p against the labelled rate, as in drift detection.
- Selective labels. Denied claims rarely reveal their true outcome, so refits drift toward approving less. Route a random 1 percent of denials to review to keep labels flowing.
- Stale costs. Store the cost matrix with the policy version and review it on a calendar.
- Imperfect reviewers. If reviewers err at rate r, referring costs C_r plus r times the error cost. Measure r with blind double review.
- Segment miscalibration. Global calibration can hide a segment that is off by 15 points. Check calibration per major segment and allow per-segment thresholds where the costs differ legitimately.
Operating it
Log the probability vector, policy version, action, expected cost, margin and later the label; then you can replay any new policy against history before shipping the winner behind a real experiment as in A/B testing for AI systems. Alert on the refer rate, mean p against the observed positive rate, and review queue age. For LLM-based classifiers, take p from the probability of a single decision token or a constrained label set, not from a number the model writes in prose, and recalibrate it like any other score; how guardrails consume such scores is covered in guardrails in production.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Cost-derived thresholds | Defensible, adapt when costs change | Need honest cost estimates |
| Abstain band | Pays for review only where it saves money | Needs capacity and reviewer quality data |
| Model cascade | Big model spend goes to ambiguous cases | Paired logs, a second calibration to maintain |
| Bounded error-rate threshold | A guarantee you can state | Lower coverage than the point estimate |
| Per-segment thresholds | Fixes segment miscalibration | More parameters, fairness review needed |
What to do next
- Write the cost matrix for one decision with the people who own the money, including a review cost.
- Check calibration on the last 30 days, per segment, before trusting any threshold.
- Compute t = C_fp / (C_fp + C_fn) and the refer band; compare them with the threshold you use today.
- Wrap the policy in a versioned, pure-function object and log probability, version, action, expected cost and margin.
- Fit an error-rate-bounded threshold with Clopper-Pearson on held-out labels and plot the risk-coverage curve.
- Rank the refer band by review saving and cap it at real reviewer capacity.
- If you run two models, collect paired scores and build the per-bucket value table before adding a cascade.
- Add a 1 percent exploration slice and a prior-shift alert.