A calibrated model tells you how likely something is. It does not tell you what to do. The step from "p = 0.83 that this refund claim is legitimate" to "approve it", "deny it" or "send it to a person" is a separate piece of software with its own inputs: what each mistake costs, what a human review costs, how much reviewer capacity exists today, and how much evidence you have that the probabilities are still right. Most teams write it as a single hard-coded if p > 0.5 and never revisit it.

This article is about that decision layer. It assumes you already have calibrated probabilities; if not, start with model calibration, which covers reliability diagrams, ECE and temperature scaling, and is not repeated here. We derive thresholds from costs, add abstain and escalate options, price model cascades, fit thresholds with a finite-sample guarantee and size the human queue, using one refund-approval example throughout.

The decision layer is its own component

Treat the decision policy as a component with a version number, not as a constant inside the model server. Its input is a probability vector over outcomes; its output is one action from a finite set; its parameters are a cost matrix and a few thresholds. Separating it lets you change behaviour without retraining, log which policy version made each decision, and test it in isolation as a pure function. It only works if a probability of 0.9 is still right about 90 percent of the time on today's traffic.

Three actions on one calibrated probabilitydenyrefer to a reviewer (cost 6)ok0.00.150.750.951.0p = calibrated probability that the refund claim is legitimatedashed line: the two-action Bayes threshold C_fp / (C_fp + C_fn) = 120 / 160Calibrated modelemits p per requestDecision policycosts + thresholds, versionedActiondeny / refer / approvepaOutcome logp, action, label laterre-fit thresholds
The refund example: two cost-derived cut points split the probability axis into deny, refer and approve. The policy is versioned, and its outcomes feed the next threshold fit.

One threshold from a cost matrix

Start with two actions. Let p be the calibrated probability that a claim is legitimate. Approving a fraudulent claim costs C_fp (the money plus handling). Denying a legitimate one costs C_fn (support contact, goodwill, some churn). Correct decisions cost zero, which loses nothing because only cost differences matter.

The expected cost of approving is (1 - p) C_fp, since you are wrong whenever the claim is fraudulent. The expected cost of denying is p C_fn. Approve when the first is smaller:

(1 - p) * C_fp  <  p * C_fn
        C_fp    <  p * (C_fp + C_fn)
           p    >  C_fp / (C_fp + C_fn)        # the Bayes threshold t

With C_fp = 120 dollars and C_fn = 40 dollars, t = 120 / 160 = 0.75. The threshold of 0.5 is correct only when the two errors cost the same, which is almost never true. Notice also that the threshold ignores the fraud base rate: a calibrated p already contains it, so "tuning for imbalance" on top double-counts it.

Adding abstain: a reject band priced in money

Now add a third action, refer to a human, with cost C_r per review. Assume for now that the reviewer is always right (we return to that). Refer when its fixed cost beats both automatic actions. This is Chow's reject rule written with costs:

cost(approve) = (1 - p) * C_fp
cost(deny)    = p * C_fn
cost(refer)   = C_r

refer  when  p * C_fn > C_r   and  (1 - p) * C_fp > C_r
       i.e.  C_r / C_fn  <  p  <  1 - C_r / C_fp

the band is non-empty only if  C_r < C_fp * C_fn / (C_fp + C_fn)

In the example C_r = 6. The lower edge is 6 / 40 = 0.15 and the upper edge is 1 - 6 / 120 = 0.95. So the policy denies below 0.15, approves above 0.95 and refers everything between. The band condition holds because 6 is below 120 x 40 / 160 = 30. If review cost 35 dollars there would be no band: no probability would make a review worth paying for, and the right policy is the plain threshold at 0.75.

The band is asymmetric, reaching far toward 1 because false approvals are expensive, and its width is set by money rather than intuition: halve review cost and it widens.

Many outcomes, many actions: the policy object

Real policies have more outcomes and more actions: approve, deny, approve with a hold, request a document, refer. The general rule is one line. With outcome probabilities p(y) and a cost table C[a][y], pick the action with the smallest expected cost. The code below is the whole policy object, including versioning and the reason string you will want in logs.

from dataclasses import dataclass
import numpy as np

@dataclass(frozen=True)
class DecisionPolicy:
    version: str
    actions: tuple            # e.g. ("approve", "deny", "refer")
    outcomes: tuple           # e.g. ("legit", "fraud")
    cost: np.ndarray          # shape (n_actions, n_outcomes), in dollars

    def expected_costs(self, probs):
        probs = np.asarray(probs, dtype=float)
        if probs.shape != (len(self.outcomes),) or abs(probs.sum() - 1) > 1e-6:
            raise ValueError("probs must be a distribution over outcomes")
        return self.cost @ probs

    def decide(self, probs):
        ec = self.expected_costs(probs)
        i = int(np.argmin(ec))
        runner_up = float(np.partition(ec, 1)[1]) if len(ec) > 1 else float("inf")
        return {
            "action": self.actions[i],
            "expected_cost": float(ec[i]),
            "margin": runner_up - float(ec[i]),   # small margin = fragile decision
            "policy_version": self.version,
        }

refund_v3 = DecisionPolicy(
    version="refund-2026-10-07",
    actions=("approve", "deny", "refer"),
    outcomes=("legit", "fraud"),
    cost=np.array([[0.0, 120.0],     # approve: free if legit, 120 if fraud
                   [40.0, 0.0],      # deny: 40 if legit, free if fraud
                   [6.0, 6.0]]),     # refer: flat review cost
)
print(refund_v3.decide([0.83, 0.17]))   # refer: approve would cost 20.4, deny 33.2

Log the margin: small-margin decisions are the ones that flip when costs or calibration move.

Escalation cascades and cost-based routing

Model escalation sends a request to a slower, more accurate model only when the cheap one is unsure. The same arithmetic applies once you count the second model's compute.

Define the value of information of calling the large model as the expected decision cost if you act now on p1, minus the expected decision cost after you see p2. Escalate when that value exceeds the large model's cost. You cannot compute p2 before calling it, but you can estimate the value per bucket of p1 from logs where both models scored the same requests and the label later arrived.

Cost-based cascade: escalate only when it lowers expected total costRequestfeatures, textSmall modelp1, cost 0.02Large modelp2, cost 0.30Humancost 6.00VOI above 0.30still in bandAct on p1confident tailAct on p2now confidentVOI = expected decision cost now minus expected cost after seeing p2estimated offline from paired (p1, p2, label) logs, per bucket of p1
A three-tier cascade. Each arrow is taken only when the estimated drop in decision cost for that bucket of the current probability exceeds the next tier's price.
def voi_table(p1, p2, y, policy, bins=20):
    """Per-bucket value of calling model 2, from paired offline scores and labels.
    p1, p2: arrays of P(legit); y: 1 if legit. Returns {bucket: dollars saved}."""
    out = {}
    edges = np.linspace(0, 1, bins + 1)
    for b in range(bins):
        m = (p1 >= edges[b]) & (p1 < edges[b + 1])
        if m.sum() < 200:            # too few samples: do not trust the estimate
            continue
        def realised(ps):
            acts = [policy.decide([q, 1 - q])["action"] for q in ps]
            idx = [policy.actions.index(a) for a in acts]
            return np.mean([policy.cost[i, 0 if yy else 1] for i, yy in zip(idx, y[m])])
        out[b] = realised(p1[m]) - realised(p2[m])
    return out

# escalate a request whose p1 falls in bucket b iff voi[b] > cost_of_model_2

Using realised costs on labelled data protects you if the large model is miscalibrated, and in the confident tails the value is near zero, so spend goes where p1 is ambiguous. For router architectures see LLM routers, and for routing by topic rather than confidence, semantic routing.

Choosing thresholds from data with a guarantee

Often you want a guarantee stated as a rate: "at most 1 percent of auto-approved claims are fraudulent". Set it from held-out data with a margin for sampling error.

The procedure: sort held-out examples by p, scan candidate thresholds from high to low, and for each compute an upper confidence bound on the error rate among examples above it. Keep the lowest threshold, the most coverage, whose bound is under the target.

from scipy.stats import beta

def clopper_pearson_upper(errors, n, delta=0.05):
    return 1.0 if n == 0 else beta.ppf(1 - delta, errors + 1, n - errors)

def lowest_safe_threshold(p, is_error, target=0.01, delta=0.05, min_n=300):
    """p: P(legit) on held-out data; is_error: 1 if auto-approving would be wrong.
    Returns (threshold, coverage) or None if no threshold meets the target."""
    order = np.argsort(-p)
    p_sorted, e_sorted = p[order], is_error[order]
    best = None
    errs = 0
    for k in range(len(p_sorted)):
        errs += e_sorted[k]
        n = k + 1
        if n >= min_n and clopper_pearson_upper(errs, n, delta) <= target:
            best = (float(p_sorted[k]), n / len(p_sorted))
    return best

Coverage against the error bound is the risk-coverage curve a product owner chooses a point on. Keeping the best of many thresholds is a mild multiple-testing problem; the Learn-then-Test framework fixes it by stopping at the first failure, one break in the loop above.

When reviewer capacity is the constraint

The refer band assumes unlimited reviewers. They are not. Suppose 10,000 claims arrive per day and the calibrated distribution of p puts 2,600 of them inside 0.15 to 0.95, but the team can review 1,200. Narrowing the band symmetrically is the wrong fix. The right one is to rank band members by how much a review saves and review the top 1,200.

The saving from reviewing one claim is min(cost(approve), cost(deny)) - C_r. It is largest near the two-action threshold of 0.75 and smallest near the band edges. Unreviewed claims fall back to the cheaper automatic action, and the saving of the last reviewed item is the shadow price of one more reviewer.

import heapq

def assign_reviews(batch, policy, capacity):
    """batch: list of (claim_id, p_legit). Returns {claim_id: action}."""
    heap, result = [], {}
    for cid, p in batch:
        ec = policy.expected_costs([p, 1 - p])
        auto = int(np.argmin(ec[:2]))                  # best of approve / deny
        saving = ec[auto] - policy.cost[2, 0]          # minus flat review cost
        result[cid] = policy.actions[auto]
        if saving > 0:
            heapq.heappush(heap, (-saving, cid))
    for _ in range(min(capacity, len(heap))):
        _, cid = heapq.heappop(heap)
        result[cid] = "refer"
    return result

Worked example: one day of refunds

Putting the refund pieces together for one day. The small model scores all 10,000 claims at 2 cents each. The value-of-information table says calling the large model at 30 cents pays only for p1 between 0.4 and 0.97, which is 3,100 claims. After the large model, 1,900 claims sit inside the refer band and are ranked by saving; the top 1,200 go to reviewers and the remaining 700 take their cheaper automatic action. The risk-coverage check on last month's labels confirms that auto-approvals above 0.95 had a Clopper-Pearson upper bound of 0.8 percent fraud, under the 1 percent target.

Compute costs 200 dollars for the small model plus 930 for the large one; review costs 7,200. The numbers are illustrative; use your own logged costs.

Failure modes

The policy is simple. The ways it goes wrong are mostly about its inputs.

  • Prior shift. If fraud doubles, a model calibrated last quarter is now overconfident about legitimacy. When only the base rate changes, correct the odds directly: multiply p / (1 - p) by the ratio of the new prior odds to the old, then convert back. Detect it by watching the mean predicted p against the labelled rate, as in drift detection.
  • Selective labels. Denied claims rarely reveal their true outcome, so refits drift toward approving less. Route a random 1 percent of denials to review to keep labels flowing.
  • Stale costs. Store the cost matrix with the policy version and review it on a calendar.
  • Imperfect reviewers. If reviewers err at rate r, referring costs C_r plus r times the error cost. Measure r with blind double review.
  • Segment miscalibration. Global calibration can hide a segment that is off by 15 points. Check calibration per major segment and allow per-segment thresholds where the costs differ legitimately.

Operating it

Log the probability vector, policy version, action, expected cost, margin and later the label; then you can replay any new policy against history before shipping the winner behind a real experiment as in A/B testing for AI systems. Alert on the refer rate, mean p against the observed positive rate, and review queue age. For LLM-based classifiers, take p from the probability of a single decision token or a constrained label set, not from a number the model writes in prose, and recalibrate it like any other score; how guardrails consume such scores is covered in guardrails in production.

Trade-offs

ChoiceGainCost
Cost-derived thresholdsDefensible, adapt when costs changeNeed honest cost estimates
Abstain bandPays for review only where it saves moneyNeeds capacity and reviewer quality data
Model cascadeBig model spend goes to ambiguous casesPaired logs, a second calibration to maintain
Bounded error-rate thresholdA guarantee you can stateLower coverage than the point estimate
Per-segment thresholdsFixes segment miscalibrationMore parameters, fairness review needed

What to do next

  1. Write the cost matrix for one decision with the people who own the money, including a review cost.
  2. Check calibration on the last 30 days, per segment, before trusting any threshold.
  3. Compute t = C_fp / (C_fp + C_fn) and the refer band; compare them with the threshold you use today.
  4. Wrap the policy in a versioned, pure-function object and log probability, version, action, expected cost and margin.
  5. Fit an error-rate-bounded threshold with Clopper-Pearson on held-out labels and plot the risk-coverage curve.
  6. Rank the refer band by review saving and cap it at real reviewer capacity.
  7. If you run two models, collect paired scores and build the per-bucket value table before adding a cascade.
  8. Add a 1 percent exploration slice and a prior-shift alert.
Key takeaway: A calibrated probability becomes a good decision only through explicit costs. Derive the threshold from the cost matrix, add a refer band only where review is cheaper than the expected error, escalate to bigger models only where the value of information beats their price, bound error rates with held-out data, and rank reviews by saving when capacity runs out.