Courts are adopting AI in two very different ways. One is quiet and administrative: transcription, translation, scheduling, routing filings to the right clerk, helping self-represented litigants fill in forms. The other touches the decision itself: risk scores read at sentencing or bail, research memos drafted by a language model, summaries of a long record prepared for a judge, and machine-generated outputs offered as evidence. The second group is where liberty, property and the legitimacy of the court are at stake, and it is where engineering choices quietly become due-process choices.
This article is written for engineers and technical leads who build or evaluate systems used by courts, law firms and legal-aid groups. It covers what the leading case on risk scores actually decided, why language models invent case law and how to build a gate that catches it, the design rules a decision-support tool needs before a judge should see it, how machine-generated evidence is being handled, and the failure modes to plan for. It is not legal advice; rules differ by jurisdiction and change, so confirm the current local rules before you rely on any of them.
Where AI touches a court
Start by sorting uses by how close they sit to the decision and how hard an error is to undo. A mistranslated word in a hearing transcript, a wrong citation in a brief and a miscalibrated risk score are all model errors, but they flow into the record by different paths and need different controls.
| Use | Who relies on it | Typical error | Primary control |
|---|---|---|---|
| Transcription, translation | Judge, parties, appeal court | Wrong word changes meaning | Certified human review of contested passages; keep audio |
| Filing triage and scheduling | Clerks | Case routed late or to wrong track | Override queue, aging alerts |
| Self-help assistants | Unrepresented litigants | Confident wrong procedure or deadline | Retrieval from court-approved content only; links to forms |
| Legal research and drafting | Lawyers, clerks, judges | Invented or misquoted authority | Citation verification gate (below) |
| Risk assessment at bail or sentencing | Judges | Biased or miscalibrated score | Validation, disclosure, contestability, judge decides |
| Machine-generated evidence | Fact finder | Unreliable or fabricated output | Reliability gatekeeping, provenance, authentication |
The rows near the bottom need the heaviest controls. In the EU, AI systems intended to assist a judicial authority in researching and interpreting facts and law are listed as high-risk under Annex III of the AI Act, which brings documentation, logging, human oversight and accuracy obligations; the application dates have shifted, so check the EU AI Act timeline before you plan around them. Public-sector deployments outside courts face similar duties, covered in AI in government.
Risk scores and what Loomis decided
The reference case is State v. Loomis, decided by the Wisconsin Supreme Court in 2016. The sentencing court had considered a COMPAS risk assessment, a proprietary tool whose method the defendant could not inspect. The defendant argued this violated due process. The court upheld the sentence but drew limits that are still the best checklist for any court-facing score: the score may be considered but must not be the determinative factor in the sentence, it must not be used to decide whether someone is incarcerated or how severe the sentence is, and reports given to judges must carry written warnings. Those warnings cover the proprietary nature of the tool, the fact that it predicts group behaviour rather than the individual, questions raised by studies about disproportionate classification of minority offenders, and the need to validate the tool on the local population because it was normed on a national sample.
The same year, a ProPublica analysis of COMPAS scores in one Florida county found that Black defendants who did not reoffend were labelled high risk more often than white defendants who did not reoffend, while the vendor argued the scores were equally calibrated across groups. Both claims can be true at once. When base rates differ between groups, a score cannot in general be calibrated and have equal false positive and false negative rates for every group simultaneously; this is a mathematical result, not a tuning problem. The engineering consequence is that someone must choose which error profile the court can defend, write it down, and monitor it. The metrics are walked through in AI fairness.
Invented citations and why models produce them
In Mata v. Avianca (S.D.N.Y. 2023), lawyers filed a brief citing cases that did not exist; a chatbot had produced them, complete with plausible docket numbers and quotations. The court imposed a $5,000 sanction. Similar incidents have since been reported in many jurisdictions, including filings by self-represented litigants and, occasionally, draft orders. Courts responded with standing orders requiring disclosure or certification of AI use, and judiciaries issued guidance for judges themselves. The judicial guidance for England and Wales, refreshed in October 2025 to replace the April 2025 version, tells judicial office holders that AI output may be inaccurate or invented and must be independently verified, and warns against entering private or confidential information into public AI tools.
The cause is structural. A language model generates the most plausible continuation. Legal citations are an extremely regular format, so the model can produce a well-formed cite whether or not it corresponds to a real decision, and it can attach a quotation that sounds like the court. Retrieval reduces the problem but does not remove it: a model given real sources can still misattribute a quote, cite a case for the opposite of its holding, or cite a case later overruled. Generic causes are covered in LLM hallucination risk; the fix in legal work is a deterministic check against an authoritative source, not a better prompt.
Architecture: a citation verification gate
The gate below treats every citation and every quotation in a draft as a claim to be verified. Steps one to four are mechanical and can be fully automated. Step five, whether the authority supports the proposition, can be assisted by retrieval and a model-based check, but the output is a flag for a human, never a pass on its own.
import re
from dataclasses import dataclass, field
# U.S. reporter citations such as "575 U.S. 320" or "925 F.3d 1291" (simplified).
CITE = re.compile(r"\b(\d{1,4})\s+(U\.S\.|S\. ?Ct\.|F\.(?:2d|3d|4th)?|F\. Supp\.(?: 2d| 3d)?)\s+(\d{1,5})\b")
QUOTE = re.compile(r"“([^”]{20,400})”|\"([^\"]{20,400})\"")
@dataclass
class Finding:
cite: str
status: str # VERIFIED / NOT_FOUND / QUOTE_MISMATCH / NEGATIVE_HISTORY / NEEDS_REVIEW
notes: list = field(default_factory=list)
def verify(draft: str, source) -> list[Finding]:
# source wraps an authoritative database: lookup(), full_text(), history()
findings = []
for m in CITE.finditer(draft):
cite = " ".join(m.groups())
case = source.lookup(cite)
if case is None:
findings.append(Finding(cite, "NOT_FOUND"))
continue
window = draft[m.end(): m.end() + 600] # quotes usually follow the cite
body = normalize(source.full_text(case))
bad = [q for q in quotes(window) if normalize(q) not in body]
if bad:
findings.append(Finding(cite, "QUOTE_MISMATCH", bad))
elif source.history(case) in {"overruled", "reversed", "vacated"}:
findings.append(Finding(cite, "NEGATIVE_HISTORY", [source.history(case)]))
else:
findings.append(Finding(cite, "NEEDS_REVIEW", ["confirm it supports the proposition"]))
return findings
def quotes(s):
return [a or b for a, b in QUOTE.findall(s)]
def normalize(s):
return re.sub(r"\s+", " ", s.replace("’", "'")).strip().lower()
def gate(findings) -> bool:
return not any(f.status in {"NOT_FOUND", "QUOTE_MISMATCH"} for f in findings)Three design points matter more than the regex. First, the resolver must query a source of record, such as an official reporter, a court's own opinion archive or a licensed citator, and never the same model that wrote the draft. Second, block on failure: a report that lists problems but lets the document through will be skimmed. Third, keep the report. A court that asks whether AI was used, and how it was checked, gets a concrete answer with the model version, the sources consulted and the reviewer who signed off.
Worked example: a bench memo
A clerk drafts a two-page bench memo with an assistant. The draft contains three citations; the citations and output below are hypothetical, for illustration. The gate returns:
cite status notes
999 U.S. 101 NEEDS_REVIEW confirm it supports the proposition
998 F.3d 2002 QUOTE_MISMATCH "the standard is necessarily flexible and ..."
997 F.4th 3003 NOT_FOUND
gate: BLOCKED (1 not found, 1 quote mismatch)The first citation exists and its quotation matches; it goes to the reviewer only for the support check. The second case exists, but the quoted sentence appears nowhere in the opinion; the model paraphrased and put quotation marks around the paraphrase, a common pattern. The third does not resolve at all. The clerk removes the third, replaces the second with the exact language from the opinion, and reruns the gate. Total cost: a few minutes, compared with a sanction, a corrected order or an appeal.
Design rules for judicial decision support
Any tool whose output a judge reads before deciding, whether a risk score, a case summary or a recommended outcome, should meet these rules before deployment:
- The judge decides, visibly. The interface shows the score as one input among the factors the law requires, with the Loomis-style warnings attached, and it records the judge's reasons independently of the score. Design against automation bias: do not pre-fill a decision field from the model.
- Parties can see and contest the inputs. The defendant receives the input data and the score early enough to correct errors, such as a misrecorded prior conviction. A tool whose inputs cannot be disclosed is a poor fit for an adversarial process.
- Validate locally and per group. Report calibration, false positive rate and false negative rate for each relevant group on the court's own population, and repeat on a schedule. Check proxies: postcode and arrest history can carry protected attributes even when race is excluded.
- Prefer models the court can explain. A points-based checklist or a logistic regression with published weights is easier to defend than a black box with marginally better accuracy.
- Version and log everything. Each output stores the model version, input snapshot and timestamp, so a decision can be reconstructed on appeal. See human-in-the-loop approval gates for the general pattern.
Machine-generated evidence and deepfakes
Evidence is the newest front. Two problems overlap: outputs from complex systems offered as proof (a model's analysis of trading records, software that reconstructs an accident) and fabricated media (a synthetic audio recording of a party). In the United States, the federal Advisory Committee on Evidence Rules voted in May 2025 to publish a proposed Rule 707 for public comment, and the comment period closed on 16 February 2026. As proposed, machine-generated evidence offered without an expert witness would have to meet reliability standards modelled on Rule 702: sufficient facts or data, reliable principles and methods, and reliable application to the case. The proposal has been contested and its final status was not confirmed at the time of writing, so check it before relying on it.
For deepfakes, authentication still runs through existing rules, but engineering can help the honest party. Hash recordings at capture and keep the hash in an append-only log, sign media at the device where supported, preserve original files and metadata rather than exports, and document the chain of custody. Detector scores are weak evidence on their own: detectors generalize poorly to generators they were not trained on, so treat a score as a reason to investigate, not as proof either way.
Failure modes
- Automation bias. Judges and clerks anchor on a number or a fluent summary. Mitigate with independent reasoning fields and periodic audits comparing decisions with and without the tool.
- Invented or misquoted authority. Caught by the gate above; missed when teams check only that a case exists, not that the quotation and holding match.
- Confidential data in public tools. Sealed filings, juvenile records or witness details pasted into a consumer chatbot. Use an enterprise deployment with no training on inputs and data-loss prevention on the input path.
- Drift. A risk tool validated years ago on a different population and charging practice. Schedule revalidation and alert on shifts in score distribution.
- Vendor opacity. Trade-secret claims block the scrutiny the adversarial process requires. Put disclosure, audit access and validation data rights into procurement contracts.
- Self-help assistants giving wrong deadlines. Ground answers in court-approved content, show the source, and route anything about a deadline to an official page.
Trade-offs
Courts face a backlog, and AI can help clear it: transcription and triage alone save real time. The tension is that efficiency gains come cheapest exactly where errors are hardest to see. Transparency competes with vendor intellectual property; simple, explainable models compete with slightly more accurate opaque ones; disclosure rules for AI use add friction for honest lawyers while doing little against those who conceal it. A workable line is to automate freely where a human reviews every output against the record, gate mechanically where outputs can be checked against a source of record, and keep decisions about liberty with a judge who can explain them without the tool.
What to do next
- Inventory every AI use in your court or firm and place each in the table above by closeness to the decision.
- Put a blocking citation and quotation gate in front of every filing or order drafted with AI assistance.
- Resolve citations only against a source of record, never against the drafting model.
- For any risk score, publish per-group calibration and error rates on the local population and revalidate yearly.
- Attach Loomis-style warnings to score reports and record judicial reasons separately from the score.
- Write disclosure, audit access and validation data rights into every vendor contract.
- Hash and log recordings at capture, and track the status of evidence rules on machine-generated outputs.
- Read the current local standing orders and judicial guidance on AI before deploying anything.