Police agencies now buy AI in three main forms. Language models draft incident reports from body camera audio. Face recognition searches a photo database for people who resemble a suspect. Plate readers and patrol forecasting software tell officers where to look and whom to stop. Each tool produces output that can end up in a case file, and so in front of a prosecutor, a defence lawyer and a judge.

That changes what a mistake costs. A recommender's error loses a click. A police AI error can put the wrong person in a cell, or send patrols back to the same streets for years. This article is for engineers and security teams who build, buy or audit these systems. It covers the main pipelines, the failure modes each one has, the laws now attached to them, and the controls that keep model output in its place as a lead or a draft. It does not repeat judicial risk scores, which are covered in AI in courts.

Where AI enters a police workflow

Every police AI pipeline has the same four stages: a sensor, a model step, an output and a human decision. What differs is how much legal weight the output carries after that decision, and how easily a busy person can skip the decision.

Three police AI pipelines, the human gate on each, and where the output ends upBody camera audioSpeech to text + LLMDraft narrativeOfficer edits, certifiesCCTV still (probe)1:N face searchRanked candidatesIndependent corroborationALPR cameraPlate OCR + hotlistHit alertVisual plate and state checkCase filebecomes evidenceProsecutor and defencedisclosure, cross-examinationAppend-only audit logmodel and version, input hash, operator, purpose, case number, every human decisionBlue: sensors. Amber: model output, which is a lead or a draft. Green: the human step that law or policy requires.
Each pipeline ends in a human gate, and everything after the gate is discoverable. The audit log under the pipelines is what lets anyone reconstruct, months later, what the model said and what a person did with it.

Two design rules follow from the picture. First, the output must say what it is. A draft must look like a draft, and a list of candidates must look like a list of candidates and not like an identification. Second, the human gate must be enforced by the software. A policy that says an officer should corroborate a match is weak. A case workflow that cannot move to a warrant request until a corroboration field is filled in is much stronger.

LLM-drafted reports

Report drafting is the fastest-growing use. The product transcribes body camera audio, then a language model writes a first-person narrative in the officer's voice. The officer edits it and signs. The appeal is plain: officers spend a large share of every shift writing reports.

The risks are particular to this setting. The model can add details nobody said, such as a weapon, a direction of travel or a quotation. The transcript can mishear a name or a number. The officer, tired at the end of a shift, tends to accept fluent text, and the draft can replace the officer's own memory, which is what a report is meant to record. In court, a defence lawyer can ask which sentences the officer wrote, and an officer who cannot answer loses credibility on everything.

Legislatures have started to respond. Utah's SB 180, in effect since 7 May 2025, requires agencies to adopt a policy on generative AI. Any report written wholly or partly with it must carry a disclaimer and the author's certification that they reviewed it for accuracy. California's SB 524, signed on 10 October 2025 and in force since 1 January 2026, requires disclosure of AI use and the program used, the officer's review and signature, and retention of the first AI draft and an audit trail. Read the statutes for exact wording before you design a compliance feature around them.

The engineering answer is sentence-level provenance. Every sentence in the draft carries the transcript segments it came from, and a checker flags any sentence whose names, numbers or quotations do not appear in those segments. Flagged sentences must be edited or confirmed before the report can be signed.

import re
from dataclasses import dataclass

@dataclass
class Sentence:
    text: str
    sources: list          # transcript segment ids, e.g. ["seg-0412"]
    officer_edited: bool = False

CHECKABLE = re.compile(r"\b(\d[\d:./-]*|[A-Z][a-z]+)\b|\"[^\"]+\"")

def unsupported(draft, segments):
    # Return sentences whose names, numbers or quotes are absent from their cited audio.
    flagged = []
    for s in draft:
        cited = " ".join(segments[i]["text"] for i in s.sources if i in segments)
        if not s.sources:
            flagged.append((s, "no source segment"))
            continue
        for m in CHECKABLE.finditer(s.text):
            if m.start() == 0:              # a capitalised first word is not a name
                continue
            token = m.group(0).strip('"')
            if token.lower() not in cited.lower() and not s.officer_edited:
                flagged.append((s, f"'{token}' not in cited audio"))
                break
    return flagged

def can_sign(draft, segments, first_draft_stored):
    # Signing is blocked until every flag is resolved and the raw draft is retained.
    return first_draft_stored and not unsupported(draft, segments)

Keep the first draft in write-once storage with a hash, next to the transcript and the model version. Mark officer edits so a reviewer can see which words are machine text. Axon's Draft One can add friction on purpose by inserting obviously absurd sentences that the officer must remove before submitting, and 2025 reporting found that some departments switch that safeguard off. A test of attention that an agency can disable is not a control; sentence provenance enforced at signing is.

Face recognition: a search, not an identification

A one-to-many face search compares a probe image, often a poor CCTV frame, against a gallery of licence or booking photos and returns a ranked list of candidates with scores. It is a search, not an identification. The arithmetic shows why.

Worked example. The gallery holds 10,000,000 photos. At the chosen threshold, the algorithm's false match rate is one in a million per comparison, which is good. One search therefore returns about 10 wrong people above the threshold. Suppose the real suspect is in the gallery and the low-quality probe still finds them 90 percent of the time. Then a candidate above the threshold is the suspect with probability of about 0.9 divided by 10.9, roughly 8 percent. If the suspect is not in the gallery at all, every candidate is wrong. A better algorithm does not remove the base-rate problem, because gallery size multiplies the false match rate.

The next step makes things worse. An analyst chooses the candidate who looks most like the probe, and a witness is then shown a photo lineup that contains that person. The lineup now contains someone selected for looking like the offender, so a confident wrong pick is likely. That is the pattern in the wrongful arrest of Robert Williams in Detroit. The city's settlement in June 2024 paid Williams $300,000. It barred arrests based solely on face recognition results or on a photo lineup generated from such a search, and required a review of cases from 2017 to 2023.

Controls that follow:

  • Label every result as an investigative lead that is not probable cause, in the interface and in exported PDFs.
  • Require a case number, a stated purpose and the operator's identity before a search can run. Log the probe image hash and every edit made to the probe, such as cropping, enhancement or a composite.
  • Return a candidate list rather than a top-1 answer, and record the scores and the gallery size so a reader can do the base-rate arithmetic.
  • Block the warrant or arrest workflow until an independent corroboration field, such as phone records, a fingerprint or a location alibi check, is filled in and reviewed.
  • Test demographic performance on probes of the quality you actually get, not on passport-grade photos.

Predictive patrols and feedback loops

Place-based forecasting predicts where crime will be recorded and sends patrols there. The trap is in the word recorded. Many offences, such as drug possession, are recorded mainly when police are present. More patrols mean more records, and more records mean more predicted crime. Lum and Isaac showed this in 2016 with Oakland drug arrest data: a forecasting algorithm would have concentrated patrols in neighbourhoods that were already heavily policed, while public health survey data suggested drug use was spread much more widely. Ensign and colleagues later described these as runaway feedback loops.

A ten-line simulation shows the effect. Two districts have the same true rate. Patrols follow past records, and records follow patrols.

import random

def simulate(days=365, true_rate=0.3, seed=1):
    random.seed(seed)
    records = [6, 5]                      # one extra historical record in district 0
    for _ in range(days):
        target = 0 if records[0] >= records[1] else 1   # patrol the 'hotter' district
        if random.random() < true_rate:   # crime is only seen where police are
            records[target] += 1
    return records

print(simulate())   # district 0 collects nearly every record; district 1 stays at 5

The fixes are statistical, not cosmetic. Train on victim-reported offences, such as burglary reports, which depend much less on where officers stand. Cap how far the allocation can move from one period to the next. Spend a fixed fraction of patrol time on random allocation so you keep collecting unbiased data. Evaluate against outcomes that patrols cannot create. Forecasts about individuals are a different category: the EU AI Act prohibits assessing a person's risk of offending based solely on profiling or personality traits.

Plate readers and evidence integrity

Automatic licence plate readers (ALPR) are the most widely deployed police AI and the least discussed. The OCR can misread characters, such as 8 and B or 0 and D, or read the right characters on a plate from another state. Hotlists go stale when a stolen car is recovered but the entry stays. A hit then leads to a high-risk stop of an innocent driver. The controls are simple and often skipped: show the camera image with the alert, require the officer to confirm the plate characters and state visually, expire hotlist entries and record their source, and set retention limits on non-hit reads, which build a location history of everyone who drives past.

Across all three pipelines, treat AI output as evidence that may need to be disclosed. Pin model versions, store inputs and outputs with hashes in an append-only log, and keep enough to replay a result later. The methods in AI forensics and audit logging for LLM systems apply directly. Remember that a deepfaked video submitted as evidence is now a realistic threat as well; the deepfakes article covers provenance controls.

The rules that apply

In the European Union, the AI Act's prohibitions have applied since 2 February 2025. Three matter most here: real-time remote biometric identification in publicly accessible spaces for law enforcement, except in narrow listed cases with prior authorisation; individual crime-risk prediction based solely on profiling; and building face databases by untargeted scraping of images from the internet or CCTV. Several listed uses are high-risk systems under Annex III instead, such as remote biometric identification that is not prohibited, evaluating the reliability of evidence and profiling during investigations. As amended by the Digital Omnibus, those obligations apply from 2 December 2027: risk management, logging, human oversight, accuracy testing and registration. See the EU AI Act guide for the full timeline.

In the United States the rules are a patchwork of state statutes, city ordinances, consent decrees and settlements such as Detroit's. Procurement is often the strongest lever a jurisdiction has. AI in government explains how to write security and audit requirements into contracts.

Failure modes

FailureWhat it looks likeControl
Hallucinated report detailA weapon or quote nobody said appears in a signed reportSentence provenance, sign-off blocked on flags
Search treated as identityArrest on a top-1 face match plus a lineupLead labelling, corroboration gate
Probe tamperingAn edited or composite probe returns a confident matchLog probe edits and hash every version
Feedback loopPatrol map never changes, records rise where patrols areVictim-reported targets, random allocation share
Stale hotlistStop of a car that was recovered weeks agoEntry expiry, source field, visual confirmation
Unreplayable outputVendor updated the model; nobody can reproduce the resultVersion pinning, stored inputs and outputs

Trade-offs

Friction against adoption: every gate in this article costs officer time, and tools that are too slow get bypassed, for example with a personal phone app. Put the friction where the legal risk is (signing, warrant requests, stops) and nowhere else. Retention against privacy: keeping drafts, probes and plate reads enables audits and defence access, but builds a surveillance archive. Keep what accountability needs, for a fixed time, and restrict who can query it. Vendor convenience against inspection: closed systems rarely expose scores, versions or logs. Make these export rights a contract condition.

What to do next

  1. List every AI tool in use, including vendor features switched on by default, and draw its sensor, model, output and human gate.
  2. For report drafting, enable draft retention, edit tracking and the required disclaimers; add a provenance checker that blocks signing.
  3. For face search, add purpose and case fields, probe-edit logging, lead labelling and a corroboration gate in the case system.
  4. Run the base-rate arithmetic for your gallery size and threshold, and print it in analyst training.
  5. For patrol forecasting, compare the forecast's targets with victim-reported data and add a random allocation share.
  6. For ALPR, require visual confirmation, expire hotlist entries and set a retention limit for non-hit reads.
  7. Write disclosure procedures so prosecutors can hand AI outputs and logs to the defence.
Key takeaway: Police AI produces drafts, candidate lists and alerts, not facts. Make every output say what it is, enforce the human step in software instead of policy, and keep version-pinned, hashed records that a defence lawyer can examine. Do the base-rate arithmetic before you trust a face match, train forecasts on data that patrols cannot create, and confirm every plate read by eye.