Amazon Rekognition is a set of pretrained computer vision models that AWS runs for you behind ordinary API calls. You send an image or point at a video in S3, and you get back JSON: labels with confidence scores, bounding boxes, detected text, face geometry, moderation categories, or matches against a face collection you built earlier. You do not train, host or scale a model. You do own everything around the call: what you send, which thresholds you apply, how you handle throttling, what you store, and what decisions you let the output drive.

This article explains the service from first principles: which API family to use for which job, the hard input limits, how labels and face search actually behave, how asynchronous video analysis works, and how to build a production image-moderation pipeline that survives traffic spikes. It ends with failure modes, trade-offs against training your own model, and a checklist. If your input is documents rather than photos, Amazon Textract is the better tool; if it is text, look at Amazon Comprehend.

One availability note matters before you design anything. AWS's Rekognition documentation states that Streaming Video and Bulk Image Analysis are no longer available to new customers from 30 April 2026; accounts that used them in the previous 12 months keep access. A new account should therefore plan around the Image APIs and the stored-video APIs, and process live video by sampling frames into the Image APIs, which is the migration path AWS itself points to.

Three API families, three operational shapes

Rekognition is easiest to reason about as three families with very different operational shapes.

Stateless image analysis. One request, one response, nothing stored by the service. DetectLabels finds objects, scenes and concepts; DetectModerationLabels flags unsafe or suggestive content; DetectText reads short text such as signs and captions (up to 100 words per image); DetectFaces returns face boxes, landmarks, pose and quality; CompareFaces compares a source face against faces in a target image; DetectProtectiveEquipment looks for face, hand and head covers on people. These calls are latency-bound and throttled per second.

Stateful face search. A collection is a server-side index of face feature vectors. IndexFaces detects faces and stores a vector for each one; SearchFacesByImage and SearchFaces look up the nearest stored vectors. Rekognition does not keep the image itself, only the vector plus the IDs you attach. A newer layer lets you group several face vectors under one user with CreateUser and AssociateFaces, then search by user with SearchUsersByImage.

Asynchronous stored video. Operations such as StartLabelDetection, StartFaceDetection, StartContentModeration and StartSegmentDetection take a video in S3, return a JobId at once, publish a completion status to an SNS topic, and leave the results to be paged out with the matching Get* call.

Image (bytes)up to 5 MBImage in S3up to 15 MBVideo in S310 GB, 6 h, H.264Image APIssynchronous, statelessFace collectionstateful vectorsStart* video jobsasynchronousLabels, text, moderationJSON + confidenceFaces, PPE, compareboxes + attributesFace / user matchessimilarity scoresSNS completionthen Get* with JobIdIndexFacesSearch*Your code owns thresholds, retries, storage and the decision the output drives
Figure 1. Inputs, the three API families and what comes back. Only collections hold state; everything else is request-response or job-based.

Input limits you must design around

Most production bugs with Rekognition are input bugs, so learn the fixed limits before you write code. These are set quotas from the developer guide and cannot be raised:

LimitValueWhat it means for you
Image passed as raw bytes5 MB (4 MB for DetectProtectiveEquipment)Phone photos often exceed this; resize or use S3
Image referenced in S315 MBPrefer S3 for large images; bucket must be in the same Region
FormatsPNG and JPEG onlyConvert HEIC, WebP and GIF before calling
Minimum dimensions80 px each side (64 px for PPE)Thumbnails fail; send the original
Maximum dimensions10K px for DetectLabels and DetectModerationLabels; 4096 px for PPEVery large panoramas need downscaling
Smallest detectable face40x40 px in a 1920x1080 image, scaling up for larger imagesCrowd shots lose background faces
Faces per collection20 million face vectors; 10 million user vectors by defaultShard collections by tenant long before this
Stored video10 GB, 6 hours, H.264 in MPEG-4 or MOV, 20 concurrent jobs per accountQueue job starts; do not fire them blindly
Pagination token TTL24 hoursCollect Get* results promptly after completion

Transactions per second are a separate, adjustable default quota per API and per Region. The guide's own sizing rule is simple: peak concurrent calls divided by the window you can spread them over, with five seconds as the minimum window. If 1,000 users try to authenticate in the first ten seconds of your busiest hour, you need 100 TPS of CompareFaces in that Region. Smooth spiky traffic with a queue, then size the quota for the smoothed rate.

Labels: hierarchy, filters and thresholds

Labels form a hierarchy. A detected car comes back as Car with parents Vehicle and Transportation, and each ancestor also appears in the list as its own label. Common objects carry Instances with bounding boxes; scenes and concepts do not. Each label also carries Aliases and Categories. Two request options keep responses small and decisions clean: Features (GENERAL_LABELS, IMAGE_PROPERTIES, or both) and Settings with inclusion and exclusion filters by label or by category.

import boto3
from botocore.config import Config

rek = boto3.client("rekognition", config=Config(
    retries={"max_attempts": 8, "mode": "adaptive"}))   # backoff + client-side rate limiting

def listing_labels(bucket, key):
    resp = rek.detect_labels(
        Image={"S3Object": {"Bucket": bucket, "Name": key}},
        Features=["GENERAL_LABELS", "IMAGE_PROPERTIES"],
        MinConfidence=70,                    # API default is 55
        MaxLabels=25,
        Settings={
            "GeneralLabels": {"LabelInclusionFilters": ["Car", "Bicycle", "Motorcycle", "Person"]},
            "ImageProperties": {"MaxDominantColors": 3},
        },
    )
    labels = {l["Name"]: l["Confidence"] for l in resp["Labels"]}
    sharp = resp["ImageProperties"]["Quality"]["Sharpness"]     # 0-100
    return labels, sharp, resp["LabelModelVersion"]

Three habits make label output safe to act on. First, a confidence score is not a probability comparable across labels or model versions; calibrate a threshold per label on a few hundred of your own images. Second, log LabelModelVersion with every stored result, so you can explain a shift in your metrics. Third, use the quality scores: a blurry, dark photo yields low-confidence labels, and the right response is asking for a better photo.

Face collections and search

Face search is a nearest-neighbour lookup over vectors, and its quality is decided at indexing time. Index one good, frontal, well-lit photo per person, attach your own identifier with ExternalImageId (letters, digits and _.-: only), and keep the returned FaceId in your database. IndexFaces indexes at most the 100 largest faces in an image and reports the rest in UnindexedFaces with reasons such as too small, too blurry, too dark or extreme pose.

There is a default mismatch worth knowing. IndexFaces defaults to QualityFilter=AUTO, so poor faces are quietly not indexed. SearchFacesByImage and SearchUsersByImage default to NONE, so a poor face is still searched. Search also uses only the largest face in the input image, and the match threshold defaults to 80. Set all three explicitly:

def enroll(collection, person_id, bucket, key):
    r = rek.index_faces(CollectionId=collection,
                        Image={"S3Object": {"Bucket": bucket, "Name": key}},
                        ExternalImageId=person_id, MaxFaces=1, QualityFilter="HIGH")
    if not r["FaceRecords"]:
        reasons = [u["Reasons"] for u in r["UnindexedFaces"]]
        raise ValueError(f"enrolment photo rejected: {reasons}")
    return r["FaceRecords"][0]["Face"]["FaceId"], r["FaceModelVersion"]

def identify(collection, image_bytes, threshold=95.0):
    r = rek.search_faces_by_image(CollectionId=collection, Image={"Bytes": image_bytes},
                                  FaceMatchThreshold=threshold, MaxFaces=3, QualityFilter="MEDIUM")
    return [(m["Face"]["ExternalImageId"], m["Similarity"]) for m in r["FaceMatches"]]

Treat a face match as evidence, not a verdict. A similarity score is a measure of how close two vectors are, and the error rate depends on lighting, pose, image quality and the population you serve. For anything consequential, require a high threshold, keep a human in the loop, and record which threshold and face model version produced the match. Face data is biometric data in many jurisdictions; get informed consent, keep a deletion path (DeleteFaces plus your own records), and do not build identification features you could not defend to the people being identified.

Stored video as asynchronous jobs

Stored-video analysis is a job, so build it like one. The start call takes a ClientRequestToken that makes retries idempotent (the same token returns the same JobId), a JobTag that comes back in the completion notification, and a NotificationChannel naming an SNS topic plus an IAM role Rekognition can use to publish. Starting more than 20 concurrent jobs gets LimitExceededException, so put job starts behind a queue with a concurrency limit rather than starting one per upload.

def start(bucket, key, upload_id):
    return rek.start_label_detection(
        Video={"S3Object": {"Bucket": bucket, "Name": key}},
        ClientRequestToken=upload_id[:64],          # retry-safe
        JobTag=upload_id,
        MinConfidence=60,                           # video default is 50
        NotificationChannel={"SNSTopicArn": TOPIC_ARN, "RoleArn": PUBLISH_ROLE_ARN},
    )["JobId"]

def on_completion(job_id):                          # called from the SNS -> SQS consumer
    token, labels = None, []
    while True:
        kw = {"JobId": job_id, "SortBy": "TIMESTAMP"}
        if token:
            kw["NextToken"] = token
        r = rek.get_label_detection(**kw)
        if r["JobStatus"] != "SUCCEEDED":
            raise RuntimeError(r.get("StatusMessage", r["JobStatus"]))
        labels.extend(r["Labels"])                  # each has Timestamp in ms
        token = r.get("NextToken")
        if not token:
            return labels

Subscribe an SQS queue to the SNS topic rather than a function directly: the queue buffers bursts of completions, gives you a dead-letter queue for jobs whose results cannot be fetched, and lets you replay. SNS fan-out and SQS redrive are covered in their own articles. Fetch results within the 24-hour pagination window and copy them to your own store; the job is not your archive.

Worked example: a marketplace photo pipeline

Consider a second-hand marketplace that receives 600,000 listing photos a day, with a peak hour carrying 12 percent of the day: 72,000 photos, or 20 per second on average and bursts of perhaps 60 per second when a promotion goes live. Each photo needs a moderation check and a category check (does a listing in Bicycles actually show a bicycle?).

  1. The client uploads to S3 with a presigned URL; the object-created event goes to an SQS queue, never straight to a function, so bursts queue instead of throttling.
  2. A consumer with a fixed concurrency of 20 pulls messages and calls DetectModerationLabels first. If any label in the blocked set exceeds 90, the listing is rejected and no further call is made; between 60 and 90 it goes to a human review queue.
  3. Photos that pass call DetectLabels with an inclusion filter for the listing's category, and the result plus both model versions is written to a table keyed by photo ID. A conditional write on that key makes redelivered messages harmless.
  4. Throttling exceptions return the message to the queue with a longer visibility timeout; after five receives it lands in a dead-letter queue that pages nobody but is reviewed daily.

Sizing: two calls per photo at a smoothed 20 photos per second is 20 TPS per API, plus headroom for retries, so request a quota of roughly 30 TPS for each of the two APIs in the Region, cap consumers at that rate, and let the queue absorb the 60-per-second bursts. If a burst lasts ten minutes, the backlog grows by (60 minus 30) times 600, or 18,000 photos. When arrivals fall back to 20 per second, the consumers drain it at a net 10 per second, so it clears in about 30 minutes. That delay is acceptable for listings and costs nothing extra; a synchronous design would have needed twice the quota for the same work.

Failure modes and operations

SymptomCauseFix
ProvisionedThroughputExceededException or ThrottlingExceptionAbove quota, or a sudden spikeQueue, smooth, retry with exponential backoff and jitter; then size the quota
InvalidS3ObjectExceptionWrong Region, missing object permission, KMS key the caller cannot useSame-Region bucket; grant read on the object and its key
ImageTooLargeExceptionBytes over 5 MB, or PPE over its limitsResize client-side or switch to an S3 reference
InvalidImageFormatExceptionHEIC, WebP, GIF or a corrupt fileConvert to JPEG at upload time
InvalidParameterException on searchNo face found in the input imageRun DetectFaces first; show the user a retake prompt
Boxes in the wrong placeEXIF orientation applied by your viewer but not your drawing codeNormalize orientation before calling, then draw on the same pixels
Silent metric driftModel version changed underneath youStore LabelModelVersion and FaceModelVersion; alert on change

Collections deserve one more warning: a collection is tied to the face model version it was created with. Moving to a newer model means creating a new collection and re-indexing from your stored enrolment photos, which is only possible if you kept them. Keep the source images (encrypted, access-logged) for as long as your consent allows.

Trade-offs: managed vision versus your own model

Rekognition wins when your categories match its taxonomy, when you value zero model operations, and when request volumes are moderate. It loses when you need classes it does not know (a specific part number, a defect type), when you need to explain or tune the model, or when volume is so high and steady that a self-hosted model on reserved GPUs costs less. A practical middle path: use Rekognition as the first filter and route the uncertain band, say 50 to 80, to a human or a specialist model. Multimodal foundation models return prose rather than per-class scores and cost more per image; keep them for the long tail.

If you orchestrate many steps (moderation, labels, a human review, a notification), Step Functions gives you retries, timeouts and a visible execution history for each photo without writing that plumbing yourself.

What to do next

  1. List the decisions your product will make from vision output, and the cost of a false positive and a false negative for each.
  2. Collect 300 to 500 of your own images per decision and measure precision and recall at several thresholds before choosing one.
  3. Convert uploads to JPEG or PNG and enforce the 80 px minimum and the 5 MB or 15 MB limits at the edge.
  4. Put every call behind a queue with fixed consumer concurrency, adaptive retries and a dead-letter queue.
  5. Store results with LabelModelVersion or FaceModelVersion, threshold used and request time.
  6. For face features, write the consent, retention and deletion policy first, set QualityFilter and FaceMatchThreshold explicitly, and keep enrolment images so you can re-index.
  7. For video, gate job starts to fewer than 20 concurrent, use ClientRequestToken, and fetch results within 24 hours.
  8. Do not design new work around Streaming Video or Bulk Image Analysis on a new account; sample frames into the Image APIs instead.
Key takeaway: Rekognition gives you managed vision models behind simple APIs, but the reliability of a vision feature comes from what surrounds the call: input normalization against fixed limits, a queue that smooths traffic to a sized TPS quota, thresholds calibrated on your own images, stored model versions, explicit face-search settings, and a clear policy for biometric data.