Amazon Bedrock is AWS's managed service for running foundation models from Amazon and other providers behind one AWS API. You do not provision GPUs or deploy containers. You call an endpoint with a model or inference profile ID, IAM decides whether you may, and you pay per token or for reserved capacity. That simplicity hides a set of decisions that determine whether a production system is reliable and affordable: which API to call, how requests are routed across Regions, how quotas are consumed, which service tier serves a request, and what to do with each way a response can end.
This article works through those decisions from first principles, with code for the Python SDK. Every API field, error and quota rule below was checked against the AWS documentation on 2026-10-01. Prices, per-model multipliers and limits change quickly, so they are left to the AWS pricing and quota pages. Knowledge bases, agents and customisation are out of scope.
The moving parts
Bedrock exposes several endpoints. The control plane, the bedrock client, manages resources: listing models, creating batch jobs, guardrails and provisioned capacity. The runtime, bedrock-runtime, serves inference through four operations: Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. The documentation also describes a separate bedrock-mantle endpoint for some models, with its own quota scheme; check the model's page to see which endpoint serves it.
The modelId can be a base model, an inference profile, a Provisioned Throughput ARN, a custom model deployment, a Marketplace endpoint or a prompt ARN. Most production calls use an inference profile: a routing policy over one model and a set of Regions.
Converse first, InvokeModel when you must
InvokeModel takes each provider's native request body. Converse defines one message format for every messages-capable model: role-tagged messages with content blocks, an optional system list, and a common inferenceConfig with maxTokens, temperature, topP and stopSequences. Model-specific parameters go in additionalModelRequestFields, and requestMetadata tags let you filter invocation logs.
import os
import boto3
from botocore.config import Config
# Adaptive mode backs off client-side when the service throttles.
cfg = Config(retries={"max_attempts": 6, "mode": "adaptive"}, read_timeout=120)
brt = boto3.client("bedrock-runtime", region_name="us-east-1", config=cfg)
MODEL_ID = os.environ["BEDROCK_MODEL_ID"] # an inference profile ID, e.g. one starting "us."
resp = brt.converse(
modelId=MODEL_ID,
system=[{"text": "You summarise support tickets in two sentences."}],
messages=[{"role": "user", "content": [{"text": ticket_text}]}],
inferenceConfig={"maxTokens": 300, "temperature": 0.2},
requestMetadata={"team": "support", "feature": "ticket-summary"},
)
stop = resp["stopReason"]
usage = resp["usage"] # inputTokens, outputTokens, totalTokens (+ cache fields if used)
if stop == "end_turn":
summary = resp["output"]["message"]["content"][0]["text"]
elif stop == "max_tokens":
raise TruncatedOutput(usage) # do not ship half an answer
elif stop in ("guardrail_intervened", "content_filtered"):
summary = None # handled as a policy outcome, not an error
else:
raise UnexpectedStop(stop)The response carries the model's message in output, token counts in usage, the latency in metrics.latencyMs, and a stopReason. The documented values are end_turn, tool_use, max_tokens, stop_sequence, guardrail_intervened, content_filtered, malformed_model_output, malformed_tool_use and model_context_window_exceeded. Handle every one explicitly. The commonest production bug in LLM code is treating any HTTP 200 as success and shipping a truncated or filtered answer.
Tool use: the loop you own
With a toolConfig, the model can answer with a toolUse block instead of text, and stopReason becomes tool_use. Bedrock does not run your tools. Your code executes the call, appends a toolResult with the matching toolUseId, and calls Converse again. The loop is yours, and so are its limits:
tools = {"tools": [{"toolSpec": {
"name": "get_order_status",
"description": "Look up the status of an order by its ID.",
"inputSchema": {"json": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"]}}}}]}
messages = [{"role": "user", "content": [{"text": "Where is order 8812?"}]}]
for _ in range(5): # hard cap on tool round trips
resp = brt.converse(modelId=MODEL_ID, messages=messages, toolConfig=tools,
inferenceConfig={"maxTokens": 500})
msg = resp["output"]["message"]
messages.append(msg) # keep the assistant turn verbatim
if resp["stopReason"] != "tool_use":
break
results = []
for block in msg["content"]:
if "toolUse" in block:
call = block["toolUse"]
try:
out = {"json": dispatch(call["name"], call["input"])} # your code, validated
status = "success"
except ToolError as e:
out, status = {"text": str(e)}, "error"
results.append({"toolResult": {"toolUseId": call["toolUseId"],
"content": [out], "status": status}}) # status: some models only
messages.append({"role": "user", "content": results})Three rules keep this loop safe. Cap the number of round trips, because a model can call tools repeatedly. Validate tool inputs against the schema in your own code before acting on them, since the schema guides the model but does not guarantee its output. And return tool failures as results with an error status rather than raising, so the model can recover or explain.
Inference profiles and cross-Region inference
Capacity for a popular model in one Region is finite. Cross-Region inference lets Bedrock route a request to another Region with spare capacity, through an inference profile. There are two kinds. A geographic profile, such as one for the US or the EU, keeps processing inside that geography. A global profile may route to any supported commercial Region.
| Geographic profile | Global profile | |
|---|---|---|
| Where requests run | Regions within one geography | Any supported commercial Region |
| Pricing (per AWS docs) | Standard pricing | Approximately 10% savings |
| SCP needs | Allow every destination Region in the profile | Allow aws:RequestedRegion of unspecified |
| Choose when | Data residency rules apply | Cost and capacity matter more than location |
Price is set by the Region you call from, with no routing charge. Requests can reach Regions not enabled in your account, and data stays on the AWS network, encrypted between Regions. CloudTrail logs the call in the source Region, with additionalEventData.inferenceRegion showing where it ran, the evidence auditors will ask for. Inference profiles do not support Provisioned Throughput, which is called by its own ARN. A Region-restricting SCP blocks a geographic profile until every Region in it is allowed; see IAM condition keys.
Quotas: how tokens burn down
On-demand inference is limited per model by tokens per minute and tokens per day, as well as by request rates. The accounting is not simply input plus output. At the start of a request, Bedrock reserves the input tokens plus your max_tokens, and throttles the request if that would exceed the quota. At the end it settles the real charge: input tokens, plus tokens written to the prompt cache, plus output tokens multiplied by a burndown rate, and returns the unused reservation. Cache reads are not counted toward the quota.
The burndown rate depends on the model. It is one for many models and higher for some, notably Anthropic models; the current per-model values are on the quota documentation page and have changed as new models arrived. The documentation's own example uses a rate of five: 1,000 input tokens and 100 output tokens consume 1,500 tokens of quota, while you are billed for 1,100. So a maxTokens far above real output cuts concurrency, because every in-flight request holds the full reservation. Size maxTokens from the observed output-token distribution in CloudWatch, with headroom, not from the model's maximum.
Errors and retries
| Exception | HTTP | Meaning and response |
|---|---|---|
| ThrottlingException | 429 | Over account quota. Back off with jitter; reduce maxTokens; raise quotas or change tier. |
| ModelNotReadyException | 429 | Model not ready to serve; the SDK retries automatically up to 5 times. |
| ModelTimeoutException | 408 | Processing exceeded the model timeout. Shorten input or output, or stream. |
| ModelErrorException | 424 | The model failed while processing. Retry a bounded number of times, then fail the request. |
| ServiceUnavailableException | 503 | Retry with backoff; alert if sustained. |
| ValidationException | 400 | Bad request shape or parameters. Never retry; fix the code. |
| AccessDeniedException | 403 | IAM, SCP or model access. Never retry. |
| ResourceNotFoundException | 404 | Wrong model or profile ID for this Region. |
Keep retries bounded, and put idempotency in your own system: a retry that runs a tool twice is a bug no retry policy can fix.
Service tiers
The serviceTier field selects scheduling. Reserved is capacity bought for one or three months, targeting 99.5% uptime, with minimums of 100,000 input and 10,000 output tokens per minute and overflow to Standard; it is arranged through your AWS account team. Priority costs more than standard on-demand pricing and is served first. Standard (default, or no field) is the normal path. Flex trades longer processing for a discount, suiting evaluations and background work.
# Interactive path: default tier. Overnight enrichment: flex, if the model supports it.
resp = brt.converse(
modelId=MODEL_ID,
messages=messages,
inferenceConfig={"maxTokens": 400},
serviceTier={"type": "flex"},
)
served = resp.get("serviceTier", {}).get("type") # which tier actually served itPriority, Standard and Flex share your on-demand quota, while Reserved capacity is separate. Not every model supports every tier, so check the model card. The tier that actually served a request is returned in the response and recorded in CloudWatch as ResolvedServiceTier, which is the metric to watch if you depend on Priority or Reserved.
Prompt caching
Prompt caching lets a supported model reuse a long, repeated prefix such as a policy manual. Mark the end of the reusable part with a cachePoint block; some models also accept checkpoints in system or tools.
messages = [{"role": "user", "content": [
{"text": long_policy_manual}, # identical on every request
{"cachePoint": {"type": "default"}}, # everything above is a cache candidate
{"text": user_question}, # varies per request
]}]
resp = brt.converse(modelId=MODEL_ID, messages=messages, inferenceConfig={"maxTokens": 400})
u = resp["usage"]
read = u.get("cacheReadInputTokens", 0)
hit_ratio = read / max(1, u["inputTokens"] + read + u.get("cacheWriteInputTokens", 0))Usage reports cacheReadInputTokens and cacheWriteInputTokens; the CloudWatch metrics are CacheReadInputTokenCount and CacheWriteInputTokenCount. Any change before the checkpoint is a miss. Minimum lengths, lifetimes and cache pricing vary by model, so check the prompt caching page.
Batch inference
For work that does not need an immediate answer, batch inference processes a set of prompts asynchronously from S3. Each JSONL line holds a recordId and a modelInput in either InvokeModel or Converse format, chosen when the job is created:
# input.jsonl, one record per line (Converse-format modelInput shown):
# {"recordId": "T-000001", "modelInput": {"messages": [{"role": "user",
# "content": [{"text": "Summarise: ..."}]}], "inferenceConfig": {"maxTokens": 300}}}
bedrock = boto3.client("bedrock", region_name="us-east-1") # control plane
job = bedrock.create_model_invocation_job(
jobName="ticket-summaries-2026-10-01",
roleArn=BATCH_ROLE_ARN, # can read input, write output
modelId=MODEL_ID,
modelInvocationType="Converse", # default is InvokeModel
inputDataConfig={"s3InputDataConfig": {"s3Uri": "s3://my-bucket/batch/in/"}},
outputDataConfig={"s3OutputDataConfig": {"s3Uri": "s3://my-bucket/batch/out/"}},
)
# Then react to the job's state-change event (EventBridge) instead of polling,
# and join outputs back to inputs by recordId: output order is not guaranteed.The documented constraints shape the design. Batch does not support tool calling or structured output, because each record is processed independently with no back-and-forth. The order of output records is not guaranteed, so always join on recordId. Batch is not supported for provisioned models. Record counts and file sizes are quota-limited, and batch pricing is on the pricing page. Use EventBridge for job state changes rather than a polling loop.
Security and data handling
Converse is authorised by bedrock:InvokeModel, and ConverseStream by the streaming action, bedrock:InvokeModelWithResponseStream. To block a model completely, deny both bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream on it, since denying one leaves the other open. Scope policies to the specific profile and model ARNs a workload needs, as in AWS IAM. Keep traffic off the internet with an interface endpoint, covered in AWS PrivateLink.
The Converse API reference states that Bedrock does not store the text, images or documents you send as content, and uses it only to generate the response. Invocation logs, if enabled, are your copy, so treat them as sensitive. Guardrails attach to a request through guardrailConfig with a guardrail identifier, version and optional trace. When one blocks content, stopReason is guardrail_intervened and the trace explains which policy fired. Guardrails are one layer; the broader hardening checklist is in LLM deployment hardening.
Failure modes and trade-offs
- Throttling at modest traffic. Usually
maxTokensset to the model maximum, multiplied by burndown. Measure output length and lower it. - Truncated answers shipped as complete.
max_tokensormodel_context_window_exceededtreated as success. Branch on every stop reason. - Residency surprise. A global profile used for regulated data. Choose geographic profiles by policy and alert on unexpected inferenceRegion values in CloudTrail.
- Runaway agent loops. Tool use with no round-trip cap or budget. Cap turns and tokens per session.
- Cache that never hits. A timestamp or user ID placed before the checkpoint. Keep the prefix byte-identical.
- Model ID drift. Hard-coded IDs break when models are retired or renamed. Keep the ID in configuration and test a switch before you need it.
The central trade-off is convenience against control: no serving infrastructure and one API across providers, at the price of per-token costs, shared quotas and less control over latency than self-hosting. Make each choice per workload.
What to do next
- Move one InvokeModel integration to Converse and branch explicitly on all nine stop reasons.
- Replace hard-coded model IDs with an inference profile ID from configuration, and decide geographic or global per workload based on data rules.
- Pull the output-token distribution from CloudWatch and set
maxTokensfrom it, then recompute how many concurrent requests your quota supports. - Configure retries with backoff for 429 and 503, none for 400 and 403, and put idempotency in your own layer.
- Add
requestMetadatatags for team and feature so invocation logs can be attributed. - Mark stable prompt prefixes with a cachePoint where the model supports it and track the cache read ratio.
- Move non-urgent jobs to Flex or batch after checking model support and the pricing page.
- Audit IAM to deny both invoke actions for unapproved models and confirm CloudTrail shows inferenceRegion for every call.