Amazon Polly turns text into speech and Amazon Transcribe turns speech into text. Each is simple to call once, but a production voice feature also has to deal with input limits, per-engine quotas, synchronous versus asynchronous versus streaming calls, custom vocabularies and pronunciation, and audio formats that silently break accuracy. This article explains both services from the API up, then builds one pipeline that uses them together.

Every limit quoted here comes from the AWS quotas pages as checked on 3 October 2026. Quotas change and many are adjustable, so read the Service Quotas console for your account before you size anything. Prices are deliberately left out.

Advertisement

Three calling styles for each service

Both services offer more than one calling style, and choosing the wrong one causes most early problems.

NeedPollyTranscribe
Short, interactiveSynthesizeSpeech: audio streamed back in the responseStreaming: StartStreamTranscription over HTTP/2 or WebSocket
Long, offlineStartSpeechSynthesisTask: writes to S3, optional SNS noticeStartTranscriptionJob: reads from S3, writes JSON to S3
StatusGetSpeechSynthesisTaskGetTranscriptionJob or EventBridge events

Polly synchronous requests take up to 3,000 billed characters (6,000 total including SSML tags, which are not billed) and the returned audio is capped at 10 minutes. The asynchronous task takes up to 100,000 billed characters (200,000 total). Transcribe batch jobs take files up to 28,800 seconds (eight hours) and 2 GB, and at least 500 ms long. The quotas page also lists a newer StartSpeechSynthesisStream operation for generative voices; check the API reference for its current shape before relying on it.

Polly: engines, speech marks, SSML and lexicons

A Polly request names a voice, an engine and an output format. The engines are standard, neural, long-form and generative; not every voice supports every engine, and the newer engines are available in fewer regions, so call DescribeVoices with an engine filter rather than hard-coding a list. Audio formats include mp3, ogg_vorbis, ogg_opus, mulaw, alaw and pcm (raw 16-bit signed little-endian, mono); json returns speech marks.

Speech marks are the feature people miss. Asking for json output with SpeechMarkTypes such as word, sentence, viseme or ssml returns a line of JSON per event with its time offset in the audio, which is how you highlight words as they are read or drive a lip-synced avatar. It is a second request with the same text and voice, so the timings match the audio.

import boto3

polly = boto3.client("polly", region_name="us-east-1")

ssml = """<speak>
  Welcome to lesson three. <break time="400ms"/>
  We configure <sub alias="Kubernetes">K8s</sub> pod disruption budgets.
</speak>"""

audio = polly.synthesize_speech(
    Text=ssml, TextType="ssml", VoiceId="Joanna", Engine="neural",
    OutputFormat="mp3", LexiconNames=["devops-terms"],
)
with open("lesson3.mp3", "wb") as f:
    f.write(audio["AudioStream"].read())

marks = polly.synthesize_speech(
    Text=ssml, TextType="ssml", VoiceId="Joanna", Engine="neural",
    OutputFormat="json", SpeechMarkTypes=["word", "sentence"],
)
for line in marks["AudioStream"].read().decode().splitlines():
    print(line)   # {"time":0,"type":"sentence","start":...,"end":...,"value":"Welcome to lesson three."}

SSML controls pauses, emphasis, rate and substitutions, but support varies by engine, and some tags are rejected outright: <audio>, <lexicon>, <lookup> and <voice> are not supported, a <break> is capped at 10 seconds, and <prosody> rate cannot go below -80 percent. Test your SSML against the engine you deploy on.

Lexicons fix pronunciation once for every request. They are W3C PLS documents uploaded with PutLexicon: up to 100 per account, 40,000 characters each, five applied per request. A lexicon entry can give a phoneme string or an alias, so kubectl can be read as 'cube control' everywhere without editing every script. Larger lexicons add synthesis latency, so keep them focused per domain.

Advertisement

Transcribe batch jobs

A Transcribe batch job points at an audio file in S3 and writes a JSON transcript, and optionally subtitle files, to an output bucket. The settings that most affect quality are the language, a custom vocabulary or custom language model, and how speakers are separated.

transcribe = boto3.client("transcribe", region_name="us-east-1")

transcribe.start_transcription_job(
    TranscriptionJobName="lesson3-2026-10-03",
    Media={"MediaFileUri": "s3://course-audio/lesson3.mp3"},
    MediaFormat="mp3",
    LanguageCode="en-US",
    OutputBucketName="course-transcripts",
    OutputKey="lesson3/",
    Settings={
        "VocabularyName": "devops-terms",
        "ShowSpeakerLabels": True,
        "MaxSpeakerLabels": 2,
    },
    Subtitles={"Formats": ["vtt", "srt"]},
    ContentRedaction={"RedactionType": "PII", "RedactionOutput": "redacted"},
)

  • Custom vocabulary lists domain words and phrases. The quota is 51,200 bytes per vocabulary, 256 characters per phrase and 100 vocabularies per account by default. Create it once with CreateVocabulary and wait for it to reach a ready state before referencing it.
  • Custom language models are trained on your own text so the model learns which word sequences are likely in your domain. They take hours to train, there are default limits of 10 in total and 3 training at once, and they suit large, stable domains better than a short list of product names.
  • Speaker diarization (ShowSpeakerLabels) separates voices in a single channel. Channel identification transcribes each channel separately, up to two channels by default, which is far more reliable for call recordings where agent and customer are on separate channels. Pick the one that matches how the audio was recorded.
  • Redaction with ContentRedaction replaces detected personal data in the transcript; redacted_and_unredacted writes both versions, which then need different access controls.

The output JSON has a full transcript string, a list of items with start and end times and a confidence score for each word, and speaker or channel segments. Use the per-word confidence to flag low-confidence regions for human review rather than trusting the whole file equally.

Transcribe streaming

Streaming transcription keeps a bidirectional connection open: you send small audio chunks, Transcribe sends back partial results that change as more audio arrives and final results that do not. Turning on partial-result stabilisation makes the early words of a partial result stop changing sooner, at a small cost in accuracy, which matters for live captions where flicker is distracting.

boto3 does not expose the streaming operation. The long-used Python route was the awslabs amazon-transcribe package, but its README now marks it deprecated in favour of a new official aws-sdk-transcribe-streaming client in the aws-sdk-python project. The sketch below follows the old package's documented pattern because it shows the shape clearly; for new code, check the new client's API, which differs.

import asyncio
from amazon_transcribe.client import TranscribeStreamingClient
from amazon_transcribe.handlers import TranscriptResultStreamHandler

class Printer(TranscriptResultStreamHandler):
    async def handle_transcript_event(self, event):
        for result in event.transcript.results:
            print(result.alternatives[0].transcript)   # partial and final results both arrive

async def run(chunks):                      # chunks: async iterator of 16 kHz PCM bytes
    client = TranscribeStreamingClient(region="us-east-1")
    stream = await client.start_stream_transcription(
        language_code="en-US", media_sample_rate_hz=16000, media_encoding="pcm",
    )
    async def send():
        async for chunk in chunks:          # 100 ms of 16 kHz 16-bit mono = 3,200 bytes
            await stream.input_stream.send_audio_event(audio_chunk=chunk)
        await stream.input_stream.end_stream()
    await asyncio.gather(send(), Printer(stream.output_stream).handle_events())

The sample rate you declare must match the audio you send, and the audio must be mono for a single channel. A stream declared as 16 kHz that actually carries 8 kHz telephony audio does not fail; it returns plausible nonsense. Chunks of 50 to 200 ms are a good range: much smaller chunks add per-event overhead, much larger ones add latency to every caption. The default limits are 25 concurrent streams and 25 stream starts per second per account and region, both adjustable.

Worked example: narrated lessons with captions

Consider an online course with 400 lessons whose scripts are written by engineers and full of terms such as kubectl, etcd and PostgreSQL. The team wants narrated audio, captions for accessibility and a searchable transcript, regenerated whenever a script changes.

Lesson scriptSSML in S3Step Functionsorchestrates per lessonPolly StartSpeechSynthesisTaskMP3 + speech marks to S3S3 eventTranscribe StartTranscriptionJobcustom vocab, VTT subtitlesaudio URIRound-trip checkword error vs scripttranscriptReview queuemispronounced termsfailPublishaudio + captions to CDNpassLexicon updatePutLexiconre-runLive captions (streaming)StartStreamTranscription
The worked example pipeline. Polly narrates each lesson asynchronously, Transcribe produces captions and a transcript, and comparing the transcript with the source script catches mispronounced technical terms, which are then fixed with a pronunciation lexicon.

  1. An S3 upload of a lesson's SSML triggers a Step Functions execution. Lessons average 9,000 characters, beyond the 3,000-character synchronous limit, so the workflow calls StartSpeechSynthesisTask with an output bucket and an SNS topic, then waits for completion. A second task with OutputFormat=json writes sentence speech marks for the player's highlighting.
  2. When the MP3 lands, the workflow starts a Transcribe job on it with a custom vocabulary of the course's terms and VTT subtitles enabled. Transcribing the synthetic audio gives captions timed to the actual narration.
  3. A round-trip check compares the transcript with the source script after normalising case and punctuation and computes the word error rate. Clean TTS audio should transcribe almost perfectly, so a lesson above a few percent usually means Polly mispronounced something. The differing words go to a review queue.
  4. A reviewer adds a lexicon entry, for example an alias that reads etcd as 'et-see-dee', uploads it with PutLexicon and re-runs the lesson. The same lexicon now fixes every other lesson that uses the term.
  5. Passing lessons publish audio and captions to the CDN. Live office-hours sessions use streaming transcription for captions instead, with the same custom vocabulary.

Throughput is bounded by quotas, not compute. Long-form and generative asynchronous tasks default to 1 start per second and neural to 10, so the workflow uses a Step Functions Map state with a concurrency limit rather than starting all 400 lessons at once, and retries throttling errors with backoff. See AWS Step Functions for the Map and wait-for-callback patterns used here, and Amazon EventBridge for routing Transcribe job state changes instead of polling.

Quotas and throttling

Polly throttles by both request rate and concurrency, and the limits differ sharply by engine. SynthesizeSpeech defaults are 80 transactions per second for standard voices but 8 per second (burst 10) for neural and long-form and 8 for generative. Concurrency is capped too, and throttled calls return ThrottlingException. A chat assistant that synthesises every reply with a neural voice will hit 8 per second long before compute is a concern.

Three habits keep you inside the limits. Cache synthesised audio keyed on a hash of text, voice, engine, lexicons and format; repeated prompts such as menu options should never be synthesised twice. Put a client-side limiter or a queue in front of Polly, sized below the quota, so bursts wait briefly instead of failing. And split long text at sentence boundaries into requests under 3,000 billed characters if you need synchronous playback to start quickly, playing the first chunk while later ones synthesise.

For Transcribe batch, the default of 250 concurrent jobs is generous, but StartTranscriptionJob is limited to 25 calls per second and GetTranscriptionJob to 30. Polling hundreds of jobs every second exhausts the Get quota; use EventBridge job-state events or poll with backoff. Job records are kept for 90 days, so copy anything you need from the job description into your own store.

Failure modes

These are the failures that reach production.

  • Silent truncation. Synchronous Polly output stops at 10 minutes and anything after is cut off. Long text belongs in an asynchronous task.
  • Wrong sample rate or encoding. Streaming audio that does not match the declared rate, or stereo sent as mono, degrades accuracy without an error. Validate the format at the capture point.
  • Vocabulary not ready. Referencing a vocabulary that is still being processed or has failed makes jobs fail. Gate deployments on its state.
  • Engine and region mismatch. A voice and engine pair that works in one region may not exist in another, which breaks failover. Test the exact pair in every region you deploy to.
  • Unbilled but rejected SSML. Unsupported tags or invalid XML make the whole request fail. Validate SSML in CI with a test synthesis on the target engine.
  • PII in the wrong bucket. Writing unredacted and redacted transcripts to the same prefix defeats the purpose. Use separate prefixes with separate KMS keys and IAM policies.

Trade-offs

Managed speech services remove model hosting, scaling and most tuning, and give you per-word timings, diarization and redaction without building them. In exchange you accept per-character and per-second pricing, account-level quotas, a fixed set of voices and languages, and audio leaving your network. A self-hosted open speech model can be cheaper at very high volume or necessary where data cannot leave a boundary, but you then own GPUs, batching, and accuracy testing. Within AWS, neural voices sound better than standard but have one tenth of the synchronous throughput quota by default; generative voices are the most natural but the most restricted. Pick the engine per use: standard or neural for short, high-volume prompts, long-form or generative for narration. If an LLM is producing the text, see Amazon Bedrock for that half of a voice assistant.

What to do next

  1. Read the Polly and Transcribe quota pages for your region and record the per-engine limits you depend on.
  2. Choose the calling style per feature: synchronous Polly for short prompts, asynchronous tasks for anything long, streaming Transcribe only when you need live text.
  3. Build a pronunciation lexicon and a custom vocabulary from the same term list, and version both with your content.
  4. Add a round-trip check (synthesise, transcribe, compare) to catch mispronunciations automatically.
  5. Cache synthesised audio and put a limiter below Polly's quota in front of every synchronous caller.
  6. Use EventBridge events for Transcribe job completion instead of tight polling loops.
  7. Validate audio format at capture, keep redacted and unredacted output apart, and test every voice-engine pair in every region you use.
Key takeaway: Polly and Transcribe are easy to call and easy to misuse. Match the calling style to the job, respect per-engine quotas with caching and client-side limits, fix pronunciation and recognition with lexicons and custom vocabularies built from one term list, validate audio formats, and use a synthesise-then-transcribe round trip to catch speech errors before users hear them.