What is Amazon Polly?
Amazon Polly is a text to speech API that converts written text into lifelike audio streams for applications and media. It produces MP3, OGG, and other standard audio file formats. The service draws on deep learning and generative voice engines to synthesize speech that sounds natural and emotionally engaged. It supports 100+ voices across 40+ languages and language variants. Each voice is built from native speakers. You can feed it articles, web pages, PDF documents, or any other text source. The text to speech API integrates directly into existing applications. Output can also be stored, cached, and redistributed.
Features & Benefits
- Neural Text to Speech: generate lifelike speech using a billion-parameter transformer model that produces assertive, emotionally engaged, colloquial voice output.
- Generative Voice Engine: produce synthetic speech with a generative AI engine that renders voice in an incremental, streamable manner.
- Voice Library: choose from 100+ male and female voices built from native speakers across 40+ languages and language variants.
- SSML Support: use Speech Synthesis Markup Language tags to control emphasis, intonation, phrasing, and speaking style at the sentence level.
- Custom Lexicons: define custom pronunciations for acronyms, brand names, internal terminology, or any word the default engine mispronounces.
- Speech Duration Control: adjust speech timing to support multilingual dubbing and synchronize audio to video or animation.
- Audio Output Formats: export speech as MP3, OGG, or other standard formats at sample rates of 8,000 Hz, 16,000 Hz, or 22,050 Hz.
- Caching: store and cache audio files locally for faster retrieval and redistribution.
- Data Privacy: text submissions are not retained after processing (ZDR).
What can Amazon Polly do?
- Convert written articles to audio via API
- Generate voiceovers for animations and games
- Add voice to RSS feeds and websites
- Build interactive voice response systems
- Create multilingual audio for global applications
- Produce dubbed audio in multiple languages
- Customize pronunciation of brand names and acronyms
- Control speech emphasis and intonation with SSML
- Export text to speech audio as MP3 or OGG files
- Add voice output to mobile and IoT applications
Real-World Applications
Content publishers converting long-form articles and blog posts to audio can integrate the text to speech API directly into their publishing workflow. The API processes full web pages or documents and returns a listenable audio file. Publishers can store those files and distribute them alongside written content to reach audiences who prefer listening over reading.
Game studios and animation teams that need voiceover for scripts can call the text to speech API from within their production pipeline. SSML controls let a team adjust tone and pacing on a line-by-line basis. Speech duration adjustment supports multilingual releases where dubbed audio needs to match on-screen timing.
Development teams building voice-enabled applications for global markets can pull from 40+ supported languages to reach users in their native language. The API fits into web apps, mobile apps, and IoT devices. Custom lexicons help maintain consistent pronunciation of product names and technical terms across all output.
Contact center teams building automated voice response systems can use the text to speech to generate caller prompts and responses. Stored audio files can be cached and replayed without regenerating them on every request. The generative voice engine produces speech that sounds natural enough for customer-facing interactions.