What is Vois?
Vois is text to speech software for Windows and Mac that runs 100% locally on your local machine. It combines a TTS voice engine with a full audio production suite: script editor, multi-track timeline, mastering pipeline, and platform export presets. Content creators can use it to produce podcasts, audiobooks, YouTube voiceovers, training courses, and documentary narration from written scripts.
Features & Benefits
- 63+ expressive voices – Choose from voices across 15 character categories, including narrators, hosts, and character voices.
- Three TTS engines – Fast engine for quick English drafts; Expressive engine for production-quality English with emotion and intonation; Multilingual engine for 23-language delivery.
- Voice cloning – Upload a 5–60 second audio sample to generate a custom voice model stored locally on your device.
- Multi-speaker support – Assign different voices to individual speakers within a single project for dialogues, podcasts, and audiobooks.
- Unlimited generations – Generate, preview, and re-generate audio with AI.
- Script editor – Type scripts directly or import from PDF, EPUB, DOCX, or web articles. Tag speakers and organize content into projects and series.
- Multi-track timeline – Drag and drop audio clips, add crossfades, and adjust timing across multiple tracks.
- Professional mastering pipeline – Apply LUFS normalization, de-esser, EQ, and limiter to finished audio.
- Platform export presets – Export audio formatted for Spotify, YouTube, Apple Podcasts, ACX/Audible, Google Play Books, Kobo, and Findaway.
- 23-language support – Switch languages mid-project without switching voices.
- 100% local processing – All generation happens on-device. Scripts and voice models are never uploaded to external servers.
- CLI & AI automation – Control Vois from the command line or integrate it with AI agents like Claude or ChatGPT to orchestrate voice workflows.
- Pronunciation dictionary – Define custom pronunciation rules for names, brands, acronyms, and technical terms.
- Project management – Organize content into projects, seasons, and chapters.
- System compatibility – Runs on macOS 12+ (Apple Silicon or Intel) and Windows 10/11 (64-bit). GPU acceleration optional, speeds up generation 2–6x on supported hardware.
Real-World Applications
Podcasters working solo can use Vois to produce multi-voice shows without hiring talent. By assigning different voices to guest roles and hosts, then arranging clips on the multi-track timeline, a single creator can build an episode that sounds fully produced. This text to speech software removes the need for separate recording, editing, and mastering tools.
Authors converting manuscripts to audiobooks may find Vois particularly useful for managing long-form projects. A 50,000-word book can be imported, narrated with a library or cloned voice, mastered to ACX standards, and exported — all inside one application. Writers who iterate heavily benefit from the flat-rate model, since re-generating revised paragraphs carries no added cost.
YouTube creators producing explainers, tutorials, or faceless channel content can generate consistent voiceovers at scale. Because Vois runs locally, sensitive scripts — unreleased product tutorials, proprietary training content — never touch a third-party server. The ability to automate generation via CLI also opens up batch workflows for high-volume channels.
Documentary producers and educators working across multiple languages can use the multilingual engine to narrate content in 23 languages while keeping a consistent voice character. A training course built in English can be re-narrated in Spanish or Japanese without switching to a different voice profile, helping teams maintain audio consistency across localized versions.