Aqua Voice is a voice to text tool that converts spoken words into clean, formatted text in real time across Mac, Windows, and iOS. It allows you to voice type inside any text box in any app without per-app configuration. The tool uses Avalon, a proprietary speech model trained on real developer and productivity workflows. Aqua Voice processes dictation across messaging apps, document editors, coding environments, and AI chat interfaces.
The voice to text tool handles dictation, text refinement, and output formatting in one layer that sits across the operating system. It reads on-screen context to improve recognition accuracy. Users set writing rules per destination and output style. The tool supports 49 languages. A separate Avalon API exposes the same speech model for developers building voice features into their own applications.
Aqua Voice Video
Features & Benefits
- Real-Time Voice Dictation: converts spoken words into formatted text inside any app on Mac, Windows, and iOS without per-app configuration.
- AI Text Refinement: corrects grammar and shapes raw speech into clean output as you talk, without a stop-and-edit step.
- Screen-Aware Transcription: reads the active application and on-screen content to improve accuracy, recognizing code syntax, app-specific terminology, and contextual vocabulary.
- Output Formatting by Destination: applies different formatting to the same spoken input depending on the target app, producing structured emails, casual chat messages, or document-style prose.
- Custom Instructions: accepts user-defined rules for tone, punctuation, filler word removal, regional spelling, and output style, with example-based formatting for consistent results.
- Custom Dictionary: stores user-defined terms including names, brands, and technical keywords for accurate and consistent transcription.
- Text Replacements: maps spoken shorthand or phrases to full text strings, reducing repetition for frequently used content.
- 49-Language Support: processes dictation in 49 languages across accents and natural speaking styles.
- Local Transcript History: stores a searchable, tagged record of past transcriptions on-device without sending data to external servers.
- Multiple Activation Keys: supports more than one keyboard shortcut to trigger dictation.
- Avalon Speech API: exposes the underlying speech model through an OpenAI-compatible endpoint with streaming, batch transcription, speaker labels, and timestamps.
- Organization Privacy Controls: lets admins enforce privacy mode and shared dictionaries across all team members from a central setting.
What can Aqua Voice do?
- voice type in any Mac or Windows app
- dictate AI prompts into ChatGPT or Claude
- transcribe speech to text on iPhone
- voice type in Google Docs or Notion
- compose emails by voice in Gmail or Outlook
- draft Slack messages by voice
- dictate code comments and prompts in VS Code or Cursor
Real-World Applications
Developers dictating into AI coding tools can speak prompts faster than they type them. Avalon recognizes framework names, CLI commands, and model identifiers like GPT versions or kubectl accurately. Spoken prompts reach the AI agent without misheard commands requiring correction, making voice a practical input method for coding sessions where speed and precision both matter.
Long-form writing projects move faster when the author speaks rather than types. Aqua Voice processes natural, conversational speech and shapes it into structured prose using custom instructions. Writers capture ideas at speaking pace and review cleaner output rather than raw dictation, with formatting rules handling punctuation, paragraph breaks, and regional spelling automatically.
High-volume team communication across messaging platforms benefits from voice input when the number of daily updates is large. A spoken message gets reformatted to match the target channel’s style before it sends. A rambling verbal update becomes a readable Slack message or a properly structured email without manual editing.
Product and engineering teams building voice features into their own apps can use the Avalon API to add the same speech model to their stack. The API mirrors the OpenAI Whisper endpoint format, so existing integrations swap over with minimal changes. Applications that handle copilot interfaces, support analytics, or live transcription gain higher accuracy on technical vocabulary without rearchitecting the pipeline.