
Pricing
Pricing information has not been added yet.
Overview
Voicebox is an open-source, local-first AI voice studio that combines voice cloning, text-to-speech, dictation, and agent voice output in one desktop application. Users can clone a voice from a short reference sample, generate speech in many languages using several TTS engines, dictate into applications with a global hotkey, and let MCP-compatible AI agents speak through a selected voice. The entire workflow can run on the user’s own machine so voice data and recordings do not need to leave local infrastructure.
The platform includes multiple speech engines, preset voices, expressive delivery controls, audio effects, long-form generation, a multi-track stories editor, Whisper-based speech recognition, a REST API, and an MCP server. Voice profiles can also include personas that a bundled local model uses for composing or rewriting responses. Voicebox is especially useful for creators, developers, accessibility workflows, and agent builders who want an open alternative combining the input side of voice AI with high-quality generated speech.
Releases
Release information has not been added yet.
Complete your AI stack
Combine voice generation with scripts, video, images, music, and publishing automation.
Best alternatives
Compare similar tools in the same category.
