VoiceStudio v0.5.6

VoiceStudio is a useful text to speech application for users who want to turn written text into natural sounding speech. It is designed for people who need voice generation without having to record every sentence manually, making it suitable for content creation, accessibility, presentations, narration, and other voice based projects.

One of the main advantages of VoiceStudio is its focus on making voice generation accessible. Instead of requiring professional recording equipment, users can generate spoken audio from text and adjust the result according to their needs. This can save considerable time when producing longer narration or repeated voice content.

VoiceStudio can also be useful for experimenting with different voices and speaking styles. Depending on the supported engine and configuration, users may have control over characteristics such as speech rate, pitch, and other voice parameters. This gives more flexibility than a basic text to speech reader.

Another appealing aspect is its potential for voice cloning and AI based voice generation. These capabilities can be useful for creators who want a consistent voice across videos, tutorials, podcasts, or other digital content. However, voice cloning should always be used responsibly and with appropriate permission from the person whose voice is being replicated.

The application can also fit well into accessibility workflows. Converting text into speech can make written information easier to consume for users who prefer listening rather than reading. It can also be useful for proofreading, allowing writers to hear how their text sounds when spoken aloud.

VoiceStudio is not necessarily aimed at replacing professional voice actors in every situation. AI generated speech can still have limitations with pronunciation, emotion, pacing, and context. Results can vary considerably depending on the voice engine and the quality of the input text.

The learning curve will also depend on how many features you want to use. Basic text to speech is relatively straightforward, while more advanced voice customization and cloning features may require additional experimentation.

Overall, VoiceStudio is an interesting option for users who want convenient access to modern speech generation tools. It combines the practicality of text to speech with more advanced voice capabilities, making it useful for both casual users and content creators.

Download VoiceStudio v0.5.6 - Software Mirrors

VoiceStudio v0.5.6 for Windows

VoiceStudio-Electron-0.5.6-win-x64.exe | 204.51 MB

VoiceStudio v0.5.6 for macOS

VoiceStudio-Electron-0.5.6-mac-x64.zip | 262.5 MB

VoiceStudio-Electron-0.5.6-mac-x64.dmg | 262.28 MB

VoiceStudio-Electron-0.5.6-mac-arm64.zip | 254.79 MB

VoiceStudio-Electron-0.5.6-mac-arm64.dmg | 254.59 MB

VoiceStudio v0.5.6 for Linux

VoiceStudio-Electron-0.5.6-linux-x64.deb | 195.01 MB

VoiceStudio-Electron-0.5.6-linux-x64.AppImage | 261.09 MB

VoiceStudio v0.5.6 Source Code

VoiceStudio v0.5.6 Source code (zip)

VoiceStudio v0.5.6 Source code (tar.gz)

VoiceStudio v0.5.6 Release Notes:

VoiceStudio now talks to your other tools. Answer Twilio phone calls in a saved voice, and connect Claude Code, Cursor, Codex CLI and the OpenAI Agents SDK to VoiceStudio with copyable setup that works, including in Docker, where MCP previously returned HTTP 405. Integration cards now say plainly which ones work with VoiceStudio and which are external links. This release also lets you replace a saved clone's reference sample, clones long reference clips on OmniVoice, finishes long audiobook chapters on 8 GB GPUs, and keeps rendered chapters through a power-off or data-folder move. Download
  • macOS Apple Silicon: DMG
  • macOS Intel: DMG
Upgrading: Install over your existing Electron app; voices, projects and settings are kept. If setup asks, choose Install local runtime to refresh its dependencies. MCP clients configured with /mcp keep working; new exports use /mcp/. OpenAI-compatible /v1 clients now get OpenAI-format errors with HTTP 400 instead of 422, and response_format="mp3" returns real MP3. Moving from Tauri: Close the app, back up its data directory, install Electron, and verify your voices and projects before removing Tauri. Follow the migration guide. Tauri v0.5.3 remains the final Tauri release; its updater cannot install Electron. Installer trust: Electron installers are unsigned or ad-hoc signed and are not Apple-notarized. Windows/macOS may show trust warnings. macOS automatic updates are unverified; use manual installer updates. See the installation guide. Highlights
  • Answer Twilio phone calls with a greeting in a saved voice; off by default, with a local phone-quality test (#2291)
  • Claude Code, Cursor, Codex CLI and the OpenAI Agents SDK connect to VoiceStudio with working setup (#2289, #2290)
  • Replace a saved clone's reference sample, and clone from references longer than 20 s on OmniVoice (#2282, #2281) — thanks @Cengokill!
  • Long audiobook chapters finish on 8 GB GPUs and survive a power-off or data-folder move (#2287, #2279) — thanks @Tran-Van-Hieu and @castlecreati!
  • Malayalam numbers, chapter cue sheets, and captions without spoken markup tags (#2280, #2278, #2295) — thanks @nikhilkilivayil, @shivsin25 and @kevin9327!

Changed

  • Voice Clone says how much of a long reference the active engine uses instead of asking you to trim to 15 s (#2281) — thanks @Cengokill!

Added

  • Edit a saved voice clone's reference sample: play it, replace it by upload or recording, and save back to the same voice (#2282) — thanks @Cengokill!
  • Answer Twilio phone calls in a saved voice: off by default, signed webhooks and single-use stream tokens, with a local phone-quality test (#2291)
  • Malayalam dubs and audiobooks speak numbers, percentages, and decimals in Malayalam instead of reading raw digits (#2280) — thanks @nikhilkilivayil!
  • Download a chapter cue sheet (HH:MM:SSTitle) after a Stories or Audiobook render; timestamps match the M4B's embedded chapters and give MP3 portable chapters (#2278) — thanks @shivsin25!
  • The OpenAI Agents integration page gives a copyable snippet that runs the Agents SDK voice pipeline on VoiceStudio (#2290)
  • The OpenAI-compatible API adds GET /v1/models and POST /v1/audio/translations (multilingual Whisper models; turbo models are refused) (#2290)
  • Speech requests accept stream_format (audio or sse) and OpenAI's {"id": ...} voice object (#2290)
  • Transcriptions pass language, prompt and temperature to Whisper engines and return words with timestamp_granularities[]=word (#2290)
  • Integrations: copyable setup for Codex CLI (config.toml), any MCP client (HTTP or stdio), the VoiceStudio API (curl and OpenAI SDK), and Docker/GHCR (#2289)

Fixed

  • Pasted and imported captions no longer speak karaoke tags, <i>/<font> tags or {\an8} alignment prefixes, while SubRip dialogue such as 2 < 3 is kept; unchanged WebVTT and SubRip exports keep the original cue markup (#2295) — thanks @kevin9327!
  • Saved voices and uploads longer than 20 s clone on OmniVoice again; the best passage is picked automatically (#2281) — thanks @Cengokill!
  • VoxCPM2 no longer pairs a capped reference with a transcript of the whole clip (#2281)
  • Speaker diarisation uses Lightning 2.6.6, which blocks code execution from a crafted checkpoint (#2296)
  • Long audiobook chapters on 8 GB GPUs no longer time out while still rendering, and a timed-out chapter stops using the GPU after its current chunk (#2287) — thanks @Tran-Van-Hieu!
  • On Windows, one reset connection no longer leaves the backend running but unreachable (#2276) — thanks @ialexbond!
  • The official openai SDK and OpenAI Agents SDK work unchanged: gpt-4o-mini-tts and current OpenAI voice names no longer fail with "Unknown model" (#2290)
  • OpenAI instructions now reaches the voice engine instead of being ignored; OmniVoice uses the voice-design tags in it (#2290)
  • Speech pcm is 24 kHz as OpenAI specifies; aac and opus return real AAC and Opus, and a missing encoder is a clear error instead of a mislabelled WAV (#2290)
  • Transcriptions report the detected language instead of always saying English, and verbose_json segments follow OpenAI's schema (#2290)
  • OpenAI-compatible routes return errors in OpenAI's format, with 400 for invalid requests (#2290)
  • Rendered audiobook chapters stay reusable after the data folder moves, and a power-off no longer empties the resume point or tears chapter audio (#2279) — thanks @castlecreati!
  • An audiobook chapter-cache miss now logs which input changed since the cached render (#2279)
  • MCP clients pointed at /mcp connect on Docker and source builds instead of getting HTTP 405 (#2289)
  • MCP tools and backend.speech_client follow a backend moved with OMNIVOICE_PORT instead of assuming port 3900 (#2289)
  • Integration cards show real capabilities: "Works with VoiceStudio" only where a setup exists, "External link" everywhere else (#2289)

Docs

  • Per-engine reference clip limits in the engine guide (#2281)

Contributors

  • @nikhilkilivayil — Malayalam number verbalization (#2280).
  • @shivsin25 and @Shivendra-Coherent — chapter cue sheet export (#2278).
  • @kevin9327 — caption markup stripping for dubbing (#2295).
  • @debpalash — integrations (Twilio, MCP, OpenAI SDK), voice cloning, audiobook reliability, Windows backend, security updates and release validation.

Bug reporters

  • @ialexbond — Windows backend unreachable after a reset connection (#2276).
  • @Tran-Van-Hieu — audiobook chapters timing out on an 8 GB GPU (#2287).
  • @castlecreati — audiobook chapter cache lost after a power-off (#2279).
  • @Cengokill — no way to replace a saved clone's reference sample (#2282).
  • @Cengokill — long reference clips failed to clone on OmniVoice (#2281).

Install

Download packages from the latest release. First launch creates a managed Python environment and downloads the default model. Later launches reuse both.

Note

On macOS, first launch needs a one-time right-click → Open approval. Intel Macs cannot run the local Python backend; use a remote backend instead.

First voice

  1. Launch VoiceStudio and open Voice Cloning.

  2. Add a clean voice sample. Three seconds works; 5–15 seconds usually gives a better prompt.

  3. Enter text, choose a language, then select Generate.

Run from source

Install the development prerequisites, then:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop

Use bun run dev for the browser UI. See Contributing for services, tests, and platform packages.

If setup fails

Pros

  • Convenient text to speech generation

  • Useful for narration and content creation

  • Can reduce the need for manual voice recording

  • Voice customization options

  • Useful for accessibility

  • Potentially useful for consistent voice production

Cons

  • Voice quality can vary depending on the selected engine

  • AI voices may still sound less natural in some situations

  • Advanced features can require experimentation

  • Voice cloning requires responsible use and proper authorization

Final Verdict

VoiceStudio is a capable voice generation tool for anyone who needs to convert text into speech quickly and consistently. Its usefulness extends from simple accessibility applications to more creative projects such as video narration and digital content.

It is not a complete replacement for professional voice recording, but its convenience makes it a practical addition to an AI focused content creation workflow.

VoiceStudio v0.5.6
Free
Software Informations:
Developer:

Operating System:
Windows / macOS / Linux
Date Added:
2026-09-23T14:03:39.768Z
Categories:

Post a Comment/Report Broken Link: