Text-to-speech and voice cloning built on the open Fish Speech models
Fish Audio is a text-to-speech platform built around the Fish Speech models, which the team also releases as open weights. You generate narration from text, clone a voice from a short reference clip, and steer delivery through the character of that reference audio rather than through a wall of sliders. Because the underlying model is open, the same pipeline can run on your own hardware or through the hosted service, and an API makes batch generation practical. It handles several languages in one model and keeps latency low enough for interactive use, which separates it from studio-first voice tools aimed only at produced narration.
Last updated: 2026-09-20. This site only provides an index; for exact features, pricing, and licensing, see the official website.
narration, voice cloning experiments, self-hosted tts and audiobook drafts
If you're comparing similar products, check the alternatives below, or browse all tools in the AI Audio & Music category.
Yes. The Fish Speech weights are published openly, so the synthesis pipeline can run locally or on your own server. You give up the hosted conveniences - a managed voice library, scaling and billing - in exchange for no per-character cost.
A short clip of clean speech is usually enough to produce a usable imitation. Longer, quieter samples improve consistency, while noisy or heavily compressed recordings carry their artefacts into every generated line.
It can be, but rights are the constraint. Only clone voices you have permission to use, and check the licence terms for the specific model weights and the hosted plan you rely on.
Text-to-speech voices that are hard to distinguish from humans
Professional audio workstation for editing, restoration and podcast mixing
Browser audio cleanup that makes phone recordings sound studio grade
Automatic audio post-production for loudness, levels and noise