Fish Audio

Text-to-speech and voice cloning built on the open Fish Speech models

AI Audio & MusicFree planOverseasβ˜…β˜…β˜…β˜…β˜† 4.0

What is Fish Audio?

Fish Audio is a text-to-speech platform built around the Fish Speech models, which the team also releases as open weights. You generate narration from text, clone a voice from a short reference clip, and steer delivery through the character of that reference audio rather than through a wall of sliders. Because the underlying model is open, the same pipeline can run on your own hardware or through the hosted service, and an API makes batch generation practical. It handles several languages in one model and keeps latency low enough for interactive use, which separates it from studio-first voice tools aimed only at produced narration.

Last updated: 2026-09-20. This site only provides an index; for exact features, pricing, and licensing, see the official website.

Key features

  • Text-to-speech generation with several languages in one model
  • Voice cloning from a short reference clip without a training run
  • Delivery and emotion guided by the character of reference audio
  • Open-weight Fish Speech models that can be self-hosted
  • Streaming API for low-latency and interactive speech output
  • Voice library with searchable community-shared voices

Pros & cons

Strengths

  • Cloning needs only a short sample rather than a training dataset
  • Open weights give a self-hosted path with no per-character fees
  • Multilingual output comes from one model, not separate voices

Watch out for

  • Cloned voices require consent and raise clear rights questions
  • The web interface is plainer than polished commercial studios
  • Fine control over pacing is limited without usable reference audio

Best for & use cases

narration, voice cloning experiments, self-hosted tts and audiobook drafts

If you're comparing similar products, check the alternatives below, or browse all tools in the AI Audio & Music category.

FAQ

Can I run Fish Audio on my own hardware?

Yes. The Fish Speech weights are published openly, so the synthesis pipeline can run locally or on your own server. You give up the hosted conveniences - a managed voice library, scaling and billing - in exchange for no per-character cost.

How much audio does cloning need?

A short clip of clean speech is usually enough to produce a usable imitation. Longer, quieter samples improve consistency, while noisy or heavily compressed recordings carry their artefacts into every generated line.

Is it suitable for commercial narration?

It can be, but rights are the constraint. Only clone voices you have permission to use, and check the licence terms for the specific model weights and the hosted plan you rely on.