Audiobox

Meta research model generating speech, sound effects and ambient scenes

AI Audio & MusicFree planOverseasβ˜…β˜…β˜…β˜†β˜† 3.0

What is Audiobox?

Audiobox is a research audio generation model from Meta that covers speech and sound inside one framework. Given text, an audio prompt, or both, it produces speech, sound effects and ambient scenes, and a short voice sample can condition the output so generated speech follows that voice. Unifying speech and sound generation under one conditioning scheme is the research contribution: the same idea drives a line of dialogue and a thunderclap. Meta has published some checkpoints for research use and demos exist, but it is a research project rather than a supported product, so expect setup work and uneven results on unusual prompts.

Last updated: 2026-09-20. This site only provides an index; for exact features, pricing, and licensing, see the official website.

Key features

  • Text-prompted generation of speech, sound effects and ambience
  • Audio prompting that conditions style and content from a sample
  • Voice conditioning from a short reference clip of speech
  • One framework covering both spoken language and general sound
  • Research model weights published for non-commercial work
  • Demo interfaces for trying prompts without a local setup

Pros & cons

Strengths

  • Unusual range, covering speech and sound effects in one model
  • Audio prompting gives a more direct form of control
  • Released checkpoints support research and fine-tuning

Watch out for

  • A research release, not a supported production product
  • Running it locally needs capable hardware and setup work
  • Weight licences are research oriented, so check before commercial use

Best for & use cases

audio research, sound effect prototyping, speech experiments and academic work

If you're comparing similar products, check the alternatives below, or browse all tools in the AI Audio & Music category.

FAQ

Can I use Audiobox in a product?

Not straightforwardly. It is a research release, and while the code is openly available the model weights carry research-oriented terms. Review the licence on the specific checkpoint before building anything commercial.

What makes it different from a text-to-speech tool?

It generates sound as well as speech, and it accepts audio prompts alongside text. That means you can condition output on a reference clip rather than describing everything in words.

Do I need a GPU?

For anything beyond a demo, yes. Local inference on published checkpoints expects a capable GPU, and CPU-only runs are slow enough to be impractical for iteration.