Meta research model generating speech, sound effects and ambient scenes
Audiobox is a research audio generation model from Meta that covers speech and sound inside one framework. Given text, an audio prompt, or both, it produces speech, sound effects and ambient scenes, and a short voice sample can condition the output so generated speech follows that voice. Unifying speech and sound generation under one conditioning scheme is the research contribution: the same idea drives a line of dialogue and a thunderclap. Meta has published some checkpoints for research use and demos exist, but it is a research project rather than a supported product, so expect setup work and uneven results on unusual prompts.
Last updated: 2026-09-20. This site only provides an index; for exact features, pricing, and licensing, see the official website.
audio research, sound effect prototyping, speech experiments and academic work
If you're comparing similar products, check the alternatives below, or browse all tools in the AI Audio & Music category.
Not straightforwardly. It is a research release, and while the code is openly available the model weights carry research-oriented terms. Review the licence on the specific checkpoint before building anything commercial.
It generates sound as well as speech, and it accepts audio prompts alongside text. That means you can condition output on a reference clip rather than describing everything in words.
For anything beyond a demo, yes. Local inference on published checkpoints expects a capable GPU, and CPU-only runs are slow enough to be impractical for iteration.
Text-to-speech voices that are hard to distinguish from humans
Professional audio workstation for editing, restoration and podcast mixing
Browser audio cleanup that makes phone recordings sound studio grade
Automatic audio post-production for loudness, levels and noise