AudioLDM 2

Latent diffusion model for text-to-audio, music, and speech generation.

Open SourceSelf HostedOffline CapableGPU Required (8GB+ VRAM)
0.0 (0)

About

AudioLDM 2 generates audio from text using a unified latent diffusion approach that covers text-to-audio, text-to-music, and text-to-speech within one framework built on a shared representation of sound. Developed by Haohe Liu and colleagues at CVSSP, University of Surrey, and published in IEEE/ACM Transactions on Audio, Speech, and Language Processing, it ships seven checkpoints including a general-purpose full model, a 48kHz high-fidelity variant, a music-focused model, and speech models trained on GigaSpeech and LJSpeech. Inference runs on CUDA, CPU, or Apple MPS devices, and a Hugging Face Diffusers integration speeds generation up roughly three times while enabling arbitrary-length output. Users work through a command-line interface, a Gradio web app, or the Diffusers API, tuning guidance scale, sampling steps, and seed. The code and models are open source, and the typical audience is audio ML researchers, sound designers, and developers prototyping generative audio features.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
Minimum VRAM
8 GB
Added
Apr 3, 2026

Related Tools

Featured

Audio generation framework by Meta including MusicGen for text-to-music.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)

Audio super-resolution model for upsampling audio to higher sample rates.

Open SourceSelf HostedOfflineGPU 6GB+
Intermediate
0.0 (0)
Featured

State-of-the-art music source separation model by Meta for splitting tracks.

Open SourceSelf HostedOffline
Easy
0.0 (0)

Fast music generation model producing full songs with lyrics in seconds.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)

Audio diffusion model by Harmonai for generating music samples.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

PyTorch library for deep learning research on audio generation including MusicGen and AudioGen.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)
Browse all Music & Audio Generation tools