AudioLDM
Original latent diffusion model for text-to-audio generation.
About
AudioLDM is a text-to-audio latent diffusion model from the University of Surrey's CVSSP group that generates sound effects, environmental audio, and short music from text descriptions. It works in a latent space learned from audio and is the foundation that AudioLDM 2 later built on. Pretrained checkpoints, a Hugging Face demo, and Colab notebooks are provided, and generation quality varies with the random seed.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Music & Audio Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- Minimum VRAM
- 8 GB
- Added
- Apr 3, 2026
Related Tools
Audio generation framework by Meta including MusicGen for text-to-music.
Latent diffusion model for text-to-audio, music, and speech generation.
Audio super-resolution model for upsampling audio to higher sample rates.
State-of-the-art music source separation model by Meta for splitting tracks.
Fast music generation model producing full songs with lyrics in seconds.
PyTorch library for deep learning research on audio generation including MusicGen and AudioGen.