ACE-Step 1.5
Updated music generation model with improved quality and longer generation.
About
ACE-Step 1.5 is an open foundation model for music generation that combines diffusion with a deep-compression autoencoder and a linear transformer to synthesize up to four minutes of music in roughly twenty seconds on an A100. It uses MERT and m-hubert for semantic representation alignment, supports lyric input, and aims to balance generation speed, musical coherence, and controllability. Released as an open-source research checkpoint.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Music & Audio Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- Minimum VRAM
- 8 GB
- Added
- Apr 3, 2026
Related Tools
Audio generation framework by Meta including MusicGen for text-to-music.
Latent diffusion model for text-to-audio, music, and speech generation.
Audio super-resolution model for upsampling audio to higher sample rates.
State-of-the-art music source separation model by Meta for splitting tracks.
Fast music generation model producing full songs with lyrics in seconds.
PyTorch library for deep learning research on audio generation including MusicGen and AudioGen.