Favicon of ACE-Step

ACE-Step

An open-source AI music model that runs on NVIDIA GPUs and Apple Silicon, with text-to-music generation, adjustable duration and localized lyric editing.

Screenshot of ACE-Step website

ACE-Step is an open-source music generation model for musicians, producers and developers who want to create and edit music on their own hardware. It generates songs with vocals or instrumental tracks from text descriptions and supplied lyrics. You can choose the duration and describe the sound with genre tags, longer prompts or a scene description.

Its editing controls let you work on an existing track. Variations produce alternative takes, while localized audio editing can change selected passages and preserve the rest. Lyric editing works with generated music and uploaded audio, keeping the melody, vocal character and accompaniment while replacing words. Edits are limited to short passages because larger changes can cause distortion.

For songwriting, the Lyric2Vocal model generates vocal samples from lyrics, useful for hearing a draft as a sung performance or making guide tracks. Text2Samples generates instrument loops, sound effects and other production material from descriptions. ACE-Step also supports singing in languages including English, Chinese, Japanese and Spanish, though less common languages may produce weaker results.

The model combines diffusion generation with Sana's Deep Compression AutoEncoder and a lightweight linear transformer. It runs on NVIDIA GPUs and Apple Silicon, with performance measurements covering RTX 3090, RTX 4090, A100 and MacBook M2 Max hardware. The code uses the Apache 2.0 license.

Similar to ACE-Step