
ACE-Step is an open-source music generation model for musicians, producers and developers who want to create and edit music on their own hardware. It generates songs with vocals or instrumental tracks from text descriptions and supplied lyrics. You can choose the duration and describe the sound with genre tags, longer prompts or a scene description.
Its editing controls let you work on an existing track. Variations produce alternative takes, while localized audio editing can change selected passages and preserve the rest. Lyric editing works with generated music and uploaded audio, keeping the melody, vocal character and accompaniment while replacing words. Edits are limited to short passages because larger changes can cause distortion.
For songwriting, the Lyric2Vocal model generates vocal samples from lyrics, useful for hearing a draft as a sung performance or making guide tracks. Text2Samples generates instrument loops, sound effects and other production material from descriptions. ACE-Step also supports singing in languages including English, Chinese, Japanese and Spanish, though less common languages may produce weaker results.
The model combines diffusion generation with Sana's Deep Compression AutoEncoder and a lightweight linear transformer. It runs on NVIDIA GPUs and Apple Silicon, with performance measurements covering RTX 3090, RTX 4090, A100 and MacBook M2 Max hardware. The code uses the Apache 2.0 license.
Claim this page and we'll verify you by hand. ACE-Step gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find ACE-Step?Promote it
Something wrong or outdated on this page?
2.6KUpdated 2 years ago
macOS · Linux · Web#Batch processing#Hugging Face integration
AudioLDM 2 generates sound effects, music and speech on your own hardware. It's a Python tool for people experimenting with synthetic audio, including sound designers and researchers who want to work with pretrained models. A Gradio browser interface and command-line tools provide access to local generation; a hosted Hugging Face demo is also available.
2.3KUpdated 10 months agoApache-2.0
macOS · Windows · Linux · Docker#Hugging Face integration#Multimodal input
1.8KUpdated 2 months agoApache-2.0
macOS#MLX#Streaming inference
10.6KUpdated 2 days agoApache-2.0
Linux · Web#Hugging Face integration#Multilingual#Multimodal input
23.7KUpdated 2 years agoMIT
#Multimodal input
MusicGen is Meta AI's music generation model within AudioCraft, a PyTorch library for developers and audio researchers who want to generate music in their own computing environment. It creates music from text descriptions and can use a melody to guide the result. AudioCraft includes both inference and training code, so it's suited to people building audio tools or studying music generation.
135.6KUpdated 47 minutes agoGPL-3.0
macOS · Windows · Linux · Web#ControlNet#Inpainting#LoRA
DiffRhythm is a local AI music generation model for musicians, developers and researchers who want to create full-length songs on their own hardware. It uses latent diffusion to generate songs with vocals and accompaniment, and can also produce instrumental music. The full model supports songs up to 4 minutes and 45 seconds.
Magenta RealTime 2 is a local AI music model and synthesis engine for musicians and developers who want to play or build AI musical instruments on a laptop. It generates streaming audio in real time, with open weights and code under the Apache 2.0 license.
YuE is an open-source music generation project for musicians, songwriters and developers who want to turn lyrics and a style prompt into songs with vocals and accompaniment. Its YuE2 models create an editable melody and chord score before generating the recording, so you can review the composition and change musical details before hearing the result.
ComfyUI is a local visual AI workspace for artists and technical teams who want to control how images, video, audio, 3D models and text are made. Its node canvas shows each model and processing step, so users can build and adjust workflows without writing code. It runs on your hardware.