9.7KUpdated 1 day ago
macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF
Wan2GP brings video, image, music and speech generation to your own computer, with particular attention to GPUs with limited memory. It's for creators who want several media models in one browser interface. The project builds on Wan-Video/Wan2.1.
13.2KUpdated 2 days agoApache-2.0
#ControlNet#Image-to-image#Inpainting
DiffSynth-Studio is a Python diffusion model engine for developers and researchers who want to generate media and train models on their own hardware. It supports large models on consumer GPUs through memory offloading and quantization, with inference and training in the same framework. It's open source under Apache 2.0.
1.3KUpdated 3 years agoApache-2.0
Magenta Studio brings AI music generation into Ableton Live as a Max for Live MIDI plugin. It's for producers who want to develop melodies and drum patterns within their existing Live sessions. The interface runs inside Live on your computer and works with MIDI clips in Session View.
3.9KUpdated 2 years agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#Multimodal input
Riffusion is a Python library for generating music and audio on your own hardware using Stable Diffusion. It's for developers and musicians who want to experiment with text-driven sound generation or build it into an app. The hobby project is no longer actively maintained.
2.6KUpdated 2 years ago
macOS · Linux · Web#Batch processing#Hugging Face integration
AudioLDM 2 generates sound effects, music and speech on your own hardware. It's a Python tool for people experimenting with synthetic audio, including sound designers and researchers who want to work with pretrained models. A Gradio browser interface and command-line tools provide access to local generation; a hosted Hugging Face demo is also available.
2.3KUpdated 10 months agoApache-2.0
macOS · Windows · Linux · Docker#Hugging Face integration#Multimodal input
DiffRhythm is a local AI music generation model for musicians, developers and researchers who want to create full-length songs on their own hardware. It uses latent diffusion to generate songs with vocals and accompaniment, and can also produce instrumental music. The full model supports songs up to 4 minutes and 45 seconds.
4.8KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
Lollms WebUI is a local, single-user AI interface for people who want text chat and media generation in one place. It runs on Windows, macOS and Linux, with Docker support, and lets writers, developers and other users choose models and task-specific personalities. It's free and open source under Apache 2.0. The project receives minimal maintenance.
4.6KUpdated 20 hours agoMIT
macOS · Windows · Linux · Docker · Web#Visual workflows
SwarmUI is a self-hosted AI image generation interface that combines a straightforward generation screen with direct access to ComfyUI's node-based workflows. It's for people who want to run creative models on their own hardware, with room to build more detailed workflows as their needs grow.
23.7KUpdated 2 years agoMIT
#Multimodal input
MusicGen is Meta AI's music generation model within AudioCraft, a PyTorch library for developers and audio researchers who want to generate music in their own computing environment. It creates music from text descriptions and can use a melody to guide the result. AudioCraft includes both inference and training code, so it's suited to people building audio tools or studying music generation.
1.8KUpdated 2 months agoApache-2.0
macOS#MLX#Streaming inference
Magenta RealTime 2 is a local AI music model and synthesis engine for musicians and developers who want to play or build AI musical instruments on a laptop. It generates streaming audio in real time, with open weights and code under the Apache 2.0 license.
3.3KUpdated 3 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
10.6KUpdated 2 days agoApache-2.0
Linux · Web#Hugging Face integration#Multilingual#Multimodal input
YuE is an open-source music generation project for musicians, songwriters and developers who want to turn lyrics and a style prompt into songs with vocals and accompaniment. Its YuE2 models create an editable melody and chord score before generating the recording, so you can review the composition and change musical details before hearing the result.
4.9KUpdated 7 months agoApache-2.0
macOS#LoRA#Multilingual
ACE-Step is an open-source music generation model for musicians, producers and developers who want to create and edit music on their own hardware. It generates songs with vocals or instrumental tracks from text descriptions and supplied lyrics. You can choose the duration and describe the sound with genre tags, longer prompts or a scene description.