2.1KUpdated 6 hours agoAGPL-3.0
macOS · Windows · Linux#Agent Skills#MCP#Multilingual
OpenChatCut is a local-first AI video editor for creators who want conversational editing with control over the finished cut. AI changes become editable clips, captions, effects and audio tracks in the same project you can adjust manually. It's a free, open-source ChatCut alternative under AGPL-3.0, with a desktop app for macOS, Windows and Linux.
742Updated 4 hours agoAGPL-3.0
Web#MCP#Tool calling#Works offline
CozyClay is a local 3D previsualization studio for filmmakers and AI video creators who want to plan framing and movement before generating a finished clip. You build a rough scene, pose characters, and edit camera moves and cuts on a timeline. Those shots become visual references for models such as Seedance, Kling and Veo, so you can specify the camera through a clip rather than relying on text alone.
9.7KUpdated 1 day ago
macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF
Wan2GP brings video, image, music and speech generation to your own computer, with particular attention to GPUs with limited memory. It's for creators who want several media models in one browser interface. The project builds on Wan-Video/Wan2.1.
13.2KUpdated 2 days agoApache-2.0
#ControlNet#Image-to-image#Inpainting
DiffSynth-Studio is a Python diffusion model engine for developers and researchers who want to generate media and train models on their own hardware. It supports large models on consumer GPUs through memory offloading and quantization, with inference and training in the same framework. It's open source under Apache 2.0.
34.6KUpdated 1 day agoApache-2.0
macOS#ControlNet#Hugging Face integration#Image-to-image
Diffusers is an open-source Python library for developers and researchers who want to run diffusion models on their own hardware or build generation features into an application. It uses PyTorch and supports image, video and audio generation. The library is licensed under Apache 2.0 and supports Apple Silicon.
7.3KUpdated 1 week agoApache-2.0
macOS · Windows · Linux · Docker · Web#ControlNet#Image-to-image#Inpainting
SD.Next is a self-hosted web interface for artists, researchers and people who want to generate and edit images or videos on their own hardware. It builds on Automatic1111 WebUI's original codebase and supports Stable Diffusion alongside other diffusion models. It's open source under Apache 2.0.
30.1KUpdated 1 day ago
macOS · Windows#Image-to-image
FaceFusion is an AI face manipulation tool for photos and videos that processes media on your own machine. It's aimed at content creators, VFX artists, and film studios that want control over their footage and personal data. Local processing keeps that media on your hardware rather than sending it to a cloud service.
135.6KUpdated 49 minutes agoGPL-3.0
macOS · Windows · Linux · Web#ControlNet#Inpainting#LoRA
ComfyUI is a local visual AI workspace for artists and technical teams who want to control how images, video, audio, 3D models and text are made. Its node canvas shows each model and processing step, so users can build and adjust workflows without writing code. It runs on your hardware.
17.7KUpdated 1 week agoApache-2.0
#Hugging Face integration
Wan2.2 is an open source family of video generation models for creators and developers who want to make clips on their own GPUs. Licensed under Apache 2.0, it covers text-to-video and image-to-video generation, with separate models for speech-driven video and character animation. Generation runs on your hardware when you use the downloadable models.
77KUpdated 21 hours agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Image-to-image
Unsloth brings model training and everyday AI use into a desktop app for people who want to run models on their own hardware. Its no-code interface covers chat, fine-tuning and media generation on macOS, Windows and Linux. The Unsloth software is open source under Apache 2.0.
17.1KUpdated 2 weeks ago
Windows#Batch processing#Works offline
Waifu2x-Extension-GUI is a local AI upscaler for Windows users working with anime, photos and video. It combines image enlargement, noise reduction and video frame interpolation in a desktop interface. Media processing stays on your PC; the app doesn't upload your files or collect user data.
sindresorhus.comOn-Device and In-Browser AI
macOS · iOS#Batch processing#Multilingual
Aiko is a paid, native transcription app for macOS, iOS and visionOS that processes speech on your device with OpenAI's Whisper model. It's for people turning meetings, lectures or other recordings into text while keeping the audio local, including sensitive recordings.
1.2KUpdated 9 months agoMIT
macOS · Windows · Linux#Batch processing#Image-to-image
AI Render is a Blender add-on that uses your scene and a text prompt to generate an image with Stable Diffusion. It's for Blender artists who want to use their 3D work as the basis for AI images and explore different treatments within the application they already use.
13.1KUpdated 2 years agoApache-2.0
macOS · Windows#Batch processing#Hugging Face integration#Multilingual
insanely-fast-whisper is a command-line tool for people who want to transcribe audio on their own hardware, with a focus on processing long recordings quickly. It runs OpenAI's Whisper locally on NVIDIA GPUs or Apple Silicon Macs, including support for Windows with CUDA. The project is open source under the Apache 2.0 license.
520Updated 13 hours ago
macOS · Windows · Linux#Visual workflows#Works offline
Comfy Desktop installs and launches ComfyUI on your computer, taking care of the Python environment and dependencies. You can manage several independent ComfyUI instances, each with its own version, custom nodes and settings. It suits people who want to run local image and media workflows without setting up every environment by hand.
3.2KUpdated 2 years agoMIT
macOS · Web#Hugging Face integration#Multimodal input
Pyramid Flow is an open-source AI video generator for people who want to create clips on their own hardware, and for researchers working on video models. It turns text prompts into video or animates a supplied image with guidance from text. The Python code uses PyTorch and carries MIT. Model licenses are separate: the SD3-derived weights use the Stability AI Community License. Check the chosen checkpoint before deployment.
12.3KUpdated 2 years agoApache-2.0
Web#Hugging Face integration#LoRA#Multimodal input
AnimateDiff adds text-driven animation to personalized Stable Diffusion models without requiring separate training for each model. It's for artists and developers who want to generate motion while keeping the visual style of a chosen image model. The Python implementation runs locally and includes a Gradio browser interface.
29.9KUpdated 6 months agoApache-2.0
#Hugging Face integration#Multimodal input
Open-Sora is an open source AI video generation project for developers, researchers, and creators who want to run and adapt a model on their own hardware. Its model focuses on turning reference images into video, with text prompts guiding the result. It also generates video directly from text. The code and Open-Sora 2.0 weights use Apache 2.0.
7.6KUpdated 8 months ago
Web#Hugging Face integration#Multimodal input
SkyReels-V2 is a family of downloadable video models for developers and researchers who want to run generation on their own GPUs. It supports text-to-video, image-to-video and extended video generation. Its autoregressive Diffusion Forcing architecture generates longer sequences by continuing video segments.
3.7KUpdated 11 months agoApache-2.0
Web#Hugging Face integration#LoRA
Mochi 1 is a text-to-video model for creators and developers who want to generate videos on their own hardware or adapt a model to their own footage. Genmo releases it under Apache 2.0, with downloadable weights and code for local use. Genmo also offers a hosted playground for trying the model in a browser.
13KUpdated 11 months agoApache-2.0
Windows · Web#Hugging Face integration#LoRA#Multimodal input
CogVideoX is a family of downloadable video generation models for developers, researchers and creators who want to generate clips on their own hardware. It turns English text prompts into video, animates a supplied image and can continue an existing video. A local Gradio web interface provides a browser front end for generation.
4.8KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
Lollms WebUI is a local, single-user AI interface for people who want text chat and media generation in one place. It runs on Windows, macOS and Linux, with Docker support, and lets writers, developers and other users choose models and task-specific personalities. It's free and open source under Apache 2.0. The project receives minimal maintenance.
8.7KUpdated 2 years agoMIT
Linux#Hugging Face integration#Multimodal input#ONNX
Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.
4.6KUpdated 20 hours agoMIT
macOS · Windows · Linux · Docker · Web#Visual workflows
SwarmUI is a self-hosted AI image generation interface that combines a straightforward generation screen with direct access to ComfyUI's node-based workflows. It's for people who want to run creative models on their own hardware, with room to build more detailed workflows as their needs grow.
2.1KUpdated 4 months agoGPL-3.0
Windows
Flowframes is a Windows desktop app that uses local AI to generate intermediate frames for smoother video and animation. It's for people working with camera footage, 2D animation or rendered video who want a graphical interface for interpolation. The app is open source under GPL-3.0 and available as donationware, with free builds.
606Updated 4 days agoApache-2.0
Windows#ControlNet#GGUF#Inpainting
Amuse combines AI generation with media editing in a Windows app that runs models on your own hardware. It's for people who want to create images, video, audio and text locally, then work on the results in the same application. Its editor also accepts existing local video.
2.3KUpdated 2 years ago
Windows · Web#Hugging Face integration
Deforum is a free Stable Diffusion tool for artists and developers who want to generate images and animations on their own hardware. The project is no longer maintained. Its code uses MIT, while downloaded Stable Diffusion models retain their own weight licenses. Its Python script and Jupyter notebook support local generation, while Google Colab and Replicate offer cloud execution.
9.2KUpdated 2 weeks agoApache-2.0
#ControlNet#LoRA#Multimodal input
Sana is an open-source framework for running image and video generation on your own hardware, with image models small enough for laptop GPUs. It's aimed at creators who want local AI generation and developers who need training and inference pipelines for their own models. The code uses the Apache 2.0 license.
13.2KUpdated 1 year ago
Linux · Docker#Multilingual#Multimodal input
Wav2Lip is a local AI lip-sync tool that changes a face's mouth movements in an existing video to match supplied speech. It's for researchers and people making academic or personal video projects who want to process their own files. Commercial use is prohibited under the project's stated terms because its models were trained on the LRS2 dataset.
3.1KUpdated 1 year agoAGPL-3.0
Web#Multilingual#Works offline
Whishper is a self-hosted speech-to-text app for people who need transcripts or translated subtitles from audio and video. Its browser interface brings transcription, translation and subtitle editing together, with all three running on your own machine. It can work offline, so local media doesn't need to go to a cloud transcription service.
2.1KUpdated 3 days ago
Windows · Linux#LoRA#Quantization
Musubi Tuner is a Python toolkit for training LoRA adapters for image and video generation models on your own hardware. It's aimed at people who want to customize these models using their own datasets and are comfortable working with training scripts. It also includes image and video generation scripts for supported architectures.
4.3KUpdated 6 months agoApache-2.0
Linux · Web#Hugging Face integration#Multimodal input
EchoMimic turns a portrait image and an audio recording into an animated talking-head video on your own hardware. Its distinguishing feature is editable facial landmark control: you can drive animation with audio, landmarks, or audio combined with selected landmarks. It's intended for academic research and suits people comparing methods for speech animation and facial motion control.
goodsnooze.gumroad.comDictation and Voice Typing
macOS · iOS#Batch processing#Multilingual#Ollama integration
MacWhisper is a native macOS transcription app for people working with interviews, lectures, meetings and other recorded audio. It runs speech recognition on your own Mac, so local transcription keeps audio on your device. It also offers cloud transcription through services such as OpenAI, ElevenLabs and Deepgram, which send audio off your machine.
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
immersivetranslate.comOCR and Document Scanning
macOS · iOS · Android · Browser Extension#Inpainting#Multilingual
Immersive Translate is an AI translation extension and mobile app that keeps original text alongside its translation. It's aimed at students, researchers and people who read foreign-language material for work. The bilingual page layout lets readers compare passages without replacing the source text.
14.1KUpdated 3 years ago
macOS · Windows · Linux · Web#Multimodal input#Works offline
SadTalker generates talking head videos from a single portrait and an audio recording. The project states an Apache 2.0 license and removal of its earlier noncommercial restriction. It runs locally on Windows, Linux and macOS, and suits creators who want to animate a face without recording a person on camera. Its animation includes facial expressions and head movement, with examples covering speech and singing in different languages.