2.9KUpdated 9 months agoApache-2.0
Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization
Whisper WebUI turns audio into transcripts and subtitles through a browser interface that runs on your own machine or a self-hosted server. It's for people captioning videos, transcribing recordings or translating spoken content who want local speech processing. The project is open source under Apache 2.0 and supports Docker and Pinokio.
18.2KUpdated 10 months ago
#Image-to-image#Inpainting
CodeFormer restores degraded faces in photos and videos using a learned library of facial details and a Transformer model. It's for people restoring images on their own hardware, developers adding face enhancement to a project, and researchers working on image restoration. Its main distinction is an adjustable balance between visual quality and fidelity to the input face.
6.6KUpdated 1 year ago
Windows · Linux · Web#Batch processing#Inpainting#Multilingual
MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
24.3KUpdated 4 days agoBSD-2-Clause
macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual
WhisperX is an open source speech-to-text tool for people transcribing interviews, meetings, and long recordings on their own computer. It builds on OpenAI's Whisper to produce transcripts with word-level timestamps and optional speaker labels.
17.3KUpdated 11 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.
21.8KUpdated 1 week agoMIT
macOS · Windows · Linux#Hugging Face integration#Multilingual#Speaker diarization
Buzz transcribes and translates speech on your own computer using OpenAI's Whisper. It's for people who need transcripts or subtitles from recordings, plus live captions from a microphone. Local transcription works offline; the optional OpenAI Whisper API sends audio to a cloud service.
575Updated 1 day agoGPL-3.0
macOS · Linux · iOS · Docker#Image-to-image#Inpainting#LoRA
Draw Things is an AI image generation app for iPhone, iPad and Mac that keeps generation on your device and works offline. It's for people who want to create and edit images without sending that work to a cloud service, including artists developing character concepts or trying out apparel designs.
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
19.1KUpdated 4 months ago
macOS · Windows · Linux · Web#Hugging Face integration
LivePortrait animates a still portrait using facial expressions and head movements from a driving video. It's for creators who want to animate faces or edit motion in portrait videos on their own hardware. It works with realistic photos as well as portraits in oil paintings, sculptures and 3D renders.
13.6KUpdated 3 years agoAGPL-3.0
macOS#ControlNet#Image-to-image#Inpainting
DiffusionBee is an open-source AI art app for Mac users who want to generate and edit images on their own computer. It runs Stable Diffusion offline, with image generation processed on the device. Model downloads require network access, and optional image uploads can send images externally. Its visual interface suits artists and designers who want local image tools without working through code.
881Updated 3 days ago
Docker#Multilingual#Ollama integration
Lingarr is a self-hosted subtitle translator for people who maintain a media library and want subtitles in another language. It automates translation of subtitle files using a service you choose, with support for both local AI and hosted translation providers.
1.5KUpdated 2 months agoMIT
Docker#Batch processing#Multilingual#OpenAI-compatible API
subgen generates subtitles on your own hardware for personal media libraries, including films and shows that don't have usable subtitles available. It's an open source, MIT-licensed Python service that runs in Docker or as a standalone application. Speech recognition runs locally using Whisper models through faster-whisper and stable-ts, with support for CPU processing and NVIDIA GPUs through CUDA.
7.7KUpdated 4 weeks agoMIT
macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration
Vibe is an open source desktop app for people who need transcripts or subtitles without uploading their recordings to a transcription service. It runs on macOS, Windows and Linux under the MIT license. Audio transcription works fully offline, with processing on your own computer.
36.9KUpdated 2 years agoBSD-3-Clause
macOS · Windows · Linux#Batch processing
Real-ESRGAN is a local AI upscaler and restoration toolkit for people enlarging images or improving video, with dedicated models for anime illustrations and animation. It builds on ESRGAN and uses models trained entirely on synthetic data to address degraded images. The code is open source under the BSD 3-Clause license.
1KUpdated 1 month agoMIT
macOS
Practical-RIFE is a local AI frame interpolation tool for engineers and developers working with video. It generates intermediate frames to increase frame rates, with models intended for ordinary footage, animation, and post-processing videos made by diffusion models. The Python project builds on RIFE and SAFA, with an emphasis on how the output looks rather than improvements in numerical image-quality scores alone.
3.3KUpdated 2 months agoMIT
Docker · Web#Multilingual#Speaker diarization#Voice activity detection
Whisper ASR Webservice turns Whisper speech recognition into a self-hosted API for developers adding transcription to their apps or services. It runs in Docker on your own machine or server, with CPU processing or CUDA GPU acceleration. The Python project is open source under the MIT license.
6KUpdated 2 months agoGPL-3.0
macOS · Windows · Linux#Batch processing#ONNX#Visual workflows
chaiNNer is a desktop image processing editor for people who want AI upscaling and repeatable editing workflows without writing scripts. It runs models locally on Windows, macOS, and Linux. Connected nodes let you combine model processing with ordinary image edits in the same workflow.
1.7KUpdated 1 week agoMPL-2.0
Linux#Multilingual#Works offline
Speech Note combines offline dictation, reading aloud and translation in a desktop app for Linux and Sailfish OS. It's for people who want to take multilingual notes, type by voice or listen to text without sending their words to a cloud service. Speech and text processing stay on your device; models are downloaded separately through the app's graphical browser.
21.9KUpdated 7 months agoAGPL-3.0
macOS · Windows · Linux · Docker
Video2X runs AI video upscaling and frame interpolation on your own hardware. It's for people who want to increase a video's resolution or generate extra frames for smoother motion, including those working with anime footage. The project is open source under AGPL-3.0.
14.4KUpdated 1 day agoMIT
macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual
Subtitle Edit is an MIT-licensed subtitle editor for Windows, macOS and Linux. It's for people creating captions, translating dialogue or fixing subtitles that don't match the video. Its core editing, conversion and video playback work offline on your device, with optional AI tools for transcription and translation.
12.6KUpdated 3 months ago
Web#Multimodal input#Quantization
HunyuanVideo is an AI video generation model for creators and developers who want to generate footage on their own hardware. Tencent provides model weights and inference code for text-to-video and image-to-video generation, alongside a hosted web experience. Local inference runs on your GPUs; the web offering runs through Tencent's service.
11KUpdated 9 months agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
LTX-Video is an AI video generation model for creators building controlled animations and developers adding video tools to their own products. You can run it locally or on your own servers using publicly available weights. The LTX family also offers a managed cloud API; local deployments can run in isolated environments without a cloud dependency.