7.3KUpdated 1 week agoApache-2.0
macOS · Windows · Linux · Docker · Web#ControlNet#Image-to-image#Inpainting
SD.Next is a self-hosted web interface for artists, researchers and people who want to generate and edit images or videos on their own hardware. It builds on Automatic1111 WebUI's original codebase and supports Stable Diffusion alongside other diffusion models. It's open source under Apache 2.0.
91.5KUpdated 7 hours agoApache-2.0
macOS · Windows · Linux · Docker#Hybrid search#MCP#Multi-agent workflows
RAGFlow is an Apache 2.0 licensed RAG engine for teams building AI agents that need to answer questions from their own documents. It can run on a self-hosted server through Docker on Windows, macOS or Linux. A separate hosted cloud service is available.
80.8KUpdated 2 days ago
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
MinerU parses documents locally into structured text for AI agents, RAG systems and knowledge bases. It's for people working with scanned PDFs, academic papers and Office files whose tables, formulas or page layouts need more care than plain text extraction.
24.9KUpdated 1 week agoMIT
Windows · Docker · Web#Batch processing#ONNX
rembg removes image backgrounds on your own hardware, with batch processing and a Python library for developers building image workflows. It's also useful for people preparing product cutouts or portraits who want control over processing and model choice. Local models keep images on your machine and can work offline once downloaded.
5.7KUpdated 2 days agoGPL-3.0
#ONNX
Piper turns text into spoken audio on local hardware. It's an open-source, GPL-3.0 text-to-speech engine for developers building voice features, accessibility tools and self-hosted AI projects. Speech generation runs locally, giving people a way to add a voice to software they control.
8.9KUpdated 10 hours agoMIT
macOS · Windows · Linux · iOS#Batch processing#MCP#Multilingual
OpenWhispr is a free, MIT-licensed dictation and meeting transcription app for people who want voice input across their apps with control over where processing happens. It's available on macOS, Windows, Linux and iOS. Local transcription works offline and keeps audio on your device; optional cloud transcription sends audio to the selected provider, whose retention policies apply.
28.1KUpdated 3 years agoAGPL-3.0
#Hugging Face integration#ONNX#Voice conversion
so-vits-svc is an offline AI framework for changing the voice in an existing singing recording while preserving its pitch and intonation. It's aimed at developers and researchers who want to train their own singing voices, including fictional character voices. The project is archived and no longer maintained.
857Updated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API
OpenedAI Speech is a self-hosted text-to-speech server for developers who want local speech generation in apps built around OpenAI's speech API. The project is archived and no longer maintained. It's open source under AGPL-3.0, and it generates audio on your own hardware without an OpenAI API key.
966Updated 9 months agoGPL-3.0
Windows#Batch processing#Image-to-image#Inpainting
NMKD Stable Diffusion GUI is a local AI image generator for people who want to create and edit images on a Windows PC. It combines Stable Diffusion generation with inpainting, LoRA training and image post-processing in a desktop interface. It's open source under GPL-3.0.
1.9KUpdated 11 months ago
Docker#Batch processing#Hugging Face integration#ONNX
Nomic Embed Text v1.5 is an English text embedding model for developers building semantic search, document retrieval, and RAG applications on their own hardware or servers. It turns text into numerical representations that applications can compare by meaning. Its main distinction is adjustable embedding size: you can use smaller vectors when storage matters, with a tradeoff in retrieval quality.
5.1KUpdated 22 hours agoApache-2.0
Windows · Docker#Agent Skills#Hugging Face integration#ONNX
TensorRT Model Optimizer, called NVIDIA Model Optimizer or ModelOpt, is a Python library for developers preparing models for local or self-hosted inference. It reduces model size and memory use and can speed up inference through compression and other optimization techniques. It's open source under Apache 2.0.
11.8KUpdated 4 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
Ludwig is an open-source Python framework for developers and researchers who want to train custom AI models on their own hardware. A YAML file describes the model and training pipeline, while Ludwig handles preprocessing, training and evaluation. It uses the Apache 2.0 license. Install the Python package with the optional LLM dependencies for fine-tuning; current source requires Python 3.12 or later.
16.2KUpdated 1 day agoApache-2.0
Windows · iOS · Android#Image-to-image#Multimodal input#ONNX
MNN is a lightweight C++ framework for developers who want AI models to run on phones, PCs and embedded devices. It handles inference and training on the device, with a focus on small application footprints and hardware acceleration. The project is open source under Apache 2.0, and Alibaba uses it in apps including Taobao, Youku and DingTalk.
8.1KUpdated 1 day agoMIT
Windows · Linux · iOS · Android · Docker · Web#Multi-user access#ONNX#Semantic search
LibrePhotos is a self-hosted photo manager for people who want to organize a photo library on their own server, with AI tools for finding images and grouping faces. It scans files already on your storage and supports RAW photos as well as videos. The project is open source under the MIT license.
8.7KUpdated 2 years agoMIT
Linux#Hugging Face integration#Multimodal input#ONNX
Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.
606Updated 4 days agoApache-2.0
Windows#ControlNet#GGUF#Inpainting
Amuse combines AI generation with media editing in a Windows app that runs models on your own hardware. It's for people who want to create images, video, audio and text locally, then work on the results in the same application. Its editor also accepts existing local video.
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
31.4KUpdated 1 week agoApache-2.0
macOS#Distributed execution#ONNX
PyTorch Lightning is a Python framework for researchers and developers who want to pretrain or fine-tune models on their own hardware or in the cloud. It handles repetitive training code while leaving model logic under your control. The framework is open source under Apache 2.0.
5.7KUpdated 2 weeks agoMIT
macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows
GLaDOS is a local AI voice assistant modeled on the sarcastic character from Valve's Portal games. It's for people who want a conversational companion on their own hardware, with camera awareness and connections to home automation. The Python project is open source under the MIT license and runs on Linux and Windows. macOS support is experimental.
2.8KUpdated 9 months agoApache-2.0
Windows · Linux#Batch processing#ONNX#Voice activity detection
openWakeWord is a Python library for developers building voice interfaces that listen locally for a chosen word or phrase. It includes English models for triggers such as "hey jarvis" and "alexa", plus phrases for weather and timers. The code uses Apache 2.0. Included pretrained models use CC-BY-NC-SA-4.0, which restricts commercial use.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
3.5KUpdated 6 days agoApache-2.0
#Hugging Face integration#ONNX#Quantization
Optimum is a collection of Python packages for developers who want to train or run Hugging Face models more efficiently on specific hardware. It extends Transformers, Diffusers, TIMM and Sentence Transformers, with integrations for local machines, mobile and edge devices, and cloud accelerators. It's open source under Apache 2.0.
38.6KUpdated 2 months agoMIT
Windows · Linux · Web#Hugging Face integration#ONNX#Voice conversion
RVC WebUI is a local AI voice conversion tool for people who want to train a custom voice, change the voice in a recording, or use a live voice changer. It runs on Windows and Linux, including Ubuntu servers, with a browser interface for training and conversion and a separate interface for live use. It's free and open source under the MIT license.
9.2KUpdated 7 days agoApache-2.0
Docker#Image-to-image#Inpainting#Multimodal input
ModelScope combines a hosted model and dataset hub with a Python library you can run locally. It's for developers and researchers who want to use AI models in their own applications, fine-tune them on their own data, or compare their performance. The library is open source under Apache 2.0.
44KUpdated 18 hours agoApache-2.0
#Batch processing#Hugging Face integration#ONNX
Ray Serve is a self-hosted Python library for developers building inference APIs that combine models with application logic. It runs on a laptop, on-premise servers, Kubernetes, or cloud infrastructure you choose. It's open source under Apache 2.0.
214Updated 3 weeks agoMIT
Linux · Docker · Web#Home Assistant integration#Hugging Face integration#Multilingual
Wyoming Piper connects Piper's local text-to-speech engine to Home Assistant and other clients that use the Wyoming protocol. It's for people building a voice assistant on their own hardware who need speech generation as a self-hosted service. The project is open source under the MIT license.
29.9KUpdated 3 weeks ago
macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization
InsightFace is a face analysis toolkit for developers and teams building identity verification, access control, or face editing software. The code uses the MIT license. Its Python tools and self-hosted recognition server run inference on your own hardware. It also offers commercial models and API access for face swapping and deepfake detection.
4.9KUpdated 3 months ago
Linux · Docker#llama.cpp backend#Multimodal input#Ollama integration
jetson-containers is a Docker container build system for developers running local AI and robotics workloads on NVIDIA Jetson hardware. It supplies prebuilt images and lets you combine AI packages into custom containers, reducing the work of assembling compatible GPU software for JetPack/L4T.
21.1KUpdated 4 days ago
macOS · Windows · Linux · Docker#ONNX#Voice conversion
Voice Changer (w-okada), also called VCClient, converts your voice as you speak using AI voice models. It's for people who want live voice conversion on their own computer, including those recording gaming commentary while running demanding software. Processing can stay local.
8.2KUpdated 2 years ago
#Batch processing#GGUF#Hugging Face integration
GOT-OCR2.0 is an OCR model for developers and researchers who want to extract text from images on their own hardware. It handles both plain text and formatted output through a single model, with recognition modes for selected regions and documents spanning multiple pages. The Python codebase builds on Vary.
1.6KUpdated 9 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
rerankers is a Python library for developers building search and retrieval systems who want to compare reranking models without rewriting their integration each time. It takes a query and candidate documents, then ranks their relevance through a shared interface across local models and hosted services. It's open source under Apache 2.0.
16.8KUpdated 1 day agoMIT
Docker · Web#Hugging Face integration#Multi-user access#ONNX
CVAT is a browser-based data annotation platform for teams building computer vision datasets. Its open-source Community edition runs on your own infrastructure with Docker and uses the MIT license. CVAT Online is hosted by CVAT, while the Enterprise offering runs in an organization's own cloud or internal environment.
11KUpdated 1 week agoBSD-3-Clause
Windows · Linux · Docker#Batch processing#ONNX
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.
22KUpdated 20 minutes agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
9.1KUpdated 1 year agoApache-2.0
macOS · Windows#Batch processing#Multilingual#ONNX
Kokoro is a text-to-speech model and inference library for developers who want to generate speech on their own hardware or servers. Its compact Kokoro-82M model suits personal projects and production applications, with Apache 2.0 licensing for both the library and model weights.