17.8KUpdated 2 years agoMIT
Web#Batch processing#Hugging Face integration#Multimodal input
Janus-Pro is a multimodal AI model from DeepSeek that answers questions about images and creates pictures from text prompts. It runs on your own hardware and suits developers and researchers who want both capabilities in one model. A local Gradio demo provides a browser interface, while Hugging Face hosts a separate online demo.
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.
1.1KUpdated 7 months agoAGPL-3.0
Web · Browser Extension#Multilingual#Multimodal input#Ollama integration
NativeMind brings a local AI assistant into Chrome for people who want help with webpages, documents, and writing without sending that content to a cloud model. The browser extension connects to Ollama on your machine, keeping prompts and AI processing on-device. It requires no account and uses the AGPL-3.0 open-source license.
19.4KUpdated 10 months agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input#Voice cloning
Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.
2KUpdated 7 months agoMIT
macOS · Windows · Linux · Web#MCP#RAG
NotebookLlama is a self-hosted NotebookLM alternative for people who want to work with documents through an app they can host and modify. Its browser interface runs locally, while LlamaCloud provides document extraction and indexing. It requires cloud services.
3.5KUpdated 20 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
1.2KUpdated 12 months agoMIT
macOS · Windows · Linux · Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API
Hollama is an open-source LLM chat app whose interface runs entirely in your browser. It's for people who want to chat with local AI through Ollama or connect to OpenAI servers, with support for multiple server connections. The interface stores data locally in the browser; the connected server handles model requests, so where inference runs depends on the server you choose.
4.4KUpdated 20 hours agoMIT
macOS · Windows · Linux · Web · JetBrains#Agent Client Protocol#Code execution#Git integration
gptme is a self-hosted AI agent that works directly in your terminal, with access to your files and installed tools. It's for developers who want a coding assistant in their own environment, and people who want an agent for data analysis or other knowledge work. The software is free under the MIT license and doesn't require a gptme account.
6.4KUpdated 1 day agoApache-2.0
Docker · Web
docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.
3.3KUpdated 2 weeks agoApache-2.0
Docker · Web#LLM tracing#MCP
Laminar is an open-source platform for developers who need to see why an AI agent failed and check whether a fix worked. You can self-host it with Docker or on Kubernetes, including AWS and GCP, or use its managed cloud service. It uses the Apache 2.0 license.
699Updated 2 weeks agoAGPL-3.0
Linux · Docker · Web
Nextcloud Recognize automatically categorizes media on your self-hosted Nextcloud server. It's for people who want to organize photo, video and music collections without sending sensitive data to a cloud recognition service. Image processing runs on your Nextcloud machine.
5.1KUpdated 21 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
3.1KUpdated 1 year agoAGPL-3.0
Web#Multilingual#Works offline
Whishper is a self-hosted speech-to-text app for people who need transcripts or translated subtitles from audio and video. Its browser interface brings transcription, translation and subtitle editing together, with all three running on your own machine. It can work offline, so local media doesn't need to go to a cloud transcription service.
4.3KUpdated 6 months agoApache-2.0
Linux · Web#Hugging Face integration#Multimodal input
EchoMimic turns a portrait image and an audio recording into an animated talking-head video on your own hardware. Its distinguishing feature is editable facial landmark control: you can drive animation with audio, landmarks, or audio combined with selected landmarks. It's intended for academic research and suits people comparing methods for speech animation and facial motion control.
8.2KUpdated 3 days agoMIT
Web · Browser Extension#LM Studio integration#Ollama integration#OpenAI-compatible API
Page Assist brings local AI chat into your browser, with a sidebar for conversations alongside a webpage and a separate tab for a ChatGPT-style interface. It's for people who already run models on their own hardware and want to ask questions about what they're reading without leaving the page.
11KUpdated 1 day agoApache-2.0
Docker · Web#llama.cpp backend#MCP#Multi-user access
HuggingChat UI is the open-source chat application behind Hugging Face's hosted HuggingChat. You can run it on your own computer or server and connect it to a local LLM backend or a cloud provider. It's for people and teams who want a browser-based ChatGPT alternative with control over the chat service and where its data lives. The code uses the Apache 2.0 license.
752Updated 3 weeks agoApache-2.0
Linux · Docker · Web · Browser Extension#Batch processing#OpenAI-compatible API#Quantization
Meme Search is a free, self-hosted web app for people who want to find memes in their own collection by image content or text. It runs locally through Docker and uses AI descriptions to make images searchable, even when their filenames aren't useful. The code uses the Apache 2.0 license.
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
37.6KUpdated 19 hours agoMIT
iOS · Android · Web#Human approval#MCP#Persistent memory
CopilotKit is a self-hostable SDK for developers building AI agents into web and mobile apps, Slack, or Microsoft Teams. Agents can display interactive charts and forms using an app's own components, read shared app state, and take actions through frontend tools, APIs, or MCP tools.
1.5KUpdated 3 weeks agoMIT
Docker · Web#Ollama integration#OpenAI-compatible API#Streaming inference
Unmute adds spoken conversation to text LLMs using Kyutai's speech recognition and speech synthesis models. It's for developers who want a self-hosted voice interface while keeping their choice of language model. The project uses the MIT license, and a hosted browser demo is available at Unmute.sh.
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
2.6KUpdated 21 hours agoApache-2.0
Web#OpenAI-compatible API#Prompt caching
vLLM Production Stack is an open source inference stack for teams serving LLMs on their own Kubernetes GPU clusters. It brings request routing and monitoring around vLLM, so applications can move from one serving instance to a distributed deployment without changing their code. It requires a GPU-enabled Kubernetes environment.
3.4KUpdated 23 hours agoApache-2.0
Docker · Web#Multilingual#Multimodal input
MTEB is an Apache 2.0 Python toolkit for evaluating embedding models and retrieval systems. It runs evaluations through Python or a command-line interface and publishes an interactive leaderboard.
7.7KUpdated 2 years agoMIT
Docker · Web#Multilingual
MeloTTS is a Python text-to-speech library for developers who want to generate speech locally, including on machines without a dedicated GPU. It supports real-time inference on a CPU. Its language and accent choices make it relevant for applications that need spoken output across different audiences.
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
2.7KUpdated 6 days agoMIT
Docker · Web#Batch processing#Human approval#Multimodal input
Paperless-GPT adds AI text extraction and document organization to an existing paperless-ngx library. It runs in Docker on your own server and suits people who want less manual sorting of scanned paperwork. The project is open source under the MIT license.
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
8.3KUpdated 3 years agoApache-2.0
macOS · Windows · Linux · Docker · Web#Multi-user access#Role-based access
CompreFace is a self-hosted face recognition service for developers who want to add facial identification to an application without building or training their own machine learning system. It runs as a Docker-based server on your hardware or in a cloud deployment you manage. It's free and open source under the Apache 2.0 license.
2.9KUpdated 20 hours agoAGPL-3.0
macOS · Docker · Web#ControlNet#Distributed execution#Human approval
SimpleTuner is an open-source toolkit for fine-tuning image, video and audio generation models on your own hardware or GPU servers. It's for creators and researchers adapting models to their datasets, and teams sharing training infrastructure. A web dashboard manages training jobs.
2.7KUpdated 3 months agoMIT
Docker · Web#MCP#Multi-user access#Single sign-on
MetaMCP is a self-hosted gateway for developers and teams who want to give AI clients access to several MCP servers through one endpoint. It runs on your own machine or server with Docker and is open source under the MIT license. You choose which tools clients see.
37.1KUpdated 20 hours agoApache-2.0
iOS · Android · Web
MediaPipe is an open-source toolkit for developers adding on-device AI to applications on Android, iOS, the web, desktop and edge devices. It pairs pretrained models with APIs for specific tasks, so developers can use existing solutions or customize them for their applications. The project uses the Apache 2.0 license.
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
38.6KUpdated 2 months agoMIT
Windows · Linux · Web#Hugging Face integration#ONNX#Voice conversion
RVC WebUI is a local AI voice conversion tool for people who want to train a custom voice, change the voice in a recording, or use a live voice changer. It runs on Windows and Linux, including Ubuntu servers, with a browser interface for training and conversion and a separate interface for live use. It's free and open source under the MIT license.