6.6KUpdated 1 year ago
Windows · Linux · Web#Batch processing#Inpainting#Multilingual
MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.
40.3KUpdated 22 hours ago
macOS · Windows · Linux · Docker · Web
PhotoPrism is a self-hosted photo and video library for people who want to organize personal media on their own hardware or server. Its AI recognizes faces and labels pictures by content and location, so finding a photo doesn't depend entirely on folders or tags you've added yourself. You browse and share the library through a web app.
13KUpdated 1 year agoAGPL-3.0
Windows · Web#ControlNet#GGUF#Image-to-image
Stable Diffusion WebUI Forge runs image generation on your own hardware through a browser interface. It builds on Stable Diffusion WebUI and suits people who want its image creation tools with more control over GPU memory use, as well as developers extending those tools. It's open source under AGPL-3.0.
22KUpdated 2 hours agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
165.2KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Web#Batch processing#Code execution#Image-to-image
Stable Diffusion web UI (AUTOMATIC1111) is a browser interface for generating and editing images with models running on your own hardware. It's for artists and anyone who wants control over prompts, models and image variations. The software is open source under AGPL-3.0.
12.6KUpdated 3 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#ControlNet#Hugging Face integration#LoRA
Kohya's GUI lets you train and fine-tune image generation models on your own GPU-equipped computer through a browser interface. It's for artists and model makers who want to teach a model a particular style or subject while controlling the training settings. The interface builds on Kohya's Stable Diffusion training scripts, with a command-line interface available too.
11KUpdated 1 week ago
Linux · Web#MCP
MCP Inspector is a locally run developer tool for testing Model Context Protocol (MCP) servers. It's for developers building or checking the servers that connect AI applications to tools and data. You can inspect a local server or connect to a remote HTTP endpoint.
4.4KUpdated 2 days agoMIT
Web#LoRA#Multimodal input#Ollama integration
Ollama JavaScript connects Node.js and browser applications to models running through Ollama. It's for developers building chat interfaces, AI agents or other apps that need a local LLM backend. The library is open source under the MIT license, with TypeScript types and an API that follows Ollama's REST interface.
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
10.8KUpdated 8 months agoMIT
Windows · Docker · Web#Multi-user access#Multilingual
Doccano is a self-hosted text annotation tool for machine learning practitioners who need labeled training or evaluation data. It runs on your own machine or server, with a browser interface and Docker support. The software is open source under the MIT license.
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
28.2KUpdated 2 hours agoApache-2.0
macOS · Windows · Linux · Web · VS Code · JetBrains#Code execution#Git integration#MCP
Qwen Code is an Apache 2.0 licensed AI coding agent for developers who want help working through a codebase, changing code and checking the result. It runs on macOS, Windows and Linux, with a terminal interface and a desktop app. It builds on Google Gemini CLI and has developed into an agent that can use Qwen models alongside other model providers.
10.4KUpdated 6 days ago
Web#Code execution#Visual workflows
Typebot is a visual chatbot builder you can host on your own server or use through its managed cloud service. It's for businesses building conversational forms and chatbots for websites, mobile apps, and WhatsApp, with developers able to extend flows through APIs and JavaScript. The project uses Fair Source licensing.
21.7KUpdated 23 hours agoApache-2.0
Docker · Web#Human approval#MCP#Multi-agent workflows
Google ADK is an open-source AI agent framework for developers building applications that carry out multi-step tasks. You can run agents locally in Docker or on your own infrastructure, and connect them to locally running models through adapters. The framework is optimized for Gemini but supports other models and providers. Its Python repository uses the Apache 2.0 license.
20.4KUpdated 3 months agoMIT
Web#Multimodal input#Tool calling
SWE-agent is an open-source AI coding agent that lets a language model use tools to attempt fixes for issues in GitHub repositories. It's aimed at developers and researchers studying how agents handle real software tasks. The project is in maintenance-only mode: mini-swe-agent has superseded it, and the maintainers recommend that successor for new users.
17.8KUpdated 1 day ago
Web#Git integration#Guardrails#LLM tracing
Wren AI is a self-hosted data agent for teams and agent builders who need answers based on agreed business definitions. It turns plain-language questions into SQL and interactive dashboards, using the same definitions when someone asks directly or through Claude, ChatGPT or Gemini over MCP.
25.8KUpdated 4 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Hybrid search#llama.cpp backend#Multi-user access
kotaemon is a self-hosted document chat app for people who want to ask questions across their files and check where the answers came from. It runs in a browser on Windows, macOS or Linux, with Docker also supported. The project uses the Apache 2.0 license.
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
16.3KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Image-to-image#Multilingual
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
10.5KUpdated 7 months agoApache-2.0
macOS · Windows · Linux · Android · Web#Code execution#MCP#Multimodal input
aichat brings Ollama and cloud AI services into the same terminal interface for developers and people who work at the command line. It runs locally on macOS, Linux and Windows, with Android support through Termux. Model processing happens through the backend you choose: Ollama supports local models, while providers such as OpenAI, Claude and Gemini process requests in the cloud.
17.3KUpdated 12 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.
395Updated 4 weeks ago
Docker · Web#Multi-user access#OpenAI-compatible API
Feeds Fun is a news reader for people whose RSS subscriptions produce more articles than they want to read. It assigns tags automatically, then uses rules you define to score articles by topic. You can self-host it with Docker or use the hosted service at feeds.fun.
33.9KUpdated 2 weeks agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#Multilingual
SillyTavern is a locally installed LLM frontend for AI hobbyists who want detailed control over character chats and prompts. It builds on TavernAI as an independently developed fork and brings text models, image generation and voice into one interface. It's open source under AGPL-3.0.
12.4KUpdated 24 hours ago
iOS · Android · Docker · Web#Batch processing
Inbox Zero is an AI email assistant you can self-host with Docker or use as a hosted service. It's for people handling busy Gmail, Google Workspace, or Microsoft Outlook inboxes who want help sorting mail and preparing replies. It works alongside your existing email client.
6.2KUpdated 2 weeks agoApache-2.0
Web#LLM tracing#Ollama integration#Prompt versioning
Helicone combines an AI gateway with LLM observability for engineers building agents, chatbots, and document processing apps. You can self-host the open source observability platform under Apache 2.0 or use the hosted service. It also integrates with Ollama for apps that run models locally. Its hosted gateway routes requests to external AI providers, so those requests leave your machine.
2.9KUpdated 1 day agoMIT
Web · VS Code#Code execution#Hugging Face integration#MCP
Inspect AI is a Python framework for researchers and developers testing language models and AI agents. Developed by the UK AI Security Institute and Meridian Labs, it evaluates coding, reasoning, knowledge, behavior and multimodal understanding, including tasks where agents must take actions to succeed.
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
19.1KUpdated 4 months ago
macOS · Windows · Linux · Web#Hugging Face integration
LivePortrait animates a still portrait using facial expressions and head movements from a driving video. It's for creators who want to animate faces or edit motion in portrait videos on their own hardware. It works with realistic photos as well as portraits in oil paintings, sculptures and 3D renders.
23.2KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
41.9KUpdated 6 days agoGPL-3.0
macOS · Windows · Linux · iOS · Android · Web#MCP#Multimodal input#Ollama integration
Chatbox is an AI chat client for people who want local models and cloud providers in the same app. It connects to Ollama for local LLM use and supports GPT, Claude, Gemini, Grok and DeepSeek with your own API keys. Chatbox also offers its own hosted model service.
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
18.3KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Human approval#Hybrid search#llama.cpp backend
DocsGPT is an MIT-licensed, open-source platform for teams that want AI search, assistants and agents over their own documents. It can run on your servers with local models, including fully air-gapped deployments where documents and questions stay inside your network. Answers include the source title and page number so readers can check the evidence.
37.3KUpdated 1 day agoAGPL-3.0
macOS · Windows · Docker · Web#Batch processing#Hugging Face integration#MCP
PDFMathTranslate translates scientific PDFs while keeping their page layout, formulas, charts, contents pages and annotations. It's for researchers, students and others who need to read papers in another language without losing the relationship between the text and its figures. It produces both translated PDFs and bilingual documents for comparison with the original.
91Updated 1 day agoAGPL-3.0
Web#Multi-user access#Multilingual#Multimodal input
Nextcloud Assistant brings AI into Nextcloud Hub's documents, email, chat and calendar. It's for teams that want help with shared work while choosing where AI processing happens. With an on-premises model, data stays on your server. The app is open source under AGPL-3.0.
13.7KUpdated 11 months agoMIT
Linux · Web#Hugging Face integration#Multimodal input
TRELLIS is a local AI model for generating 3D assets from images or text prompts, aimed at 3D artists and researchers exploring asset creation. It can produce meshes, radiance fields and 3D Gaussians from the same underlying representation, so you can choose an output suited to your rendering or editing work.