68.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux#Agent Client Protocol#Code execution#Human approval
Open Interpreter is a terminal coding assistant built around open-weight models, with model-specific behavior for Kimi, Qwen and DeepSeek. It's a fork of OpenAI's Codex for developers who want to choose their model provider while keeping a familiar agent interface. The project uses the Apache 2.0 license.
10.6KUpdated 2 years agoMIT
macOS · Windows · Linux · Docker#Hugging Face integration
Petals lets developers and researchers use large language models that won't fit on a single consumer GPU by sharing the work across a network of machines. It supports text generation and fine-tuning from a desktop computer or Google Colab. Each participant holds part of the model, while other computers handle the remaining parts.
18.5KUpdated 1 day agoApache-2.0
Linux · Docker#Batch processing#Hugging Face integration
Parakeet is NVIDIA's speech recognition model family. The linked parakeet-tdt-0.6b-v2 is its English speech-to-text model for developers and researchers building transcription services, subtitles or voice applications. It runs locally through NeMo on Linux, with NVIDIA GPUs recommended for inference. It's a model you can embed in an application, rather than a desktop transcription app.
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
49.1KUpdated 2 days agoAGPL-3.0
Linux · Docker · Web#Multi-user access#Multimodal input#OpenAI-compatible API
New API is a self-hosted AI gateway for developers and teams that want several model providers behind one service. It builds on One API and converts between OpenAI Chat Completions, Responses, Anthropic Messages and Gemini formats, so apps and agents can switch providers without changing each client's connection settings.
5.8KUpdated 20 hours agoBSD-3-Clause
Linux#Distributed execution#Hugging Face integration
torchtitan is an open-source training platform for researchers and developers building generative AI models on their own GPU machines or server clusters. It uses PyTorch's distributed training tools and keeps the model code relatively simple when spreading work across GPUs. The Python codebase has extension points and replaceable components for experiments with model architectures and training infrastructure.
1.9KUpdated 12 months agoGPL-3.0
Linux#Works offline
nerd-dictation is an open-source dictation utility for desktop Linux that recognizes speech locally through VOSK. It's for people who want voice input in their existing applications and are comfortable with a command-line tool. Audio processing stays on your machine, and recognition works offline.
38.7KUpdated 11 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API
Langchain-Chatchat is a self-hosted application for asking questions about your own documents and using AI agents. It focuses on Chinese-language use and open models, with a fully offline setup that can keep documents and model processing on your hardware. Its code is open source under Apache 2.0.
1.4KUpdated 2 days agoAGPL-3.0
Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.
3.1KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization
Distributed Llama runs a local LLM across several computers, sharing both the computation and the model's memory use. It's for people who want to use their own networked hardware for inference rather than keep the entire workload on one machine. The C++ project is open source under the MIT license.
17.5KUpdated 1 day agoMIT
macOS · Windows · Linux · Web#Persistent memory#Tool calling
Leon is a self-hosted personal AI assistant for Linux, macOS and Windows. The project focuses on the 2.0 Developer Preview on its develop branch. Its MIT-licensed runtime combines tools, context and memory with a local web interface.
1.7KUpdated 2 days ago
Linux#Multimodal input#Quantization
RKLLM is a software stack for developers building local AI applications on Rockchip hardware. It uses the chip's neural processing unit (NPU) to run language and multimodal models on development boards, with support for the RK3588, RK3576, RK3562 and RV1126B series.
22.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Hybrid search#Ollama integration
localGPT is a self-hosted AI document chat app for people who want to question and summarise files on their own hardware. Its local Ollama setup keeps documents and conversations on your machine. Answers include source passages, so you can check what the model used.
1.8KUpdated 2 weeks agoMIT
macOS · Windows · Linux#Human approval#Tool calling
OpenAdapt turns a demonstrated task into a repeatable program for browser, desktop, and remote applications. Its local, MIT-licensed open-source engine is for teams and AI agent builders who need to automate work in interfaces their APIs can't reach. It checks the outcome independently before reporting success and stops when verification fails.
12.4KUpdated 4 weeks agoApache-2.0
macOS · Windows · Linux#Code execution#Multimodal input#Tool calling
Agent S is an open-source AI agent framework that controls ordinary desktop and web applications through the screen, mouse and keyboard. It runs on macOS, Windows and Linux and is aimed at developers, computer-use researchers and people building automation for their own desktops. You describe the task in natural language; the agent clicks, types and scrolls without requiring an API integration or a separate script for each application.
10.5KUpdated 3 weeks ago
macOS · Windows · Linux · Web#Batch processing#ControlNet#Image-to-image
Easy Diffusion runs Stable Diffusion on your own computer through a browser interface. It's for people who want to generate and edit images locally without assembling the software components themselves. The free distribution bundles the required software and works on Windows, Linux and macOS.
3KUpdated 5 days ago
Linux · iOS#Hugging Face integration#LoRA#Quantization
TorchAO is a PyTorch library for developers who want to train or run models on their own hardware with less memory and faster computation. It reduces the precision of model weights and activations, with options for language models and image or video generation. Its PyTorch integration works with torch.compile and FSDP2 across most Hugging Face PyTorch models.
47.7KUpdated 1 month agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#GGUF#llama.cpp backend#LoRA
text-generation-webui, also called TextGen, runs language models on your own hardware through a desktop app or a self-hosted browser interface. It's for people who want private chat and writing tools, and developers who need a local model API. It works offline without telemetry; web search and page fetching use the internet.
5.5KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API
Kokoro-FastAPI runs the Kokoro-82M speech model on your own machine or server and exposes an OpenAI-compatible speech API. It's for developers adding local text-to-speech to assistants, reading apps or audiobook workflows. Speech generation runs locally, and the API doesn't require an OpenAI account.
6.6KUpdated 1 year ago
Windows · Linux · Web#Batch processing#Inpainting#Multilingual
MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.
11KUpdated 1 week agoBSD-3-Clause
Windows · Linux · Docker#Batch processing#ONNX
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.
40.3KUpdated 21 hours ago
macOS · Windows · Linux · Docker · Web
PhotoPrism is a self-hosted photo and video library for people who want to organize personal media on their own hardware or server. Its AI recognizes faces and labels pictures by content and location, so finding a photo doesn't depend entirely on folders or tags you've added yourself. You browse and share the library through a web app.
26.1KUpdated 5 hours ago
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llamafile puts an LLM and the software that runs it into a single executable. It’s for people who want to run models locally or share them with others without asking each recipient to set up a separate runtime. The project is open source under Apache 2.0.
22KUpdated 1 hour agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
165.2KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Web#Batch processing#Code execution#Image-to-image
Stable Diffusion web UI (AUTOMATIC1111) is a browser interface for generating and editing images with models running on your own hardware. It's for artists and anyone who wants control over prompts, models and image variations. The software is open source under AGPL-3.0.
12.6KUpdated 3 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#ControlNet#Hugging Face integration#LoRA
Kohya's GUI lets you train and fine-tune image generation models on your own GPU-equipped computer through a browser interface. It's for artists and model makers who want to teach a model a particular style or subject while controlling the training settings. The interface builds on Kohya's Stable Diffusion training scripts, with a command-line interface available too.
8.9KUpdated 2 weeks agoAGPL-3.0
macOS · Windows · Linux#Hugging Face integration#LoRA
Stability Matrix is an open source desktop app for people who use more than one Stable Diffusion interface. It manages local installations of ComfyUI, Automatic1111, Fooocus, Forge, and InvokeAI, so you can try different workflows without maintaining a separate model collection for each one. It runs on Windows, macOS, and Linux under the AGPL-3.0 license.
11KUpdated 1 week ago
Linux · Web#MCP
MCP Inspector is a locally run developer tool for testing Model Context Protocol (MCP) servers. It's for developers building or checking the servers that connect AI applications to tools and data. You can inspect a local server or connect to a remote HTTP endpoint.
13.2KUpdated 1 day agoMIT
Linux · Docker#Git integration#Ollama integration
PR-Agent is an open source AI code review agent for development teams that want to choose where their reviewer runs and which model it uses. It reviews pull requests, writes descriptions, suggests code improvements, and answers questions about proposed changes. You can run it locally through a CLI, on a self-hosted server, in Docker, or through GitHub Actions.
1.6KUpdated 2 months agoGPL-3.0
Linux#Code execution#Multimodal input#Ollama integration
Alpaca is an open source AI chat client for people who want to run models on their own device. It uses Ollama to download and manage local models, then lets you chat without an internet connection. Conversations stay on your device in SQLite files, and local models have no direct internet access.
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
6.5KUpdated 2 months agoMIT
Linux#Multilingual#Works offline
Argos Translate runs neural machine translation locally, so you can translate text without sending it to an online service. It's for people who need offline translation and developers who want to add it to their own software. The Python library and desktop interface use the same translation engine, and the project is open source under the MIT license.
24.3KUpdated 4 days agoBSD-2-Clause
macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual
WhisperX is an open source speech-to-text tool for people transcribing interviews, meetings, and long recordings on their own computer. It builds on OpenAI's Whisper to produce transcripts with word-level timestamps and optional speaker labels.
28.2KUpdated 2 hours agoApache-2.0
macOS · Windows · Linux · Web · VS Code · JetBrains#Code execution#Git integration#MCP
Qwen Code is an Apache 2.0 licensed AI coding agent for developers who want help working through a codebase, changing code and checking the result. It runs on macOS, Windows and Linux, with a terminal interface and a desktop app. It builds on Google Gemini CLI and has developed into an agent that can use Qwen models alongside other model providers.
25.8KUpdated 4 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Hybrid search#llama.cpp backend#Multi-user access
kotaemon is a self-hosted document chat app for people who want to ask questions across their files and check where the answers came from. It runs in a browser on Windows, macOS or Linux, with Docker also supported. The project uses the Apache 2.0 license.