10.8KUpdated 2 days agoMIT
Browser Extension#Multilingual#Ollama integration#OpenAI-compatible API
ChatGPTBox is a free, open-source browser extension for people who want AI help with the pages they read, search results and selected text. It works in Chrome, Edge, Firefox and Safari, with mobile support too. You can connect it to Ollama or a self-hosted model through its custom model mode, or use cloud services such as ChatGPT, Claude, Moonshot and Azure.
12KUpdated 12 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access
h2oGPT is a self-hosted ChatGPT alternative for people who want to chat with local models and ask questions about their own documents. The project is archived and no longer maintained. It's open source under Apache 2.0, with support for Linux, macOS, Windows and Docker.
8.5KUpdated 1 year agoAGPL-3.0
macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#RAG
Reor is a local AI note-taking app for people who want to search their own writing, find connections between ideas and ask questions about their notes. The project is archived and no longer maintained. It runs on macOS, Windows and Linux, and it's open source under AGPL-3.0.
14.2KUpdated 1 week agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Multilingual#Ollama integration#OpenAI-compatible API
QAnything is a self-hosted knowledge base for people and teams who want to ask questions about their own documents, including collections that mix Chinese and English. It can answer in either language regardless of the document's language, and runs locally through Docker on Windows, macOS and Linux.
857Updated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API
OpenedAI Speech is a self-hosted text-to-speech server for developers who want local speech generation in apps built around OpenAI's speech API. The project is archived and no longer maintained. It's open source under AGPL-3.0, and it generates audio on your own hardware without an OpenAI API key.
7.7KUpdated 4 months agoBSD-3-Clause
Windows · Docker · Web#Hugging Face integration#Hybrid search#Ollama integration
Verba is a self-hosted document chatbot for people who want to ask questions across their files and knowledge bases. The project is archived and no longer maintained. It uses Weaviate to find relevant passages and gives those passages to a language model to generate answers.
getzep/zepAgent Memory
Docker#Hybrid search#Knowledge graphs#OpenAI-compatible API
Zep Community Edition v1.0.2 is a deprecated, unsupported self-hosted memory service for developers building AI agents and conversational assistants. This legacy edition uses the Apache 2.0 license. It turns chat history into a knowledge graph that records how facts change over time, so an assistant can distinguish a user's current preferences from earlier ones.
9.2KUpdated 5 days agoApache-2.0
Docker · Web#Multi-user access#Ollama integration#OpenAI-compatible API
Morphic is a self-hosted AI search engine for people who want answers backed by web sources and control over the search interface they use. It combines web searches and URL reading with AI-generated responses, so you can examine the sources behind an answer. It's open source under Apache 2.0, and you can run your own instance with Docker or use the hosted website.
22.6KUpdated 3 months agoApache-2.0
macOS · Docker · Web#Ollama integration#OpenAI-compatible API
OpenUI is a self-hosted AI UI builder for developers prototyping web interfaces with models they choose. It turns plain-language descriptions into rendered interfaces, with a live preview and follow-up requests for changes. You can run the app locally through Docker or Python and use it in your browser.
33.7KUpdated 4 months ago
Windows · Docker#Code execution#Human approval#Multi-agent workflows
GPT Pilot is a self-hosted AI coding assistant for developers who want an agent to build applications under their supervision. The project is no longer maintained. It uses the FSL-1.1-MIT source-available license, with restrictions on competing uses. It runs locally as a Python CLI and is the core technology behind the Pythagora VS Code extension.
3.4KUpdated 1 day agoApache-2.0
#Batch processing#Distributed execution#Multilingual
DataTrove is an open-source Python library for teams preparing large text datasets, including LLM training corpora. It runs on your own machine or on Slurm and Ray clusters, with processing steps that carry across those environments. It uses the Apache 2.0 license.
3.7KUpdated 1 day agoMIT
Docker · Web · VS Code#Git integration#Hybrid search#llama.cpp backend
Twinny is an AI coding assistant for VS Code that lets developers choose where their models run: on their own computer, a private server or a hosted API. It's for individuals and teams who want code suggestions and repository chat with control over where their code goes. The extension and team gateway are open source under the MIT license.
241Updated 2 years agoAGPL-3.0
Docker#Multi-user access#OpenAI-compatible API
Matrix ChatGPT Bot connects Matrix rooms to OpenAI's ChatGPT API for people who want AI conversations in their existing chat client, including Element. The project is archived and no longer maintained. It's open source under AGPL-3.0, and the maintainers point users to Baibot as an alternative.
14KUpdated 5 months ago
macOS · Windows · Linux · Web#GGUF#LM Studio integration#MCP
Open-LLM-VTuber is a local AI companion for people who want a character they can talk to, with a Live2D avatar that responds through speech and expressions. It runs on Windows, macOS and Linux through web and desktop clients. With local models for speech and language processing, it works fully offline and keeps conversations on your device. Cloud APIs are optional alternatives that send the corresponding processing to external services.
1.6KUpdated 1 year agoMIT
Windows · Docker · Web#llama.cpp backend#LM Studio integration#Multimodal input
Amica is a locally runnable interface for talking with customizable 3D AI characters. It's for people who want an animated, voiced character as the face of their AI assistant, with a choice of local LLM backends or cloud services. The project builds on Pixiv's ChatVRM.
752Updated 9 months agoAGPL-3.0
Web#llama.cpp backend#OpenAI-compatible API#Persistent memory
Mikupad is a browser-based LLM frontend packaged as a single HTML file, for writers and people who want control over generated text. It connects to a model backend of your choice and supports both direct text continuation and chat with instruct models. It's open source under AGPL-3.0.
9.4KUpdated 1 day agoMIT
macOS · Windows · Linux · iOS · Android#LM Studio integration#MCP#Ollama integration
Anarlog, formerly Hyprnote, is a desktop AI meeting notetaker for people who want to keep private conversations on their own hardware. It captures audio from your device without adding a bot to the call and stays hidden during screen sharing. The app runs on macOS, Windows and Linux; its community application is open source under the MIT license.
1.1KUpdated 2 years agoApache-2.0
#OpenAI-compatible API
LLMPerf is a Python library for developers comparing LLM API performance and checking basic response correctness. The project is archived and no longer maintained. Its open-source code uses the Apache License 2.0.
10.8KUpdated 3 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.
6KUpdated 6 months agoMIT
Docker · Web#Multilingual#Ollama integration#OpenAI-compatible API
Paperless-AI is a self-hosted extension for Paperless-ngx users who want automatic document sorting and chat with their archive. It requires an existing Paperless-ngx instance and runs in Docker, with a browser interface for reviewing and processing documents. The project is no longer maintained.
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
mindmac.appDesktop Chat Apps
macOS#llama.cpp backend#LM Studio integration#MLX
MindMac is a native macOS AI chat client for people who want to use local LLMs and cloud services in the same app. Its inline mode lets you ask questions or generate text inside Notes, Mail and browsers without switching to a chat window. It runs on Intel and Apple Silicon Macs with macOS 13 or newer.
4.5KUpdated 7 months agoMIT
macOS · Windows · Linux#MCP#Ollama integration#OpenAI-compatible API
mods is a command-line AI tool for people who want to ask questions about command output or use model responses in shell pipelines. It connects to local LLMs through LocalAI as well as cloud services. The project is archived and no longer maintained.
8.4KUpdated 21 hours agoMIT
#Agent Skills#Batch processing#Guardrails
OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.
10.9KUpdated 3 years agoMIT
macOS · Docker · Web#GGUF#llama.cpp backend#OpenAI-compatible API
LlamaGPT is a self-hosted ChatGPT alternative for people who want general chat or coding help on their own computer or home server. It runs models locally and keeps conversation data on your device. After the initial model download, it works offline.
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
20.1KUpdated 2 days agoMIT
macOS · Linux · Docker · Web#Code execution#llama.cpp backend#OpenAI-compatible API
DB-GPT is a self-hosted AI data assistant for teams analyzing business data and developers building data applications. It turns plain-language requests into SQL queries and Python analysis, then produces charts, dashboards, or HTML reports. You can run it on macOS or Linux, with Docker deployment also supported.
14.1KUpdated 2 weeks agoMIT
macOS#Batch processing#GGUF#Hugging Face integration
lm-evaluation-harness lets researchers and model developers compare language models using shared academic benchmarks and public prompts. It runs evaluations against local models and benchmarks, or sends requests to a hosted model API. EleutherAI's Python framework is open source under the MIT license and powers Hugging Face's Open LLM Leaderboard.
7.1KUpdated 2 days agoApache-2.0
Docker#Guardrails#LLM tracing#Multi-agent workflows
Plano, formerly Arch Gateway, is a self-hosted AI gateway for developers building applications with multiple agents or model providers. It puts routing, guardrails and request tracing in a separate service, so each agent doesn't need its own implementation of that infrastructure. It's open source under Apache 2.0.
13.1KUpdated 4 months agoMIT
Docker · Web#Guardrails#Ollama integration#OpenAI-compatible API
Portkey Gateway is a self-hosted AI gateway for developers whose apps need to use local models and cloud providers through one OpenAI-compatible API. It routes requests to Ollama, OpenAI, Anthropic, Google Gemini and other backends, with controls for handling failures and checking model inputs and outputs.
88.8KUpdated 2 months agoMIT
macOS · Windows · Linux · Docker · Web#MCP#Multimodal input#OpenAI-compatible API
NextChat is a self-hosted AI chat interface for people who want one place to use their own LLM server and cloud models. The web and desktop project is open source under the MIT license. You can host it with Docker or on Vercel, and desktop clients run on macOS, Windows and Linux.
127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
1.3KUpdated 7 months agoApache-2.0
Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API
JoyCaption is an open-weight image captioning model for people preparing datasets to train or fine-tune diffusion models. It runs on your own GPU and covers both SFW and NSFW images, including photography, anime, digital art and furry artwork. Automated captions reduce the need to write descriptions by hand or find images that already have usable text.
10.9KUpdated 6 months agoApache-2.0
Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.
5.7KUpdated 2 weeks agoMIT
macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows
GLaDOS is a local AI voice assistant modeled on the sarcastic character from Valve's Portal games. It's for people who want a conversational companion on their own hardware, with camera awareness and connections to home automation. The Python project is open source under the MIT license and runs on Linux and Windows. macOS support is experimental.