39.6KUpdated 1 year ago
#Ollama integration#RAG#Reranking
Quivr Core is a Python framework for developers adding document-based AI answers to their own applications. It combines file ingestion with retrieval-augmented generation (RAG), so a model can answer questions using material from your documents. It supports local models through Ollama as well as cloud APIs from OpenAI, Anthropic, and Mistral.
5.4KUpdated 3 months ago
macOS · Windows · Linux#MCP#Multilingual#Ollama integration
5ire is a free desktop AI assistant that combines chat, a local document knowledge base and MCP tools. It runs on macOS, Windows and Linux, with Mac downloads for Apple Silicon and Intel. It's for people who want to use their own documents and external tools alongside conversations with local or cloud models.
61.2KUpdated 6 months agoCC-BY-4.0
Web#Code execution#MCP#Multi-agent workflows
AutoGen is a framework for developers building AI agents that work together or alongside people. Its agent runtime can run locally or across distributed systems, while model integrations such as OpenAI and Azure OpenAI send requests to external services. It's community-managed and in maintenance mode, with no further features or enhancements planned.
17.7KUpdated 2 years agoMIT
Docker · Web#Human approval#Persistent memory#Tool calling
SuperAGI is a self-hosted AI agent framework for developers who want to build task automation and manage agents through a browser interface. It runs locally in Docker and can use local LLMs with a GPU. The Python project uses the MIT license, so you can modify it for your own applications.
8.4KUpdated 21 hours agoMIT
#Agent Skills#Batch processing#Guardrails
OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.
layla-network.aiAI Characters and Roleplay
iOS · Android#Code execution#GGUF#llama.cpp backend
Layla is an offline AI assistant for Android and iOS that runs language models on your phone. It's for people who want private everyday chat or an AI companion with custom characters, memory and roleplay. Local conversations stay encrypted on your device; optional cloud mode sends requests to your chosen hosted provider.
4.8KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
Lollms WebUI is a local, single-user AI interface for people who want text chat and media generation in one place. It runs on Windows, macOS and Linux, with Docker support, and lets writers, developers and other users choose models and task-specific personalities. It's free and open source under Apache 2.0. The project receives minimal maintenance.
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
huggingface.coOCR and Document Scanning
Linux#Batch processing#Hugging Face integration#Multimodal input
Qwen2.5-VL is a vision-language model you can run on your own hardware to answer questions about images and video. It's aimed at developers building document processing tools, visual assistants and agents that interact with computer or phone screens. The instruction-tuned 7B model has Apache 2.0 licensing and works with Hugging Face Transformers, with weights available in Safetensors format.
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Multimodal input#Tool calling
Pixtral is a family of Mistral models for developers who want to run multimodal AI on their own hardware. The associated mistral-inference project is archived and no longer maintained. That status applies to the inference library.
20.1KUpdated 2 days agoMIT
macOS · Linux · Docker · Web#Code execution#llama.cpp backend#OpenAI-compatible API
DB-GPT is a self-hosted AI data assistant for teams analyzing business data and developers building data applications. It turns plain-language requests into SQL queries and Python analysis, then produces charts, dashboards, or HTML reports. You can run it on macOS or Linux, with Docker deployment also supported.
3.1KUpdated 2 days agoApache-2.0
macOS · Windows · Linux · Web#Code execution#MCP#Multi-agent workflows
BotSharp is a self-hosted framework for .NET developers building AI agents into business applications. Written in C#, it runs on Windows, Linux and macOS and is open source software under Apache 2.0. Its plugin design lets teams choose their model provider, storage and interface while keeping agent coordination in the same framework.
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Tool calling
Codestral and Devstral are Mistral coding models with downloadable weights for local or self-hosted development tools. Codestral focuses on code generation and fill-in-the-middle completion, where the model fills a gap between existing code. Devstral is designed for software engineering agents that explore a repository, use tools and edit multiple files.
huggingface.coOpen-Weight LLMs
#Guardrails#Hugging Face integration#Multilingual
Command A is an open weights language model from Cohere and Cohere Labs for researchers and developers building self-hosted chatbots, document assistants and AI agents. Its focus is business tasks that combine multilingual text, supplied documents and external tools. You can run the model on your own hardware; Cohere also offers hosted chat through a playground and Hugging Face Space.
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Multimodal input#Tool calling
Mistral Small and Large are downloadable language models for developers building chat, reasoning and tool-using applications on their own infrastructure. Capabilities and hardware requirements depend on the release. Mistral Small 3.1 adds image understanding to text generation, while Mistral Large 2 is a larger text model.
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
9.7KUpdated 9 months agoMIT
#Ollama integration#RAG#Semantic search
LangChainGo is a Go implementation of LangChain for developers building LLM applications in their own software. It connects Go programs to model backends, including Ollama for local LLM use and cloud services such as OpenAI and Gemini. It's a library, so its audience is developers who want to build an application rather than use a ready-made chat interface.
88.8KUpdated 2 months agoMIT
macOS · Windows · Linux · Docker · Web#MCP#Multimodal input#OpenAI-compatible API
NextChat is a self-hosted AI chat interface for people who want one place to use their own LLM server and cloud models. The web and desktop project is open source under the MIT license. You can host it with Docker or on Vercel, and desktop clients run on macOS, Windows and Linux.
11.1KUpdated 11 months ago
#Hugging Face integration#Quantization#Tool calling
Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.
820Updated 1 year ago
#Quantization#Tool calling
Hunyuan-A13B is Tencent's downloadable language model for developers and researchers building self-hosted AI applications. It supports general text tasks, reasoning and agent workloads, with a choice between quick responses and more deliberate reasoning. It's part of the broader Hunyuan model family, which Tencent also offers through its website.
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
3.2KUpdated 1 year agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
MiniMax-M1 is an open-source reasoning model for developers building agents or working on complex software and mathematical problems. Its million-token context window makes it a candidate for tasks with long inputs that also need extended reasoning. You can serve the model on your own infrastructure through vLLM or use it through Transformers.
7.3KUpdated 1 year agoApache-2.0
#Hugging Face integration#Tool calling#Web search
InternLM is a family of downloadable language models for developers and researchers building their own AI applications. It includes models for general conversation, complex reasoning and coding, with separate base and chat variants for customization or use in an assistant.
10.9KUpdated 6 months agoApache-2.0
Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.
1.1KUpdated 7 months agoAGPL-3.0
Web · Browser Extension#Multilingual#Multimodal input#Ollama integration
NativeMind brings a local AI assistant into Chrome for people who want help with webpages, documents, and writing without sending that content to a cloud model. The browser extension connects to Ollama on your machine, keeping prompts and AI processing on-device. It requires no account and uses the AGPL-3.0 open-source license.
3.5KUpdated 20 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
5.7KUpdated 2 weeks agoMIT
macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows
GLaDOS is a local AI voice assistant modeled on the sarcastic character from Valve's Portal games. It's for people who want a conversational companion on their own hardware, with camera awareness and connections to home automation. The Python project is open source under the MIT license and runs on Linux and Windows. macOS support is experimental.
5.1KUpdated 22 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
11KUpdated 1 day agoApache-2.0
Docker · Web#llama.cpp backend#MCP#Multi-user access
HuggingChat UI is the open-source chat application behind Hugging Face's hosted HuggingChat. You can run it on your own computer or server and connect it to a local LLM backend or a cloud provider. It's for people and teams who want a browser-based ChatGPT alternative with control over the chat service and where its data lives. The code uses the Apache 2.0 license.
37.6KUpdated 20 hours agoMIT
iOS · Android · Web#Human approval#MCP#Persistent memory
CopilotKit is a self-hostable SDK for developers building AI agents into web and mobile apps, Slack, or Microsoft Teams. Agents can display interactive charts and forms using an app's own components, read shared app state, and take actions through frontend tools, APIs, or MCP tools.
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
chatwise.appChat With Your Documents
#MCP#Multimodal input#Tool calling
ChatWise is a desktop chat app for connecting different model providers through one interface. It stores data locally; model processing follows the provider you choose. Its official documentation includes Ollama for models running on your own machine, alongside cloud providers connected with your API keys.
2.7KUpdated 3 months agoMIT
Docker · Web#MCP#Multi-user access#Single sign-on
MetaMCP is a self-hosted gateway for developers and teams who want to give AI clients access to several MCP servers through one endpoint. It runs on your own machine or server with Docker and is open source under the MIT license. You choose which tools clients see.