29.6KUpdated 1 week agoApache-2.0
#Code execution#Hugging Face integration#MCP
smolagents is an open-source Python library for developers building AI agents that carry out tasks by writing and executing Python. Its CodeAgent can combine tool calls with loops, conditionals, and calculations in one action. This approach suits tasks that need several operations, rather than a single model response.
1.3KUpdated 2 days agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
26.8KUpdated 2 months agoApache-2.0
Web#Code execution#Git integration#Multimodal input
Onlook is a visual code editor for designers who want to work directly on a website rather than hand off a separate mockup. It uses the actual app code as the design source, so visual edits change the product itself. The Apache 2.0 project can run locally or be self-hosted; Onlook also offers a hosted cloud product.
4.4KUpdated 1 day agoMIT
#Human approval#Multi-agent workflows#Multimodal input
RubyLLM is an MIT-licensed AI framework for developers building Ruby and Rails applications with local or hosted models. Its shared API lets an application switch between Ollama, cloud providers such as Anthropic and OpenAI, and OpenAI-compatible endpoints without rewriting its model integration. The framework runs in your application; model processing happens at the local or hosted backend you choose.
7.1KUpdated 1 day agoMIT
Docker · Web#LM Studio integration#Multimodal input#Ollama integration
big-AGI is a browser-based AI workspace for researchers, developers and people who want to compare model answers before relying on them. Its Beam feature sends the same prompt to several models in parallel, without letting them see each other's replies. You can examine disagreements, then use Merge to combine the responses into a single answer with a preset or custom prompt.
1.4KUpdated 2 months agoMIT
#Multimodal input#Ollama integration#Streaming inference
OllamaSharp is a C# library for developers building .NET applications around Ollama. It connects to Ollama on your own machine or a remote server and covers the full Ollama API, including model management alongside chat and embeddings. It's open source under the MIT license.
22.9KUpdated 3 days agoGPL-3.0
Docker · Web#MCP#Multimodal input#RAG
MaxKB is a self-hosted AI agent platform for organizations building customer support bots, internal knowledge assistants and business automation. It combines answers grounded in company documents with workflows that can call functions and MCP tools. You can run it on your own server through Docker and use it through a browser.
19.9KUpdated 8 months agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#Git integration#LM Studio integration
Bolt.diy is a self-hosted AI coding assistant for people building full-stack Node.js web apps who want to choose their own model provider. It's a fork of Bolt.new, with a browser workspace and an Electron desktop app for macOS, Windows, and Linux. You can also run it through Docker.
6.9KUpdated 22 hours agoApache-2.0
#Agent Client Protocol#Human approval#MCP
codecompanion.nvim brings model chat, inline code edits, and coding agents into Neovim. It's for developers who want AI assistance inside their editor, with a choice between local LLMs through Ollama and cloud providers such as Anthropic, OpenAI, and Google Gemini. The plugin runs in Neovim; where model processing happens depends on the backend you connect.
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
6KUpdated 3 months agoApache-2.0
macOS · iOS#Multimodal input#Ollama integration#Works offline
Enchanted is a ChatGPT alternative for macOS, iOS and visionOS that connects to models on your own Ollama server. It's for people who want an Apple app for chatting with privately hosted models such as Llama 2, Mistral, Vicuna and Starling. You supply the model server. The app is open source under the Apache License 2.0.
3.5KUpdated 6 days agoGPL-3.0
#Git integration#Human approval#llama.cpp backend
gptel brings LLM conversations into the Emacs buffers where you already write and work. It's for Emacs users who want AI help alongside their text or code, with a choice of local model servers and cloud providers. The client is open source under GPL-3.0.
8.2KUpdated 21 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
9.3KUpdated 2 months agoApache-2.0
#Multilingual#Multimodal input#RAG
PaperQA2 is an open source Python research assistant for people who need answers grounded in a collection of scientific papers. It searches documents on your machine and writes answers with in-text citations, including page references. Researchers can use it to summarize findings or check for contradictions across papers, while developers can build it into their own research tools.
10.9KUpdated 1 day agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.
5.1KUpdated 22 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
6.6KUpdated 5 months agoApache-2.0
Docker · Web#Hugging Face integration#Multilingual#Multimodal input
Podcastfy turns documents, websites and images into AI-generated audio conversations, with the option to write transcripts using a local LLM. It's an open-source Python alternative to NotebookLM's podcast feature for creators, educators and researchers who want control over the conversation format or need podcast generation inside their own software. It uses the Apache License 2.0.
12.6KUpdated 3 months ago
Web#Multimodal input#Quantization
HunyuanVideo is an AI video generation model for creators and developers who want to generate footage on their own hardware. Tencent provides model weights and inference code for text-to-video and image-to-video generation, alongside a hosted web experience. Local inference runs on your GPUs; the web offering runs through Tencent's service.
2.3KUpdated 7 months agoMIT
Linux · Web#Multimodal input
MMAudio generates audio that matches a video's action and timing, with text prompts available to guide the result. It also creates audio from text alone. It's for video creators who need sound for silent footage and researchers working on audio generation.
82.9KUpdated 4 hours ago
Linux · Docker · Web#MCP#Multi-agent workflows#Multi-user access
LobeHub is an AI agent workspace for people who want to assign work to several assistants and keep that work organized. It offers a hosted service and a Docker-based self-hosted version for a private device or server. A Linux download is also available.
11KUpdated 9 months agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
LTX-Video is an AI video generation model for creators building controlled animations and developers adding video tools to their own products. You can run it locally or on your own servers using publicly available weights. The LTX family also offers a managed cloud API; local deployments can run in isolated environments without a cloud dependency.
19.1KUpdated 1 week agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
Sentence Transformers is an open-source Python library for developers building semantic search and document retrieval on their own hardware. It runs embedding and reranker models locally, turning content into numerical representations for similarity comparisons and scoring results against a query. The library uses the Apache 2.0 license.
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
2.9KUpdated 6 months agoMIT
macOS · Docker#Batch processing#Hugging Face integration#Multimodal input
Infinity Embeddings is a self-hosted server for developers building semantic search and retrieval-augmented generation applications. It runs embedding and reranking models on your own hardware, with support for image and audio search alongside text. It's open source under MIT.
36.9KUpdated 4 weeks agoMIT
Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API
Vane (formerly Perplexica) is a self-hosted AI search engine for people who want answers drawn from web results, with citations they can check. It runs on your own hardware through Docker or as a server application, with a browser interface and locally stored search history. The project is free and open source under the MIT license.
24.8KUpdated 22 hours agoApache-2.0
iOS · Android#Agent Skills#Hugging Face integration#Multilingual
Google AI Edge Gallery is an open-source app for people who want to try generative AI on their own phone. It runs model inference locally, so offline chat, image analysis and audio tasks don't send your inputs to a server. It supports Android and iOS and uses the Apache 2.0 license.
922Updated 5 months agoMIT
macOS · Windows · Linux · Docker · Web#llama.cpp backend#LM Studio integration#MLX
Eclaire is a self-hosted AI assistant for notes, documents, photos, bookmarks and tasks. You can ask questions about saved material, inspect the sources behind an answer and ask the assistant to create notes or update tasks. Scheduled automations handle recurring requests such as a weekly task summary.