29.6KUpdated 1 week agoApache-2.0
#Code execution#Hugging Face integration#MCP
smolagents is an open-source Python library for developers building AI agents that carry out tasks by writing and executing Python. Its CodeAgent can combine tool calls with loops, conditionals, and calculations in one action. This approach suits tasks that need several operations, rather than a single model response.
1.3KUpdated 2 days agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
4.4KUpdated 1 day agoMIT
#Human approval#Multi-agent workflows#Multimodal input
RubyLLM is an MIT-licensed AI framework for developers building Ruby and Rails applications with local or hosted models. Its shared API lets an application switch between Ollama, cloud providers such as Anthropic and OpenAI, and OpenAI-compatible endpoints without rewriting its model integration. The framework runs in your application; model processing happens at the local or hosted backend you choose.
19.7KUpdated 6 months agoApache-2.0
Linux · Docker · Web#Batch processing#Distributed execution#OpenAI-compatible API
olmOCR is an open-source OCR toolkit for turning PDFs and image documents into text for LLM datasets and training. It suits researchers and developers who need readable document content, including pages where columns, figures, or mathematical notation make text extraction difficult. You can run it on your own GPU, including through Docker, or use a remote inference server.
1.5KUpdated 4 months agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API
Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.
19.9KUpdated 8 months agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#Git integration#LM Studio integration
Bolt.diy is a self-hosted AI coding assistant for people building full-stack Node.js web apps who want to choose their own model provider. It's a fork of Bolt.new, with a browser workspace and an Electron desktop app for macOS, Windows, and Linux. You can also run it through Docker.
3.5KUpdated 6 days agoGPL-3.0
#Git integration#Human approval#llama.cpp backend
gptel brings LLM conversations into the Emacs buffers where you already write and work. It's for Emacs users who want AI help alongside their text or code, with a choice of local model servers and cloud providers. The client is open source under GPL-3.0.
8.2KUpdated 21 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
5.1KUpdated 22 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.
6KUpdated 23 hours agoApache-2.0
#Hugging Face integration#ONNX#OpenAI-compatible API
KServe is an open source platform for teams serving LLMs and predictive machine learning models on their own Kubernetes infrastructure. It puts both kinds of workloads under a common serving API, so teams can manage different model frameworks through the same platform. It uses the Apache 2.0 license.
19.2KUpdated 2 weeks agoApache-2.0
Web · Browser Extension#OpenAI-compatible API#Streaming inference#Structured output
WebLLM runs language models directly in a user's browser, using WebGPU for GPU acceleration. It's an open-source engine for developers building web-based AI assistants and Chrome extensions that process prompts on the user's device rather than an inference server. The project uses the Apache 2.0 license.
26.6KUpdated 1 day agoApache-2.0
Docker#Guardrails#Hugging Face integration#Hybrid search
Haystack is a Python framework for developers building self-hosted AI agents, document search, and apps that answer questions using their own data. Its modular pipelines let teams control which information reaches a model and inspect how retrieval, memory, tools, and generation contribute to an answer. It's open source under Apache 2.0.
2.9KUpdated 6 months agoMIT
macOS · Docker#Batch processing#Hugging Face integration#Multimodal input
Infinity Embeddings is a self-hosted server for developers building semantic search and retrieval-augmented generation applications. It runs embedding and reranking models on your own hardware, with support for image and audio search alongside text. It's open source under MIT.
36.9KUpdated 4 weeks agoMIT
Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API
Vane (formerly Perplexica) is a self-hosted AI search engine for people who want answers drawn from web results, with citations they can check. It runs on your own hardware through Docker or as a server application, with a browser interface and locally stored search history. The project is free and open source under the MIT license.
655Updated 2 days agoApache-2.0
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.
796Updated 2 weeks agoMIT
macOS · Windows · Linux#llama.cpp backend#LM Studio integration#Multilingual
Interpreter reads Japanese text from game windows and displays English translations over your screen. It's built for people playing Japanese retro games, with MeikiOCR tuned for game text and pixel fonts. The default OCR and Sugoi V4 translation models run locally, so text stays on your computer and no internet connection is needed after setup. It's open source under the MIT license, with support for Windows, macOS and Linux.
46Updated 5 days agoAGPL-3.0
macOS · Windows · Linux · Browser Extension#Code execution#GGUF#llama.cpp backend
Vyact is an open source desktop workspace for people who want to use a local LLM with their documents, email and code. It runs on Apple Silicon Macs, Windows and Linux x64 under the AGPL-3.0 license. Intel Macs aren't supported.
922Updated 5 months agoMIT
macOS · Windows · Linux · Docker · Web#llama.cpp backend#LM Studio integration#MLX
Eclaire is a self-hosted AI assistant for notes, documents, photos, bookmarks and tasks. You can ask questions about saved material, inspect the sources behind an answer and ask the assistant to create notes or update tasks. Scheduled automations handle recurring requests such as a weekly task summary.