26.8KUpdated 2 months agoApache-2.0
Web#Code execution#Git integration#Multimodal input
Onlook is a visual code editor for designers who want to work directly on a website rather than hand off a separate mockup. It uses the actual app code as the design source, so visual edits change the product itself. The Apache 2.0 project can run locally or be self-hosted; Onlook also offers a hosted cloud product.
13.2KUpdated 1 day agoApache-2.0
#MCP#Ollama integration#RAG
LangChain4j is an Apache 2.0 open-source Java library for developers building chatbots, assistants and AI agents in JVM applications. It connects application code to local LLM backends such as Ollama as well as cloud providers such as OpenAI and Google Vertex AI. Where model requests go depends on the backend you choose.
4.4KUpdated 1 day agoMIT
#Human approval#Multi-agent workflows#Multimodal input
RubyLLM is an MIT-licensed AI framework for developers building Ruby and Rails applications with local or hosted models. Its shared API lets an application switch between Ollama, cloud providers such as Anthropic and OpenAI, and OpenAI-compatible endpoints without rewriting its model integration. The framework runs in your application; model processing happens at the local or hosted backend you choose.
28.6KUpdated 2 days agoMIT
Docker · Web · Browser Extension#Agent Skills#Code execution#MCP
Repomix turns a code repository into a single file that an AI assistant can read. It’s for developers who need to give Claude, ChatGPT, Gemini, or another model enough project context for code review or analysis. You can use its command line on your own machine, pack a repository through its website, or run it as an MCP server, including through Docker. It is open source under the MIT license.
7.1KUpdated 1 day agoMIT
Docker · Web#LM Studio integration#Multimodal input#Ollama integration
big-AGI is a browser-based AI workspace for researchers, developers and people who want to compare model answers before relying on them. Its Beam feature sends the same prompt to several models in parallel, without letting them see each other's replies. You can examine disagreements, then use Merge to combine the responses into a single answer with a preset or custom prompt.
3.9KUpdated 1 week agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
SmolLM3 is a 3B parameter language model from Hugging Face for developers and researchers who want to run an LLM on their own hardware. It comes as a base model and an instruction-tuned model for chat, reasoning and tool calling. Both run locally.
5.1KUpdated 1 year agoApache-2.0
Web#Multi-user access#Semantic search
Argilla is an open-source data annotation and feedback tool for AI engineers and domain experts who build training and evaluation datasets. You can run your own Argilla server or deploy it on Hugging Face Spaces. It's licensed under Apache 2.0.
19.7KUpdated 6 months agoApache-2.0
Linux · Docker · Web#Batch processing#Distributed execution#OpenAI-compatible API
olmOCR is an open-source OCR toolkit for turning PDFs and image documents into text for LLM datasets and training. It suits researchers and developers who need readable document content, including pages where columns, figures, or mathematical notation make text extraction difficult. You can run it on your own GPU, including through Docker, or use a remote inference server.
1.5KUpdated 4 months agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API
Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.
2.8KUpdated 2 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#Git integration#Human approval
Vexa is a meeting bot and transcription API for developers building meeting features and teams feeding calls into AI agents. Bots join Google Meet, Microsoft Teams and Zoom, then send live transcripts with speaker labels to your application or agent. You can host the full platform on your own infrastructure or use Vexa's cloud service.
21.7KUpdated 1 day agoApache-2.0
macOS#Distributed execution#Hugging Face integration#LoRA
PEFT is an open-source Python library for developers who want to adapt pretrained models on their own hardware with less compute and storage than full fine-tuning requires. It trains a small subset of parameters, often through adapters, while leaving the base model intact. It's licensed under Apache 2.0.
1.4KUpdated 2 months agoMIT
#Multimodal input#Ollama integration#Streaming inference
OllamaSharp is a C# library for developers building .NET applications around Ollama. It connects to Ollama on your own machine or a remote server and covers the full Ollama API, including model management alongside chat and embeddings. It's open source under the MIT license.
19.9KUpdated 2 years agoApache-2.0
Web#Hugging Face integration
Segment Anything 2 (SAM 2) is Meta's open-source model for selecting objects in images and tracking them through video. It's for developers and researchers who need object masks for visual applications or dataset annotation. The model and its web demo can run on your own GPU machine; Meta also provides a hosted demo.
36.9KUpdated 2 years agoBSD-3-Clause
macOS · Windows · Linux#Batch processing
Real-ESRGAN is a local AI upscaler and restoration toolkit for people enlarging images or improving video, with dedicated models for anime illustrations and animation. It builds on ESRGAN and uses models trained entirely on synthetic data to address degraded images. The code is open source under the BSD 3-Clause license.
1KUpdated 1 month agoMIT
macOS
Practical-RIFE is a local AI frame interpolation tool for engineers and developers working with video. It generates intermediate frames to increase frame rates, with models intended for ordinary footage, animation, and post-processing videos made by diffusion models. The Python project builds on RIFE and SAFA, with an emphasis on how the output looks rather than improvements in numerical image-quality scores alone.
14KUpdated 3 weeks agoMIT
#llama.cpp backend#Ollama integration#Streaming inference
Instructor is an open-source library for developers who need structured data from local LLMs or cloud models. It turns natural-language input into typed objects that applications can use, with validation and retries built into the extraction process. Its focus is data extraction.
18.2KUpdated 4 days agoApache-2.0
Windows#Agent Client Protocol#llama.cpp backend#Ollama integration
avante.nvim brings Cursor-style AI assistance into Neovim for developers who want to keep their editor and choose their own model backend. It answers questions about code, suggests changes and applies edits directly to source files. The plugin is open source under Apache 2.0.
4.8KUpdated 1 day ago
Docker · Web#Git integration#Human approval#LLM tracing
Agenta is an MIT-licensed, open source workspace for teams that want AI agents to handle tasks in chat and continue recurring work in the background. You can self-host it with Docker Compose or Helm to keep agents and workspace data on your infrastructure, or use Agenta Cloud as a hosted service.
22.9KUpdated 3 days agoGPL-3.0
Docker · Web#MCP#Multimodal input#RAG
MaxKB is a self-hosted AI agent platform for organizations building customer support bots, internal knowledge assistants and business automation. It combines answers grounded in company documents with workflows that can call functions and MCP tools. You can run it on your own server through Docker and use it through a browser.
3.3KUpdated 2 months agoMIT
Docker · Web#Multilingual#Speaker diarization#Voice activity detection
Whisper ASR Webservice turns Whisper speech recognition into a self-hosted API for developers adding transcription to their apps or services. It runs in Docker on your own machine or server, with CPU processing or CUDA GPU acceleration. The Python project is open source under the MIT license.
15.9KUpdated 1 month agoApache-2.0
#llama.cpp backend#MLX#Ollama integration
Outlines is an open-source Python library for developers who need LLM responses to match a defined structure. It constrains output during generation, reducing the need to repair malformed JSON or retry responses that don't fit an application's requirements. The library uses the Apache 2.0 license.
2.2KUpdated 3 days agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Guardrails
node-llama-cpp is an open source library for developers adding local LLM inference to JavaScript and TypeScript applications. It connects Node.js, Bun and Electron to llama.cpp, running GGUF models on your own machine. Its MIT license allows use in commercial projects.
4.8KUpdated 2 days ago
Web · Browser Extension
Whispering turns speech into text through a browser interface, a Chrome extension, or a desktop app. It's for people who prefer dictating to typing and want a shortcut to start recording. The app is open source under AGPL-3.0-or-later.
84.5KUpdated 1 day agoBSD-3-Clause
Docker#Guardrails#MCP#Tool calling
Scrapling is a Python web scraping framework for developers collecting website data or giving AI agents access to web pages. Its adaptive parser can find previously selected elements after a site's layout changes, reducing the need to repair extraction rules. It's open source under the BSD-3-Clause license and runs on your own machine or in Docker.
15.9KUpdated 1 month agoApache-2.0
Web#MLX#Multi-user access
Kubeflow is a self-hosted AI platform for teams that run machine learning workloads on Kubernetes. It brings model development, training and production workflows into a modular stack that can run on a local laptop, on-premises infrastructure or a cloud Kubernetes cluster. It's open source under Apache 2.0.
6KUpdated 2 months agoGPL-3.0
macOS · Windows · Linux#Batch processing#ONNX#Visual workflows
chaiNNer is a desktop image processing editor for people who want AI upscaling and repeatable editing workflows without writing scripts. It runs models locally on Windows, macOS, and Linux. Connected nodes let you combine model processing with ordinary image edits in the same workflow.
1.3KUpdated 1 month ago
Web#Human approval#MCP#Multi-user access
Hexabot is a self-hosted AI workflow automation platform for teams building customer service assistants and business automations. It connects conversations to actions across websites, messaging platforms, social channels and custom entry points. You can run it on-premise or in your own cloud, with workflows, memory and customer data in infrastructure you control.
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
19.9KUpdated 8 months agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#Git integration#LM Studio integration
Bolt.diy is a self-hosted AI coding assistant for people building full-stack Node.js web apps who want to choose their own model provider. It's a fork of Bolt.new, with a browser workspace and an Electron desktop app for macOS, Windows, and Linux. You can also run it through Docker.
631Updated 2 years agoMIT
macOS · Windows · Linux · Web · Browser Extension#Batch processing#Multilingual#Quantization
TranslateLocally is an open source machine translation app for Windows, macOS and Linux. It's for people who want to translate text on their own device, including those handling material they don't want to send to a remote translation service. The desktop app has an MIT license and uses Marian and Bergamot translation models.
8.2KUpdated 4 weeks agoMIT
macOS · Windows · Linux#Code execution
Pinokio is a local launcher for people who want to try AI apps on their own computers. It pairs a catalog of community projects with scripts that handle the commands each app needs. Pinokio runs on Windows, macOS, and Linux, and its code is open source under the MIT license.
6.9KUpdated 22 hours agoApache-2.0
#Agent Client Protocol#Human approval#MCP
codecompanion.nvim brings model chat, inline code edits, and coding agents into Neovim. It's for developers who want AI assistance inside their editor, with a choice between local LLMs through Ollama and cloud providers such as Anthropic, OpenAI, and Google Gemini. The plugin runs in Neovim; where model processing happens depends on the backend you connect.
29.8KUpdated 1 day agoMIT
macOS · Windows · Linux#Code execution#Guardrails#Human approval
OpenAI Agents SDK is an open-source Python framework for developers building AI apps that need to use tools, delegate tasks, or work across multiple steps. Its runtime manages agent turns and conversation state while letting developers express workflows in ordinary Python. It uses the MIT license.
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
1.7KUpdated 1 week agoMPL-2.0
Linux#Multilingual#Works offline
Speech Note combines offline dictation, reading aloud and translation in a desktop app for Linux and Sailfish OS. It's for people who want to take multilingual notes, type by voice or listen to text without sending their words to a cloud service. Speech and text processing stay on your device; models are downloaded separately through the app's graphical browser.
7.2KUpdated 2 days agoApache-2.0
Windows · Linux · Docker · Web
nvitop is an interactive terminal monitor for NVIDIA GPUs and the processes using them. It runs locally on Linux and Windows and suits people running AI workloads who need to see device usage alongside host process information. The project is open source under Apache 2.0.
10.6KUpdated 3 months agoMIT
#Hugging Face integration#Speaker diarization#Voice activity detection
pyannote.audio is a Python toolkit that separates an audio recording into timed segments labeled by speaker. It's for developers and researchers who need to track who spoke when, with pretrained models that run on their own hardware. The toolkit is open source under the MIT license.
6KUpdated 3 months agoApache-2.0
macOS · iOS#Multimodal input#Ollama integration#Works offline
Enchanted is a ChatGPT alternative for macOS, iOS and visionOS that connects to models on your own Ollama server. It's for people who want an Apple app for chatting with privately hosted models such as Llama 2, Mistral, Vicuna and Starling. You supply the model server. The app is open source under the Apache License 2.0.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
3.5KUpdated 6 days agoGPL-3.0
#Git integration#Human approval#llama.cpp backend
gptel brings LLM conversations into the Emacs buffers where you already write and work. It's for Emacs users who want AI help alongside their text or code, with a choice of local model servers and cloud providers. The client is open source under GPL-3.0.
11.7KUpdated 1 day ago
Docker · Web#LLM tracing#MCP#Ollama integration
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
8.2KUpdated 21 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
12.2KUpdated 1 month agoMIT
#Multilingual#Semantic search
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
9.3KUpdated 2 months agoApache-2.0
#Multilingual#Multimodal input#RAG
PaperQA2 is an open source Python research assistant for people who need answers grounded in a collection of scientific papers. It searches documents on your machine and writes answers with in-text citations, including page references. Researchers can use it to summarize findings or check for contradictions across papers, while developers can build it into their own research tools.
10.9KUpdated 1 day agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.
46.3KUpdated 1 day agoApache-2.0
macOS · Linux#Hybrid search#Semantic search
Milvus is an open-source vector database for developers building RAG applications, image search and recommendation systems. It stores embeddings alongside metadata so applications can retrieve related text, images or multimodal data. You can run it on your own hardware, from a laptop prototype to a distributed production cluster.
9.3KUpdated 21 hours agoAGPL-3.0
Docker#MCP#Semantic search
ParadeDB adds full-text, vector and hybrid search to the Postgres database that holds your application data. It's for developers who need search and analytics without maintaining a separate Elasticsearch cluster or synchronizing a second copy of their data. It runs as the pg_search extension in self-managed Postgres, including local Docker deployments.