31.3KUpdated 1 day agoApache-2.0
Docker#Hybrid search#Knowledge graphs#llama.cpp backend
Graphiti is a self-hosted Python framework for developers building AI agents that need to remember changing facts. It builds knowledge graphs from conversations, structured records and unstructured text, so an agent can query current information or recover what was true earlier. It's open source under Apache 2.0.
186.5KUpdated 1 day agoAGPL-3.0
Docker#Batch processing#MCP#Structured output
Firecrawl helps developers give AI agents and applications access to current web content. It searches for pages, extracts their contents, and returns data in forms an application can use. The project is open source under AGPL-3.0 and can run on a self-hosted server. Firecrawl also offers a hosted service that requires an account and API key and includes additional features. Reaching live websites requires an internet connection.
93KUpdated 2 hours agoApache-2.0
macOS · Docker#Batch processing#Distributed execution#GGUF
vLLM is an open source engine for serving large language models on hardware you control. It suits developers and teams that need to handle many requests through an API while making efficient use of memory and compute. It's licensed under Apache 2.0 and can run with GPUs or on a CPU.
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
lmstudio.aiComputer and Browser Agents
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
LM Studio is a desktop application for downloading and running language models on macOS, Windows and Linux. You can search for models, manage downloads and chat with them through the app. Downloaded models can run offline, including document chat that uses files on your computer.
130KUpdated 1 hour agoMIT
Web#Code execution#GGUF#Hugging Face integration
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
20.3KUpdated 2 hours agoMIT
#Human approval#LLM tracing#MCP
Pydantic AI is a Python SDK for developers building AI agents into their own applications. Its main draw is Pydantic validation across agent tools and results, so an agent can return structured data that application code can check and use. The SDK is MIT licensed.
206.4KUpdated 2 hours ago
#Code execution#Human approval#MCP
n8n lets technical teams build AI agents and business workflows in a visual editor, then add JavaScript or Python where they need more control. Each step shows its inputs and outputs, so teams can inspect how an agent reached a decision and what happened next.
182KUpdated 17 hours agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input
Ollama runs language models on your own computer or server. It provides a command-line runner and a local API for people building AI applications or connecting existing tools to models they host themselves. The software is distributed under the MIT license.
36.7KUpdated 2 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
SGLang is a self-hosted inference framework for teams that need to serve language and multimodal models on their own hardware. It runs on a single GPU or across distributed clusters and exposes an OpenAI-compatible API. The project is open source under the Apache 2.0 license.
116.8KUpdated 4 days agoMIT
#MCP#Ollama integration#Structured output
Browser Use is an MIT licensed browser agent for developers who want AI to carry out tasks on websites. You can run the open source agent on your own machine from Python, choose a model, and use either a local or cloud browser. A CLI is available for browser tasks too.
12KUpdated 12 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access
h2oGPT is a self-hosted ChatGPT alternative for people who want to chat with local models and ask questions about their own documents. The project is archived and no longer maintained. It's open source under Apache 2.0, with support for Linux, macOS, Windows and Docker.
12.7KUpdated 2 months agoMIT
Windows · Web#Human approval#MCP#Multi-agent workflows
Open Deep Research is a self-hosted AI agent that searches for information and writes research reports. It's for developers and teams who want to choose their own models and research tools. The project is archived and no longer maintained.
4.2KUpdated 1 year agoApache-2.0
Windows · Linux · Web · VS Code#Batch processing#Guardrails#Hugging Face integration
LMQL is a programming language for developers who need model calls and ordinary Python logic in the same program. It lets you define rules for generated text, including types, length limits, allowed answers and stopping phrases. Those rules apply during generation, so you can constrain intermediate responses as well as the final output.
47.9KUpdated 1 day agoGPL-2.0
Linux · Web#LLM tracing#Multi-user access#Structured output
Discourse AI is the official AI plugin bundled with Discourse. Forum administrators can enable its features independently, including an AI bot, semantic search, topic and chat summaries, spam detection and writing assistance. It runs inside a Discourse community, which you can host on your own server.
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
huggingface.coOCR and Document Scanning
Linux#Batch processing#Hugging Face integration#Multimodal input
Qwen2.5-VL is a vision-language model you can run on your own hardware to answer questions about images and video. It's aimed at developers building document processing tools, visual assistants and agents that interact with computer or phone screens. The instruction-tuned 7B model has Apache 2.0 licensing and works with Hugging Face Transformers, with weights available in Safetensors format.
2.1KUpdated 8 months agoMIT
macOS · iOS#llama.cpp backend#Multimodal input#RAG
LLM Farm runs large language models offline on iOS and macOS. It's for people who want on-device AI chat or need to compare how different models perform on Apple hardware before choosing one for a project. The app is open source under the MIT license, with ggml and llama.cpp handling local inference.
10.9KUpdated 6 months agoApache-2.0
Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.
3.5KUpdated 20 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
12.3KUpdated 1 year agoMIT
Linux#Multimodal input#Structured output
Zerox is an MIT-licensed OCR library for developers preparing documents for AI applications. Its Node.js and Python packages run on your own machine or server, while cloud vision models read the document pages and produce Markdown. Document conversion happens locally, but page images go to the selected model provider, so this workflow needs internet access and provider credentials.
1.7KUpdated 2 days agoApache-2.0
#Batch processing#Code execution#Multimodal input
Curator is a Python library for developers preparing LLM training datasets or extracting structured records from existing data. It supports local inference through Ollama and vLLM alongside cloud model APIs, so the same data pipeline can use models on your hardware or a hosted provider. It's open source under Apache 2.0.
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
3.4KUpdated 10 months agoApache-2.0
#Structured output
Distilabel is an open-source Python framework for engineers building datasets to train or evaluate AI models. It pairs synthetic data generation with LLM feedback, so a pipeline can create examples and judge their quality. It uses the Apache 2.0 license.
1.4KUpdated 2 days agoAGPL-3.0
Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.
5KUpdated 23 hours agoApache-2.0
#Code execution#Human approval#MCP
AG2 is an open-source Python framework for developers and researchers building systems where AI agents share work. The code uses the Apache 2.0 license.
147.3KUpdated 1 day agoMIT
#Human approval#RAG#Streaming inference
LangChain is an MIT-licensed open-source framework for developers building AI agents and applications powered by LLMs. It provides a shared interface for models, tools and data connections, so developers can change providers or test workflows without rebuilding the whole application.
9.5KUpdated 21 hours agoApache-2.0
#MCP#Ollama integration#RAG
Spring AI is an open-source Java framework for developers adding AI to Spring applications. It connects application data and APIs to models through a common interface, with Ollama support for local LLM use and integrations with cloud providers such as OpenAI, Anthropic and Amazon Bedrock. The framework runs within your application; your choice of model provider determines whether model requests stay local or go to a cloud service.
84.5KUpdated 6 days agoApache-2.0
Docker#Structured output
Crawl4AI is a self-hosted web crawler and scraper for developers building AI agents, retrieval-augmented generation (RAG) systems and data pipelines. It turns web pages into Markdown or structured JSON and runs as a Python library or a Docker server on your own hardware. The open-source code uses the Apache 2.0 license.
347Updated 2 days agoGPL-3.0
#Batch processing#LM Studio integration#Multilingual
ThunderAI brings AI writing and email processing into Thunderbird for people who want help with their inbox while choosing where their messages go. It can use local models through Ollama or an OpenAI-compatible server such as LM Studio. Cloud connections send the selected content to ChatGPT, the OpenAI API, Google Gemini or Claude instead.
4.4KUpdated 2 days agoMIT
Web#LoRA#Multimodal input#Ollama integration
Ollama JavaScript connects Node.js and browser applications to models running through Ollama. It's for developers building chat interfaces, AI agents or other apps that need a local LLM backend. The library is open source under the MIT license, with TypeScript types and an API that follows Ollama's REST interface.
23.1KUpdated 1 day agoAGPL-3.0
#Code execution#MCP#Multi-agent workflows
Skyvern uses vision models to carry out tasks in a browser, such as filling forms, collecting data, and working through logins. It's for teams automating websites whose layouts change, as well as developers who want AI actions alongside Playwright. Instead of depending on fixed page selectors, it identifies visible elements and decides how to interact with them.
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
7.5KUpdated 1 month agoApache-2.0
#Guardrails#Structured output
Guardrails AI is an open-source Python framework for developers who need to check what goes into an LLM and what comes back. It runs within your application or as a self-hosted service. The framework uses the Apache 2.0 license and helps address risks such as policy violations, hallucinations, and data leakage before outputs reach users.