12.6KUpdated 1 week agoApache-2.0
#LM Studio integration#Multimodal input#OpenAI-compatible API
LLM is an Apache-2.0 command-line tool and Python library for sending prompts to local models and remote APIs. Local model support comes through plugins; cloud providers require their own API access. It can also connect to an arbitrary OpenAI-compatible Chat Completions endpoint, including LM Studio.
3.9KUpdated 1 week agoApache-2.0
Safetensors is a file format and library for developers who store, share, or load AI model weights on their own hardware or servers. It avoids the arbitrary code execution risk of PyTorch's pickle-based files while supporting fast access to tensor data. The project is open source under Apache 2.0, with a Rust implementation and Python support.
147.3KUpdated 1 day agoMIT
#Human approval#RAG#Streaming inference
LangChain is an MIT-licensed open-source framework for developers building AI agents and applications powered by LLMs. It provides a shared interface for models, tools and data connections, so developers can change providers or test workflows without rebuilding the whole application.
9.5KUpdated 21 hours agoApache-2.0
#MCP#Ollama integration#RAG
Spring AI is an open-source Java framework for developers adding AI to Spring applications. It connects application data and APIs to models through a common interface, with Ollama support for local LLM use and integrations with cloud providers such as OpenAI, Anthropic and Amazon Bedrock. The framework runs within your application; your choice of model provider determines whether model requests stay local or go to a cloud service.
4.4KUpdated 2 days agoMIT
Web#LoRA#Multimodal input#Ollama integration
Ollama JavaScript connects Node.js and browser applications to models running through Ollama. It's for developers building chat interfaces, AI agents or other apps that need a local LLM backend. The library is open source under the MIT license, with TypeScript types and an API that follows Ollama's REST interface.
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
16.3KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Image-to-image#Multilingual
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
7.5KUpdated 1 month agoApache-2.0
#Guardrails#Structured output
Guardrails AI is an open-source Python framework for developers who need to check what goes into an LLM and what comes back. It runs within your application or as a self-hosted service. The framework uses the Apache 2.0 license and helps address risks such as policy violations, hallucinations, and data leakage before outputs reach users.
28.6KUpdated 1 day agoMIT
macOS · Windows · Linux#LM Studio integration#MCP#Multi-agent workflows
Semantic Kernel is an MIT-licensed SDK for developers adding AI agents to their applications. It supports local models through Ollama, LMStudio and ONNX, alongside cloud services such as OpenAI and Azure OpenAI. You choose the model backend. The SDK runs on Windows, macOS and Linux and supports C#, Python and Java.
25.5KUpdated 2 days agoMIT
#LLM tracing#Structured output
Stagehand is an open-source browser automation SDK for developers building AI agents that interact with websites and extract structured data. It can run with Chrome on your own machine or use Browserbase's cloud browsers. Local runs require Chrome. The project uses the MIT license and supports TypeScript, Python, and Go.
21.8KUpdated 4 months agoMIT
#llama.cpp backend#Structured output#Tool calling
Guidance is an MIT-licensed, open source Python library for developers who need language model output to follow a defined format. It works with local LLM backends including Transformers and llama.cpp, as well as OpenAI's cloud service. It's for application code.
13KUpdated 22 hours agoApache-2.0
Docker#Agent Skills#Hugging Face integration#Knowledge graphs
txtai is a Python framework for developers building search applications, chat with their data, and AI agents on their own hardware or servers. Its embeddings database combines sparse and dense vector search with graphs and relational data, so the same system can find related content and supply context to language models. It's open source under Apache 2.0.
38.4KUpdated 4 days agoMIT
#Code execution#MCP#Multimodal input
DSPy is a Python framework for developers building AI applications whose tasks need clear inputs, predictable output types, and measurable results. You define what a language model should produce, then compose those tasks into a larger program. It's open source under the MIT license.
27.9KUpdated 1 day agoApache-2.0
#MCP#Structured output#Tool calling
FastMCP is an open-source Python framework for developers connecting AI agents to their own tools and data through the Model Context Protocol (MCP). It supports locally running servers and connections to remote servers, with server development and client access in the same framework. It uses the Apache 2.0 license.
23.8KUpdated 11 months ago
Docker#Code execution#RAG
PandasAI is a Python library for people who want to ask questions about their data in plain language. It works with SQL databases, CSV and parquet data, and can answer questions across multiple pandas DataFrames. Developers can use it in their own applications; analysts can use conversational queries to reduce the code they write for individual questions.
18KUpdated 1 day ago
Docker#Distributed execution#Hugging Face integration#Quantization
Megatron-LM is a Python framework for research teams training large language models on NVIDIA GPU infrastructure. It pairs ready-made training scripts with Megatron Core, a library developers can use to build their own training systems. Its focus is distributed training, with benchmarks on H100 clusters spanning thousands of GPUs.
52.4KUpdated 2 days agoMIT
#Ollama integration#RAG#Reranking
LlamaIndex is an MIT-licensed Python framework for developers building AI agents and apps that answer questions using their own data. It connects documents and other sources to language models, then helps an app find the relevant material when a user asks something. Its open source framework can work with models served through Ollama.
15.9KUpdated 7 months agoApache-2.0
Ragas is an open-source Python library for developers who need repeatable evaluations of LLM applications and retrieval-augmented generation (RAG) systems. It combines model-based scoring with traditional metrics so teams can compare application changes using test results rather than manual judgments alone. Its license is Apache 2.0.
12.5KUpdated 1 month agoApache-2.0
Web#LLM tracing#Multi-user access#Streaming inference
Chainlit is an open-source Python framework for developers building conversational AI apps with their own application logic. You can run the app on your own server and give users a browser chat interface. It's a framework for creating an app, rather than a ready-made chatbot with a fixed model.
13.2KUpdated 1 day agoApache-2.0
#MCP#Ollama integration#RAG
LangChain4j is an Apache 2.0 open-source Java library for developers building chatbots, assistants and AI agents in JVM applications. It connects application code to local LLM backends such as Ollama as well as cloud providers such as OpenAI and Google Vertex AI. Where model requests go depends on the backend you choose.
4.4KUpdated 1 day agoMIT
#Human approval#Multi-agent workflows#Multimodal input
RubyLLM is an MIT-licensed AI framework for developers building Ruby and Rails applications with local or hosted models. Its shared API lets an application switch between Ollama, cloud providers such as Anthropic and OpenAI, and OpenAI-compatible endpoints without rewriting its model integration. The framework runs in your application; model processing happens at the local or hosted backend you choose.
1.4KUpdated 2 months agoMIT
#Multimodal input#Ollama integration#Streaming inference
OllamaSharp is a C# library for developers building .NET applications around Ollama. It connects to Ollama on your own machine or a remote server and covers the full Ollama API, including model management alongside chat and embeddings. It's open source under the MIT license.
14KUpdated 3 weeks agoMIT
#llama.cpp backend#Ollama integration#Streaming inference
Instructor is an open-source library for developers who need structured data from local LLMs or cloud models. It turns natural-language input into typed objects that applications can use, with validation and retries built into the extraction process. Its focus is data extraction.
15.9KUpdated 1 month agoApache-2.0
#llama.cpp backend#MLX#Ollama integration
Outlines is an open-source Python library for developers who need LLM responses to match a defined structure. It constrains output during generation, reducing the need to repair malformed JSON or retry responses that don't fit an application's requirements. The library uses the Apache 2.0 license.
2.2KUpdated 3 days agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Guardrails
node-llama-cpp is an open source library for developers adding local LLM inference to JavaScript and TypeScript applications. It connects Node.js, Bun and Electron to llama.cpp, running GGUF models on your own machine. Its MIT license allows use in commercial projects.
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
29.8KUpdated 24 hours agoMIT
macOS · Windows · Linux#Code execution#Guardrails#Human approval
OpenAI Agents SDK is an open-source Python framework for developers building AI apps that need to use tools, delegate tasks, or work across multiple steps. Its runtime manages agent turns and conversation state while letting developers express workflows in ordinary Python. It uses the MIT license.
3.8KUpdated 23 hours agoApache-2.0
#Hugging Face integration#Quantization
LLM Compressor is an open-source Python library for developers preparing models to run on their own hardware with vLLM. It reduces model size and memory requirements through quantization, and accepts local checkpoints or models from Hugging Face repositories. It's licensed under Apache 2.0.
19.2KUpdated 2 weeks agoApache-2.0
Web · Browser Extension#OpenAI-compatible API#Streaming inference#Structured output
WebLLM runs language models directly in a user's browser, using WebGPU for GPU acceleration. It's an open-source engine for developers building web-based AI assistants and Chrome extensions that process prompts on the user's device rather than an inference server. The project uses the Apache 2.0 license.
26.6KUpdated 1 day agoApache-2.0
Docker#Guardrails#Hugging Face integration#Hybrid search
Haystack is a Python framework for developers building self-hosted AI agents, document search, and apps that answer questions using their own data. Its modular pipelines let teams control which information reaches a model and inspect how retrieval, memory, tools, and generation contribute to an answer. It's open source under Apache 2.0.
7.2KUpdated 23 hours ago
Docker#Guardrails#LLM tracing#Tool calling
NeMo Guardrails is an open-source Python toolkit for developers who need control over how an AI assistant responds and uses tools. It runs within your application or as a self-hosted server, including in Docker. The library uses the Apache 2.0 license.
10.6KUpdated 2 days agoMIT
#Batch processing#Ollama integration#Streaming inference
Ollama Python connects Python applications to models running through Ollama on your own machine or to Ollama's cloud service. It's for developers adding local LLM features to scripts, chat applications, or other Python projects. The library is open source under the MIT license and requires a running Ollama service for local use.
7.4KUpdated 3 weeks agoLGPL-3.0
#Hugging Face integration#LoRA
mergekit combines existing language models into a single model on your own hardware, without additional training or access to the original training data. It's for developers and researchers who want to combine fine-tuned capabilities or adjust the balance between model behaviors. The Python toolkit is open source under LGPL-3.0.
27.1KUpdated 2 hours ago
#Agent Skills#Streaming inference#Structured output
Vercel AI SDK is a TypeScript library for developers building chatbots, generative interfaces, and AI agents. It gives an application one way to work with models from providers such as OpenAI, Anthropic, and Google, so teams can change providers without rebuilding the parts of their app that handle responses. It suits developers who want control over the application they build while choosing how it connects to model services.
43.6KUpdated 1 day agoApache-2.0
Web#MCP
Gradio turns a Python function, model, or API into a web interface. It runs on your computer, making it useful for developers and researchers who want people to try an AI project through a browser. It's open source under the Apache 2.0 license.