7.4KUpdated 12 hours agoApache-2.0
Linux#Agent Skills#OpenAI-compatible API#Prompt versioning
Reef is self-hosted infrastructure for developers who want AI agents to improve through feedback on actual interactions. It connects inference and learning with versioned deployment, so an agent can update its model weights or its prompts, rules, and skills while continuing to serve requests. It's open source under Apache 2.0.
22.3KUpdated 1 day agoApache-2.0
Web#Guardrails#LLM tracing#Prompt versioning
Opik is an open-source LLM observability and evaluation platform for developers building AI agents and RAG applications. Its Apache 2.0 licensed platform can be self-hosted on your own hardware or servers; Comet also offers a hosted service. Self-hosting lets teams keep their observability deployment in their own environment.
35.2KUpdated 2 hours ago
Docker#Agent Skills#LLM tracing#MCP
Langfuse is an open source observability and evaluation platform for teams building LLM applications and AI agents. It shows the steps behind a response so developers can investigate failures, slow requests, and cost. Teams can run its MIT-licensed core on their own servers with Docker Compose or Kubernetes, or use Langfuse Cloud as a hosted service.
3.3KUpdated 1 month agoApache-2.0
Docker · Web#Multi-user access#Prompt versioning
Pezzo is an open-source platform for developers and teams managing prompts and monitoring LLM applications. You can run the full stack locally with Docker Compose, keeping the prompt management and monitoring platform on infrastructure you control. Its source code uses the Apache 2.0 license.
4.2KUpdated 1 year agoApache-2.0
Windows · Linux · Web · VS Code#Batch processing#Guardrails#Hugging Face integration
LMQL is a programming language for developers who need model calls and ordinary Python logic in the same program. It lets you define rules for generated text, including types, length limits, allowed answers and stopping phrases. Those rules apply during generation, so you can constrain intermediate responses as well as the final output.
953Updated 3 weeks agoMIT
macOS · Windows · Linux#Ollama integration
Ollama Grid Search is a desktop app for comparing LLM responses across models, prompts and inference settings. It runs on macOS, Windows and Linux, and suits developers or anyone choosing a model and prompt combination for a particular task. You can inspect the responses together rather than repeat each test by hand.
lunary.aiLLM Evaluation and Testing
Docker · Web#LLM tracing#Multi-user access#Prompt versioning
Lunary is a self-hosted LLM observability and prompt management platform for teams building chatbots and AI agents. It brings production traces, user conversations and prompt versions into one place so developers can investigate errors and teams can assess response quality. You can host it in your own infrastructure or use its cloud service.
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
33.9KUpdated 2 weeks agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#Multilingual
SillyTavern is a locally installed LLM frontend for AI hobbyists who want detailed control over character chats and prompts. It builds on TavernAI as an independently developed fork and brings text models, image generation and voice into one interface. It's open source under AGPL-3.0.
6.2KUpdated 2 weeks agoApache-2.0
Web#LLM tracing#Ollama integration#Prompt versioning
Helicone combines an AI gateway with LLM observability for engineers building agents, chatbots, and document processing apps. You can self-host the open source observability platform under Apache 2.0 or use the hosted service. It also integrates with Ollama for apps that run models locally. Its hosted gateway routes requests to external AI providers, so those requests leave your machine.
38.4KUpdated 4 days agoMIT
#Code execution#MCP#Multimodal input
DSPy is a Python framework for developers building AI applications whose tasks need clear inputs, predictable output types, and measurable results. You define what a language model should produce, then compose those tasks into a larger program. It's open source under the MIT license.
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.