7.9KUpdated 3 weeks agoApache-2.0
Web
Evidently is an Apache 2.0 Python library for evaluating, testing and monitoring ML and LLM systems. It works with tabular and text data, including predictive models and RAG applications. You can run one-off evaluations or self-host its open-source monitoring UI. Evidently Cloud is a separate hosted service.
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
937Updated 2 years agoApache-2.0
#Hugging Face integration#Multimodal input#Works offline
Molmo is Ai2's family of vision-language models, with code for running and training models on your own hardware. It's for developers and researchers who need to work with images and text, adapt a model, or evaluate it against visual tasks. The Python codebase is open source under Apache 2.0 and builds on OLMo, adding image encoding and generative evaluation.
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
2.9KUpdated 23 hours agoMIT
Web · VS Code#Code execution#Hugging Face integration#MCP
Inspect AI is a Python framework for researchers and developers testing language models and AI agents. Developed by the UK AI Security Institute and Meridian Labs, it evaluates coding, reasoning, knowledge, behavior and multimodal understanding, including tasks where agents must take actions to succeed.
9.4KUpdated 2 weeks agoApache-2.0
#AI red teaming#GGUF#Hugging Face integration
Garak is an open-source LLM vulnerability scanner for developers and security teams assessing models or dialogue systems. It tests local models as well as cloud services, so you can assess a model running on your own hardware or an application exposed through an API. The Python tool uses the Apache 2.0 license.
18.5KUpdated 1 day agoApache-2.0
#LLM tracing#Multimodal input
DeepEval is a Python framework for testing AI agents, RAG pipelines, and chatbots in your own environment. It's for developers and ML teams who need to compare models or prompts and catch quality regressions before deployment. The open-source framework uses the Apache 2.0 license and fits into Pytest, Python scripts, notebooks, and CI/CD.
28.4KUpdated 1 day agoApache-2.0
macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input
Label Studio is a self-hosted platform for teams preparing training data or evaluating AI outputs through human review. It handles text, images, audio, video and time series in the same application, including tasks that combine several data types. The open source edition uses the Apache 2.0 license and runs locally or on your own server, with Docker deployment and browser access. A separate hosted cloud edition runs on the provider's infrastructure.
2.5KUpdated 2 days agoMIT
macOS · Linux#Hugging Face integration#Multilingual
LightEval is a Python toolkit from Hugging Face for evaluating LLMs running on your own hardware or through remote services. It's for developers and researchers comparing models, investigating failures, or testing performance on tasks relevant to their work. It can evaluate a model already loaded in memory as well as one served through an endpoint.
15.9KUpdated 7 months agoApache-2.0
Ragas is an open-source Python library for developers who need repeatable evaluations of LLM applications and retrieval-augmented generation (RAG) systems. It combines model-based scoring with traditional metrics so teams can compare application changes using test results rather than manual judgments alone. Its license is Apache 2.0.
5.1KUpdated 1 year agoApache-2.0
Web#Multi-user access#Semantic search
Argilla is an open-source data annotation and feedback tool for AI engineers and domain experts who build training and evaluation datasets. You can run your own Argilla server or deploy it on Hugging Face Spaces. It's licensed under Apache 2.0.
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
5.1KUpdated 22 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
7.2KUpdated 23 hours ago
Docker#Guardrails#LLM tracing#Tool calling
NeMo Guardrails is an open-source Python toolkit for developers who need control over how an AI assistant responds and uses tools. It runs within your application or as a self-hosted server, including in Docker. The library uses the Apache 2.0 license.