
Ragas is an open-source Python library for developers who need repeatable evaluations of LLM applications and retrieval-augmented generation (RAG) systems. It combines model-based scoring with traditional metrics so teams can compare application changes using test results rather than manual judgments alone. Its license is Apache 2.0.
The framework organizes evaluations around experiments, with built-in dataset management and result tracking. You can use its existing metrics or define criteria for your own application. Aspect Critique, for example, scores outputs against a chosen aspect. This lets developers evaluate qualities that matter to their use case rather than relying on a single general score.
Ragas also generates test datasets automatically, including tests based on production data. That gives teams a way to cover different scenarios when they don't already have an evaluation dataset, and to feed examples from actual use into further testing. It integrates with LangChain and LlamaIndex, as well as observability tools.
The library runs in your Python workflow. Model-based evaluations that use OpenAI send requests to its cloud API and require an API key. Ragas collects minimal anonymized usage analytics, with an opt-out available. The collection code is open source, and it publishes aggregated analytics without personal or company-identifying information.
Claim this page with an email at docs.ragas.io. Ragas gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Ragas?Promote it
Something wrong or outdated on this page?
3.4KUpdated 10 months agoApache-2.0
#Structured output
Distilabel is an open-source Python framework for engineers building datasets to train or evaluate AI models. It pairs synthetic data generation with LLM feedback, so a pipeline can create examples and judge their quality. It uses the Apache 2.0 license.
5.1KUpdated 1 year agoApache-2.0
Web#Multi-user access#Semantic search
Argilla is an open-source data annotation and feedback tool for AI engineers and domain experts who build training and evaluation datasets. You can run your own Argilla server or deploy it on Hugging Face Spaces. It's licensed under Apache 2.0.
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
27.3KUpdated 20 hours agoMIT
macOS · Windows · Linux#Code execution#MCP#Tool calling
1.7KUpdated 2 days agoApache-2.0
#Batch processing#Code execution#Multimodal input
1.1KUpdated 2 years agoMIT
#Hugging Face integration#LoRA#Quantization
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
Cua gives AI agents access to computers they can inspect and operate, with tools for desktop automation, local virtual machines, and hosted fleets. It's for developers building agents that work across native apps and browsers, or evaluating how well those agents complete computer tasks. You bring the agent and model.
Curator is a Python library for developers preparing LLM training datasets or extracting structured records from existing data. It supports local inference through Ollama and vLLM alongside cloud model APIs, so the same data pipeline can use models on your hardware or a hosted provider. It's open source under Apache 2.0.
DataDreamer connects LLM prompting, synthetic data generation, and model training in one Python library. It's for researchers and developers who want to build datasets and use them to fine-tune or align models in reproducible workflows. The library is open source under the MIT license.