Favicon of Ragas

Ragas

An open-source Python library for evaluating LLM apps and RAG systems, with custom metrics, test data generation, and LangChain and LlamaIndex integrations.

Screenshot of Ragas website

Ragas is an open-source Python library for developers who need repeatable evaluations of LLM applications and retrieval-augmented generation (RAG) systems. It combines model-based scoring with traditional metrics so teams can compare application changes using test results rather than manual judgments alone. Its license is Apache 2.0.

The framework organizes evaluations around experiments, with built-in dataset management and result tracking. You can use its existing metrics or define criteria for your own application. Aspect Critique, for example, scores outputs against a chosen aspect. This lets developers evaluate qualities that matter to their use case rather than relying on a single general score.

Ragas also generates test datasets automatically, including tests based on production data. That gives teams a way to cover different scenarios when they don't already have an evaluation dataset, and to feed examples from actual use into further testing. It integrates with LangChain and LlamaIndex, as well as observability tools.

The library runs in your Python workflow. Model-based evaluations that use OpenAI send requests to its cloud API and require an API key. Ragas collects minimal anonymized usage analytics, with an opt-out available. The collection code is open source, and it publishes aggregated analytics without personal or company-identifying information.

Similar to Ragas