
Argilla is an open-source data annotation and feedback tool for AI engineers and domain experts who build training and evaluation datasets. You can run your own Argilla server or deploy it on Hugging Face Spaces. It's licensed under Apache 2.0.
The work centers on reviewing data: experts label examples and contribute feedback, while AI suggestions help with annotation. Filters and semantic search help reviewers find relevant records and focus their attention on data that needs checking. A shared interface gives domain experts a place to contribute their judgment alongside the engineers developing the models.
Argilla supports text classification and named entity recognition, as well as LLM projects involving RAG, preference tuning and reinforcement learning from human feedback (RLHF). Teams can also collect feedback for multimodal projects such as text-to-image models. Its Python SDK connects dataset work to programmatic workflows for ongoing evaluation and model improvement.
Human review can work alongside AI-generated feedback. For example, teams have used Argilla with distilabel to curate training data, and its filters helped identify a data generation bug in the UltraFeedback dataset. These uses make it relevant to teams that need to inspect and correct generated datasets before fine-tuning models.
The project is in maintenance mode: maintainers commit to bug fixes and patches, but aren't developing new features.
Claim this page with an email at argilla.io. Argilla gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Argilla?Promote it
Something wrong or outdated on this page?
3.4KUpdated 10 months agoApache-2.0
#Structured output
Distilabel is an open-source Python framework for engineers building datasets to train or evaluate AI models. It pairs synthetic data generation with LLM feedback, so a pipeline can create examples and judge their quality. It uses the Apache 2.0 license.
28.4KUpdated 1 day agoApache-2.0
macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input
3.3KUpdated 2 weeks agoApache-2.0
Docker · Web#LLM tracing#MCP
Laminar is an open-source platform for developers who need to see why an AI agent failed and check whether a fix worked. You can self-host it with Docker or on Kubernetes, including AWS and GCP, or use its managed cloud service. It uses the Apache 2.0 license.
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
18.5KUpdated 23 hours agoApache-2.0
#LLM tracing#Multimodal input
Label Studio is a self-hosted platform for teams preparing training data or evaluating AI outputs through human review. It handles text, images, audio, video and time series in the same application, including tasks that combine several data types. The open source edition uses the Apache 2.0 license and runs locally or on your own server, with Docker deployment and browser access. A separate hosted cloud edition runs on the provider's infrastructure.
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
DeepEval is a Python framework for testing AI agents, RAG pipelines, and chatbots in your own environment. It's for developers and ML teams who need to compare models or prompts and catch quality regressions before deployment. The open-source framework uses the Apache 2.0 license and fits into Pytest, Python scripts, notebooks, and CI/CD.