Dataset and Synthetic Data Tools

Build and clean training datasets, or generate synthetic ones with a model, using tools like Distilabel and Data-Juicer.

16 tools
Open-source Python library for synthetic data generation and LLM training, with local models, API-based models, caching, and resumable workflows.

1.1KUpdated 2 years agoMIT

#Hugging Face integration#LoRA#Quantization

An open-source Python library for preparing LLM training data locally or on Slurm and Ray clusters, with filtering, deduplication and generation.

3.4KUpdated 1 day agoApache-2.0

#Batch processing#Distributed execution#Multilingual

A family of AI models you can run offline with Ollama, llama.cpp or LM Studio, with open weights and training data for building specialized agents.

2.1KUpdated 3 weeks agoApache-2.0

Linux#GGUF#Guardrails#Hugging Face integration

A self-hosted vision-language model project for image chat, with a local Gradio interface, GPU inference and Apache 2.0 code.

25KUpdated 2 years agoApache-2.0

macOS · Web#LoRA#Multimodal input#Quantization

Python dataset library for AI training and evaluation. Process local files or load data from the Hugging Face Hub, with Apache Arrow storage and streaming.

22KUpdated 2 days agoApache-2.0

#Hugging Face integration#Semantic search

An open-source LLM training and deployment platform under Apache 2.0. Build specialized models on your own infrastructure or use its hosted service.

9.4KUpdated 2 days agoApache-2.0

Docker#Distributed execution#LoRA#Multimodal input

An open-source Python library for synthetic data and structured extraction, with Ollama, vLLM, cloud APIs and managed fine-tuning integrations.

1.7KUpdated 2 days agoApache-2.0

#Batch processing#Code execution#Multimodal input

An open-source LLM training toolkit that turns documents into specialist datasets, with offline generation on macOS and Linux and optional cloud compute.

1.9KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization

A self-hosted Python toolkit for curating AI training data across text, images, video and audio, with NVIDIA GPU support and an Apache 2.0 license.

1.8KUpdated 19 hours agoApache-2.0

Linux · Docker#Distributed execution#Hugging Face integration#Multilingual

Open-source Python framework for building AI agent teams, generating training data, and simulating agent societies. Licensed under Apache 2.0.

17.8KUpdated 2 weeks agoApache-2.0

#Code execution#Hugging Face integration#Human approval

A Python framework for generating and evaluating LLM datasets, with Apache 2.0 licensing and integrations for Anthropic, Cohere and Argilla.

3.4KUpdated 10 months agoApache-2.0

#Structured output

Open-source AI data processing framework for local machines and Ray clusters, with multimodal cleaning, deduplication and Apache 2.0 licensing.

7.1KUpdated 2 days agoApache-2.0

Docker#Batch processing#Distributed execution#Multimodal input

Favicon of DeepEval

DeepEval

1 video
Open-source LLM evaluation framework in Python with local testing, explainable scores, and support for any LLM judge. Licensed under Apache 2.0.

18.5KUpdated 23 hours agoApache-2.0

#LLM tracing#Multimodal input

Favicon of Ragas

Ragas

2 videos
An open-source Python library for evaluating LLM apps and RAG systems, with custom metrics, test data generation, and LangChain and LlamaIndex integrations.

15.9KUpdated 7 months agoApache-2.0

An open-source AI data annotation tool for human feedback, fine-tuning and evaluation. Run your own server or deploy on Hugging Face Spaces.

5.1KUpdated 1 year agoApache-2.0

Web#Multi-user access#Semantic search

A local AI workbench for macOS, Windows and Linux. Evaluate agents, optimize prompts and run fully offline with Ollama or use cloud APIs.

5.1KUpdated 21 hours ago

macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows

More in Fine-Tuning and Training