Curator is a Python library for developers preparing LLM training datasets or extracting structured records from existing data. It supports local inference through Ollama and vLLM alongside cloud model APIs, so the same data pipeline can use models on your hardware or a hosted provider. It's open source under Apache 2.0.
Structured outputs let you define the fields a response should contain, then turn results into Hugging Face datasets or pandas tables. You can chain generation stages, such as producing topics before generating examples for each topic. Supported uses include reasoning traces for math and coding, question-answer pairs for domain-specific fine-tuning, product feature extraction and image-based generation.
For large datasets, Curator handles parallel requests, retries and response caching. Cached responses let you repeat a pipeline without calling the model again for the same prompts. It connects to OpenAI-compatible APIs and uses LiteLLM for providers such as Anthropic and Gemini. Provider batch APIs are also supported.
Fine-tuning integrations include Tinker and Fireworks AI for LoRA training. Fireworks training runs on its servers: it receives your uploaded chat data and serves the resulting model for inference. Generated code can run locally, in Docker, through Ray or through e2b. Cloud model calls send prompts to the chosen provider; Curator also collects anonymized usage telemetry, which you can disable.
Its optional hosted viewer uploads generated datasets to Bespoke Labs. Keep that viewer disabled when the dataset must remain in your environment.
Claim this page and we'll verify you by hand. Curator gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Curator?Promote it
Something wrong or outdated on this page?
3.4KUpdated 10 months agoApache-2.0
#Structured output
Distilabel is an open-source Python framework for engineers building datasets to train or evaluate AI models. It pairs synthetic data generation with LLM feedback, so a pipeline can create examples and judge their quality. It uses the Apache 2.0 license.
1.1KUpdated 2 years agoMIT
#Hugging Face integration#LoRA#Quantization
15.9KUpdated 7 months agoApache-2.0
Ragas is an open-source Python library for developers who need repeatable evaluations of LLM applications and retrieval-augmented generation (RAG) systems. It combines model-based scoring with traditional metrics so teams can compare application changes using test results rather than manual judgments alone. Its license is Apache 2.0.
8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
38.4KUpdated 4 days agoMIT
#Code execution#MCP#Multimodal input
DataDreamer connects LLM prompting, synthetic data generation, and model training in one Python library. It's for researchers and developers who want to build datasets and use them to fine-tune or align models in reproducible workflows. The library is open source under the MIT license.
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
DSPy is a Python framework for developers building AI applications whose tasks need clear inputs, predictable output types, and measurable results. You define what a language model should produce, then compose those tasks into a larger program. It's open source under the MIT license.