Favicon of Curator

Curator

An open-source Python library for synthetic data and structured extraction, with Ollama, vLLM, cloud APIs and managed fine-tuning integrations.

Curator is a Python library for developers preparing LLM training datasets or extracting structured records from existing data. It supports local inference through Ollama and vLLM alongside cloud model APIs, so the same data pipeline can use models on your hardware or a hosted provider. It's open source under Apache 2.0.

Structured outputs let you define the fields a response should contain, then turn results into Hugging Face datasets or pandas tables. You can chain generation stages, such as producing topics before generating examples for each topic. Supported uses include reasoning traces for math and coding, question-answer pairs for domain-specific fine-tuning, product feature extraction and image-based generation.

For large datasets, Curator handles parallel requests, retries and response caching. Cached responses let you repeat a pipeline without calling the model again for the same prompts. It connects to OpenAI-compatible APIs and uses LiteLLM for providers such as Anthropic and Gemini. Provider batch APIs are also supported.

Fine-tuning integrations include Tinker and Fireworks AI for LoRA training. Fireworks training runs on its servers: it receives your uploaded chat data and serves the resulting model for inference. Generated code can run locally, in Docker, through Ray or through e2b. Cloud model calls send prompts to the chosen provider; Curator also collects anonymized usage telemetry, which you can disable.

Its optional hosted viewer uploads generated datasets to Bespoke Labs. Keep that viewer disabled when the dataset must remain in your environment.

Similar to Curator