
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
Kiln runs on your machine and can work fully offline with Ollama. It also connects to cloud providers through your own API keys, including OpenAI, Anthropic, Gemini and OpenRouter, plus OpenAI-compatible endpoints. Cloud model calls go to the selected provider; choosing a local backend keeps model execution local.
Evaluations combine model-based judging with golden datasets and human ratings. The Eval Builder generates judges and evaluation data, while Auto-Optimize searches for better prompts and model configurations against those scores. An AI assistant can create evaluations, run experiments and propose improvements through conversation.
For document-based applications, Kiln supports RAG over PDFs, images, video and audio, with evaluations generated from your documents. Agents can use MCP tools, reusable skills and subagents. Synthetic data generation and fine-tuning share the same task data used for evaluation.
Git sync keeps datasets in your team's repository. The open-source Python library uses the MIT license and runs the app's tasks and project files in production without a rewrite.
Claim this page with an email at kiln.tech. Kiln gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Kiln?Promote it
Something wrong or outdated on this page?
38.4KUpdated 4 days agoMIT
#Code execution#MCP#Multimodal input
DSPy is a Python framework for developers building AI applications whose tasks need clear inputs, predictable output types, and measurable results. You define what a language model should produce, then compose those tasks into a larger program. It's open source under the MIT license.
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
3.1KUpdated 2 days agoApache-2.0
macOS · Windows · Linux · Web#Code execution#MCP#Multi-agent workflows
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
953Updated 3 weeks agoMIT
macOS · Windows · Linux#Ollama integration
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.
BotSharp is a self-hosted framework for .NET developers building AI agents into business applications. Written in C#, it runs on Windows, Linux and macOS and is open source software under Apache 2.0. Its plugin design lets teams choose their model provider, storage and interface while keeping agent coordination in the same framework.
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.
Ollama Grid Search is a desktop app for comparing LLM responses across models, prompts and inference settings. It runs on macOS, Windows and Linux, and suits developers or anyone choosing a model and prompt combination for a particular task. You can inspect the responses together rather than repeat each test by hand.