Ollama Grid Search is a desktop app for comparing LLM responses across models, prompts and inference settings. It runs on macOS, Windows and Linux, and suits developers or anyone choosing a model and prompt combination for a particular task. You can inspect the responses together rather than repeat each test by hand.
The app connects to an Ollama server on your own machine or a remote server. It also works with llmman, a local model runner that provides the Ollama API. The selected server runs the models, so a remote connection sends your test prompts there; the desktop app handles the experiments and result views.
An experiment can test several models against the same prompt, compare different prompts under similar settings, or vary inference parameters across combinations. Repeated runs let you examine how responses differ between attempts. You can compare timing, token counts and generation speed alongside the text, and retry individual responses when needed.
Saved experiment logs make it possible to revisit results, repeat a previous test, or change its parameters for another comparison. You can download experiments as JSON. A prompt database stores reusable prompts and examples, while custom system prompts and default parameters help keep tests consistent.
Ollama Grid Search is open source under the MIT license. Its comparisons focus on generated responses and inference settings, rather than model training. It supports sequential requests or limited concurrency to control the load on the model server.
Claim this page and we'll verify you by hand. Ollama Grid Search gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Ollama Grid Search?Promote it
Something wrong or outdated on this page?
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
35.2KUpdated 2 hours ago
Docker#Agent Skills#LLM tracing#MCP
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
Langfuse is an open source observability and evaluation platform for teams building LLM applications and AI agents. It shows the steps behind a response so developers can investigate failures, slow requests, and cost. Teams can run its MIT-licensed core on their own servers with Docker Compose or Kubernetes, or use Langfuse Cloud as a hosted service.