
PenguinHarness runs on your computer or server and uses AI agents to build runnable agent applications from plain-language requirements. It's for people developing AI apps who want the same tool to handle creation, evaluation, tuning and deployment. The project is open source under Apache 2.0.
Its self-improvement loop tests an agent against benchmarks, examines scores and execution traces, then revises its prompts and Skills. Parallel evaluators can assess the same agent, and snapshots let you restore an earlier state if a revision performs worse. Changes stay within the agent's workspace and Skills; the loop doesn't rewrite the harness core.
You can use a desktop app on macOS, Windows or Linux, or run the server yourself with Docker and access its web interface. A CLI and SDK expose the same engine for scripted work. It supports local and online models through OpenAI-compatible endpoints, with provider connections including DeepSeek, OpenAI, Anthropic and OpenRouter. You can switch models without rewriting the agent. The app and its stored data run locally; choosing a cloud model sends model requests to that provider.
The interface includes agent management, multi-session chat, scheduled tasks and isolated subagents. Built-in Skills cover browser automation, data analysis and software development. Trace views record model requests and tool calls, including failures, timing and token use. Tool calls pass through approval checks, and multi-user projects keep each user's data separate.
Claim this page with an email at penguin.ooo. PenguinHarness gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PenguinHarness?Promote it
Something wrong or outdated on this page?
294Updated 4 hours ago
macOS · Windows · Linux · Web#Agent Skills#Batch processing#Code execution
Ringer runs parallel AI coding agents on your machine and checks their output by executing tests or other validation commands. It's for developers who want to delegate implementation while keeping planning and review with a stronger model. The aim is to reduce model spending on routine coding work without relying on workers' claims that they're finished.
9.4KUpdated 2 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
42.5KUpdated 12 hours agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
483Updated 2 months agoMIT
macOS · Web#Guardrails#LLM tracing#Multi-agent workflows
28.5KUpdated 8 hours ago
Web#Human approval#LLM tracing#MCP
10.7KUpdated 4 days agoMIT
Web#Guardrails#Human approval#LLM tracing
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
HarnessX is a self-hosted Python framework for developers and researchers building AI agents for coding, research, or assistant tasks. It separates model selection from agent behavior, so you can change memory, tools, or safety checks without rewriting the agent. It's open source under MIT.
Mastra is a TypeScript framework for developers building AI agents and applications on their own servers or inside existing web apps. Its server runs locally or as a standalone deployment; Mastra Cloud provides a hosted alternative. Model routing connects to providers such as OpenAI, Anthropic, and Gemini, so running the framework locally doesn't keep those model requests on your machine.
VoltAgent is an MIT-licensed TypeScript AI agent framework for developers building assistants and automated workflows. It pairs code-based agent development with VoltOps, a console for inspecting execution and managing agents. The console can run on your own server or as a cloud service; managed deployment uses hosted infrastructure.