
Prompt flow is an MIT-licensed, open-source toolkit for developers who build LLM applications and need to test their behavior before deployment. Its development tools run locally, while an optional cloud version in Azure AI supports team collaboration. Feature development has ended.
A flow connects prompts and LLM calls with Python code and other tools in an executable workflow. That gives developers a way to develop application logic alongside prompts, inspect interactions with the model, and debug problems across the flow. The toolkit includes a Python SDK, a command-line interface and a VS Code extension with a visual flow designer.
Evaluation is a central part of the tool. You can run flows against larger datasets and calculate quality and performance metrics, rather than judge an application only by a few individual responses. Testing and evaluation can also run in CI/CD, so teams can check changes before they reach production.
For deployment, developers can use their chosen serving platform or embed a flow in an existing application's code. The supplied examples include a chatbot that answers questions about PDFs, with evaluation as part of the development process.
Local development doesn't mean all data stays on your machine. The software enables telemetry by default and can send usage information to Microsoft; you can disable it.
Claim this page and we'll verify you by hand. Prompt flow gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Prompt flow?Promote it
Something wrong or outdated on this page?
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
157.6KUpdated 3 hours ago
Docker · Web#Code execution#MCP#OpenAI-compatible API
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
61.2KUpdated 6 months agoCC-BY-4.0
Web#Code execution#MCP#Multi-agent workflows
8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
Dify is a source-available platform for teams building AI agents and apps on a visual canvas. Its Community Edition runs on your own server with Docker. Dify also offers a hosted cloud service, while Enterprise deployments can run in a VPC or on a self-hosted server. The Community Edition uses a custom Apache 2.0 derivative license.
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
AutoGen is a framework for developers building AI agents that work together or alongside people. Its agent runtime can run locally or across distributed systems, while model integrations such as OpenAI and Azure OpenAI send requests to external services. It's community-managed and in maintenance mode, with no further features or enhancements planned.
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.