Guidance is an MIT-licensed, open source Python library for developers who need language model output to follow a defined format. It works with local LLM backends including Transformers and llama.cpp, as well as OpenAI's cloud service. It's for application code.
Its main distinction is control over generation itself. Developers can restrict responses to a fixed set of choices, match a regular expression, or require structured JSON through a JSON schema or Pydantic model. With a backend that fully supports Guidance, context-free grammars can enforce more complex output syntax, such as nested markup. These constraints govern format, not the factual accuracy of the answer.
Guidance combines generation with ordinary Python logic, including loops, conditions and tool use. Reusable functions let developers compose larger workflows and grammars, while named captures make generated text available to the surrounding application. A Jupyter notebook widget displays output and distinguishes model-generated text from text supplied by the grammar.
When a grammar determines part of the output in advance, Guidance can insert that text without asking the model to generate it. This reduces model computation and GPU use, particularly for structured output with predictable punctuation or tags.
Local backends run the model on your hardware; the OpenAI backend uses a cloud service. Grammar validation and tests with the Mock model work offline without model API calls.
Claim this page and we'll verify you by hand. Guidance gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Guidance?Promote it
Something wrong or outdated on this page?
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
4.2KUpdated 1 year agoApache-2.0
Windows · Linux · Web · VS Code#Batch processing#Guardrails#Hugging Face integration
2.2KUpdated 3 days agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Guardrails
38.4KUpdated 4 days agoMIT
#Code execution#MCP#Multimodal input
27.9KUpdated 1 day agoApache-2.0
#MCP#Structured output#Tool calling
31.3KUpdated 1 day agoApache-2.0
Docker#Hybrid search#Knowledge graphs#llama.cpp backend
LMQL is a programming language for developers who need model calls and ordinary Python logic in the same program. It lets you define rules for generated text, including types, length limits, allowed answers and stopping phrases. Those rules apply during generation, so you can constrain intermediate responses as well as the final output.
node-llama-cpp is an open source library for developers adding local LLM inference to JavaScript and TypeScript applications. It connects Node.js, Bun and Electron to llama.cpp, running GGUF models on your own machine. Its MIT license allows use in commercial projects.
DSPy is a Python framework for developers building AI applications whose tasks need clear inputs, predictable output types, and measurable results. You define what a language model should produce, then compose those tasks into a larger program. It's open source under the MIT license.
FastMCP is an open-source Python framework for developers connecting AI agents to their own tools and data through the Model Context Protocol (MCP). It supports locally running servers and connections to remote servers, with server development and client access in the same framework. It uses the Apache 2.0 license.
Graphiti is a self-hosted Python framework for developers building AI agents that need to remember changing facts. It builds knowledge graphs from conversations, structured records and unstructured text, so an agent can query current information or recover what was true earlier. It's open source under Apache 2.0.