
LMQL is a programming language for developers who need model calls and ordinary Python logic in the same program. It lets you define rules for generated text, including types, length limits, allowed answers and stopping phrases. Those rules apply during generation, so you can constrain intermediate responses as well as the final output.
You can run models locally through llama.cpp or Hugging Face Transformers, including with local GPU support. Its documented Windows GPU setup uses WSL2. LMQL also works with OpenAI and Azure OpenAI APIs, which send model requests to cloud services and require API credentials. The language is open source under Apache 2.0.
Prompts can use Python variables, loops and conditional logic. Nested queries let developers reuse prompt components across larger programs, while typed outputs and regular expression constraints help produce structured results such as schema-safe JSON. Applications can combine model responses with tool calls or use LMQL to build interactive chat interfaces.
Its runtime uses speculative execution and caching to reduce repeated work and token use. An asynchronous API supports parallel queries and batching, and applications can stream responses through REST, WebSockets or Server-Sent Events.
For development, LMQL has a locally hosted browser playground and a Visual Studio Code extension. It also integrates with LangChain and LlamaIndex.
Claim this page with an email at lmql.ai. LMQL gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LMQL?Promote it
Something wrong or outdated on this page?
38.4KUpdated 4 days agoMIT
#Code execution#MCP#Multimodal input
DSPy is a Python framework for developers building AI applications whose tasks need clear inputs, predictable output types, and measurable results. You define what a language model should produce, then compose those tasks into a larger program. It's open source under the MIT license.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
318Updated 3 weeks agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Quantization
picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.
33.9KUpdated 2 weeks agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#Multilingual
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
SillyTavern is a locally installed LLM frontend for AI hobbyists who want detailed control over character chats and prompts. It builds on TavernAI as an independently developed fork and brings text models, image generation and voice into one interface. It's open source under AGPL-3.0.