Favicon of LMQL

LMQL

An open source LLM programming language that combines Python logic with output constraints and supports local llama.cpp and Transformers models or cloud APIs.

Screenshot of LMQL website

LMQL is a programming language for developers who need model calls and ordinary Python logic in the same program. It lets you define rules for generated text, including types, length limits, allowed answers and stopping phrases. Those rules apply during generation, so you can constrain intermediate responses as well as the final output.

You can run models locally through llama.cpp or Hugging Face Transformers, including with local GPU support. Its documented Windows GPU setup uses WSL2. LMQL also works with OpenAI and Azure OpenAI APIs, which send model requests to cloud services and require API credentials. The language is open source under Apache 2.0.

Prompts can use Python variables, loops and conditional logic. Nested queries let developers reuse prompt components across larger programs, while typed outputs and regular expression constraints help produce structured results such as schema-safe JSON. Applications can combine model responses with tool calls or use LMQL to build interactive chat interfaces.

Its runtime uses speculative execution and caching to reduce repeated work and token use. An asynchronous API supports parallel queries and batching, and applications can stream responses through REST, WebSockets or Server-Sent Events.

For development, LMQL has a locally hosted browser playground and a Visual Studio Code extension. It also integrates with LangChain and LlamaIndex.

Similar to LMQL