
LLM is an Apache-2.0 command-line tool and Python library for sending prompts to local models and remote APIs. Local model support comes through plugins; cloud providers require their own API access. It can also connect to an arbitrary OpenAI-compatible Chat Completions endpoint, including LM Studio.
You can run prompts and interactive chats, store prompts and responses in SQLite, and generate and save embeddings. LLM also supports extracting structured content from text and images and giving models access to tools. Available capabilities depend on the selected model and plugin.
The CLI fits terminal scripts, while the Python library exposes model access to applications. Prompt history is stored locally, but requests sent to a remote provider are processed by that provider.
Claim this page with an email at llm.datasette.io. llm (Simon Willison) gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find llm (Simon Willison)?Promote it
Something wrong or outdated on this page?
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
2.2KUpdated 3 days agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Guardrails
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
3.1KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization
7.2KUpdated 1 day agoMIT
macOS#Batch processing#Distributed execution#Hugging Face integration
10.5KUpdated 7 months agoApache-2.0
macOS · Windows · Linux · Android · Web#Code execution#MCP#Multimodal input
node-llama-cpp is an open source library for developers adding local LLM inference to JavaScript and TypeScript applications. It connects Node.js, Bun and Electron to llama.cpp, running GGUF models on your own machine. Its MIT license allows use in commercial projects.
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.
Distributed Llama runs a local LLM across several computers, sharing both the computation and the model's memory use. It's for people who want to use their own networked hardware for inference rather than keep the entire workload on one machine. The C++ project is open source under the MIT license.
MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.
aichat brings Ollama and cloud AI services into the same terminal interface for developers and people who work at the command line. It runs locally on macOS, Linux and Windows, with Android support through Termux. Model processing happens through the backend you choose: Ollama supports local models, while providers such as OpenAI, Claude and Gemini process requests in the cloud.