llamafile puts an LLM and the software that runs it into a single executable. It’s for people who want to run models locally or share them with others without asking each recipient to set up a separate runtime. The project is open source under Apache 2.0.
Its approach combines llama.cpp with Cosmopolitan Libc so the same model file can run across macOS, Linux, BSD and Windows, as well as different CPU architectures. No installation is required. Prebuilt files include a Qwen3.5 model, and people can package their own models. It also works with separate GGUF model weights, an option for Windows users whose models are too large for a single executable. Models run on the user’s machine, and the software can use a GPU when one is available.
The project also includes whisperfile, a portable speech-to-text tool built on whisper.cpp. It runs locally on the same platforms and can transcribe or translate audio files.
Claim this page and we'll verify you by hand. llamafile gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find llamafile?Promote it
Something wrong or outdated on this page?
5.8KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.
lmstudio.aiComputer and Browser Agents
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
655Updated 2 days agoApache-2.0
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
4.8KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
LM Studio is a desktop application for downloading and running language models on macOS, Windows and Linux. You can search for models, manage downloads and chat with them through the app. Downloaded models can run offline, including document chat that uses files on your computer.
Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
Lollms WebUI is a local, single-user AI interface for people who want text chat and media generation in one place. It runs on Windows, macOS and Linux, with Docker support, and lets writers, developers and other users choose models and task-specific personalities. It's free and open source under Apache 2.0. The project receives minimal maintenance.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.