Favicon of llamafile

llamafile

A local LLM runner that packages a model and runtime in one file for macOS, Linux, BSD and Windows. Open source under Apache 2.0.

llamafile puts an LLM and the software that runs it into a single executable. It’s for people who want to run models locally or share them with others without asking each recipient to set up a separate runtime. The project is open source under Apache 2.0.

Its approach combines llama.cpp with Cosmopolitan Libc so the same model file can run across macOS, Linux, BSD and Windows, as well as different CPU architectures. No installation is required. Prebuilt files include a Qwen3.5 model, and people can package their own models. It also works with separate GGUF model weights, an option for Windows users whose models are too large for a single executable. Models run on the user’s machine, and the software can use a GPU when one is available.

The project also includes whisperfile, a portable speech-to-text tool built on whisper.cpp. It runs locally on the same platforms and can transcribe or translate audio files.

Similar to llamafile