Favicon of MLC LLM

MLC LLM

An open-source local LLM compiler and deployment engine with GPU support across desktop, browser and mobile platforms, plus an OpenAI-compatible API.

Screenshot of MLC LLM website

MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.

The compiler optimizes models for native execution on the target platform. This makes it relevant to developers building on-device AI applications or deploying models on their own servers, with a shared engine across environments that use different GPU technologies.

Desktop support covers Linux, Windows and macOS. On Linux and Windows, Vulkan supports AMD, NVIDIA and Intel GPUs. Other backends include ROCm for AMD and CUDA for NVIDIA, depending on the platform. On macOS, it uses Metal with Apple GPUs, AMD discrete GPUs and Intel integrated GPUs.

Mobile execution uses Metal on Apple A-series GPUs for iOS and iPadOS, and OpenCL on Adreno and Mali GPUs for Android. Browser execution supports WebGPU and WASM. These targets let developers bring model inference into desktop software, mobile apps and web applications.

MLCEngine provides an OpenAI-compatible API. Developers can access it through a REST server, Python, JavaScript, iOS or Android interfaces, all backed by the same compiler and inference engine.

Similar to MLC LLM