Favicon of MLX LM

MLX LM

A Python package for local LLM inference and fine-tuning on Apple Silicon, built on MLX with Hugging Face model support and an MIT license.

MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.

It works with MLX-compatible models from Hugging Face, including Llama and Mistral, and supports quantized models published by the MLX Community. Inference and fine-tuning run on your hardware. Hugging Face provides model downloads and optional uploads of converted models; a separate Hugging Face Space also offers model conversion and quantization.

The package supports streaming responses, batch text generation and terminal chat that retains context during a session. Its Python API lets developers control how models generate text and incorporate generation into their own applications.

For model adaptation, MLX LM supports both low-rank and full model fine-tuning, including training with quantized models. It can also convert and quantize models for local use or sharing through Hugging Face. Distributed inference and fine-tuning use MLX's distributed capabilities.

Prompt caching reduces repeated processing when several queries share a long context. Memory controls help manage longer prompts and generations, with tradeoffs in speed or output quality. Models that are large relative to available RAM can run slowly; its memory-wiring optimization for large models requires macOS 15 or later.

Similar to MLX LM