MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.
It works with MLX-compatible models from Hugging Face, including Llama and Mistral, and supports quantized models published by the MLX Community. Inference and fine-tuning run on your hardware. Hugging Face provides model downloads and optional uploads of converted models; a separate Hugging Face Space also offers model conversion and quantization.
The package supports streaming responses, batch text generation and terminal chat that retains context during a session. Its Python API lets developers control how models generate text and incorporate generation into their own applications.
For model adaptation, MLX LM supports both low-rank and full model fine-tuning, including training with quantized models. It can also convert and quantize models for local use or sharing through Hugging Face. Distributed inference and fine-tuning use MLX's distributed capabilities.
Prompt caching reduces repeated processing when several queries share a long context. Memory controls help manage longer prompts and generations, with tradeoffs in speed or output quality. Models that are large relative to available RAM can run slowly; its memory-wiring optimization for large models requires macOS 15 or later.
Claim this page and we'll verify you by hand. MLX LM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MLX LM?Promote it
Something wrong or outdated on this page?
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
8.5KUpdated 4 weeks agoMIT
macOS · Windows · Linux#LoRA#Quantization
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
960Updated 7 months agoApache-2.0
#Hugging Face integration#LoRA#Quantization
77KUpdated 21 hours agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Image-to-image
130KUpdated 37 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
bitsandbytes is an open-source Python library for developers who need to fit large language model inference or fine-tuning into less memory on their own hardware. It works with PyTorch and carries the MIT license. Its focus is the memory cost of model weights and training, rather than a chat interface.
HQQ is a Python library that compresses language and vision models without needing a calibration dataset. It's for developers preparing models to run on their own hardware or servers, particularly when GPU memory limits the model they can use. The library is open source under Apache 2.0.
Unsloth brings model training and everyday AI use into a desktop app for people who want to run models on their own hardware. Its no-code interface covers chat, fine-tuning and media generation on macOS, Windows and Linux. The Unsloth software is open source under Apache 2.0.
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.