
MLX is a machine learning array framework for researchers and developers building models on their own hardware. Its distinctive feature on Apple silicon is shared CPU and GPU memory: both processors can work on the same arrays without copying data between them. It's open source under the MIT license.
The Python API follows NumPy closely, so developers familiar with array-based numerical code have less new syntax to learn. MLX also has C++, C and Swift APIs. Its neural network and optimizer packages use conventions similar to PyTorch for building and training more complex models.
MLX supports automatic differentiation, vectorization and computation graph optimization, and these operations can be combined. It delays computation until results are needed. Computation graphs form dynamically, and changing the shapes of inputs doesn't trigger slow recompilation, which matters when experimenting with models and different data sizes.
For local AI work, the example projects cover transformer language model training, LLaMA text generation and LoRA fine-tuning. Other examples use Stable Diffusion for image generation and OpenAI's Whisper for speech recognition. These are model development examples rather than a ready-made chat interface.
MLX runs on macOS with Apple silicon CPU and GPU support. Linux packages provide a CUDA backend or CPU-only execution.
Claim this page and we'll verify you by hand. MLX gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MLX?Promote it
Something wrong or outdated on this page?
21.9KUpdated 1 day agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
3KUpdated 5 days ago
Linux · iOS#Hugging Face integration#LoRA#Quantization
TorchAO is a PyTorch library for developers who want to train or run models on their own hardware with less memory and faster computation. It reduces the precision of model weights and activations, with options for language models and image or video generation. Its PyTorch integration works with torch.compile and FSDP2 across most Hugging Face PyTorch models.
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
9.9KUpdated 1 day agoApache-2.0
#Distributed execution
Accelerate is a Python library for developers and researchers who write their own PyTorch training loops and want to use the same code on a local machine or a distributed cluster. It handles the hardware-specific work while leaving the training logic under your control.
41.4KUpdated 3 days agoApache-2.0
#Distributed execution
Colossal-AI is a Python framework for developers and researchers training or serving large AI models on their own GPU hardware. It addresses the memory and computing demands of models that are difficult to fit on a single GPU, with tools for distributing work across a cluster. It's open source under Apache 2.0.
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.