Tools to Run AI Models Locally

The software that loads a model and runs it on your machine: one-step apps like Ollama for a laptop, and servers like vLLM for a GPU box.

Subcategories

100+ tools
A local LLM chat app for iOS, iPadOS, macOS and visionOS. Runs on Apple silicon, works offline and saves chat history locally. Licensed under MIT.

2.3KUpdated 1 year agoMIT

macOS · iOS#MLX#Quantization#Works offline

A self-hosted LLM serving stack built on vLLM, with an OpenAI-compatible API, request routing and GPU cluster monitoring. Licensed under Apache 2.0.

2.6KUpdated 21 hours agoApache-2.0

Web#OpenAI-compatible API#Prompt caching

Favicon of MeloTTS

MeloTTS

1 video
A local text-to-speech library with real-time CPU inference, multilingual voices and custom dataset training. Free under the MIT license.

7.7KUpdated 2 years agoMIT

Docker · Web#Multilingual

Open-source LLM serving infrastructure for Kubernetes with multi-node inference, demand-based autoscaling, LoRA management and vLLM integration.

5.1KUpdated 21 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

A free, open-source face recognition server that runs in Docker, uses FaceNet and InsightFace, and supports CPU and GPU processing.

8.3KUpdated 3 years agoApache-2.0

macOS · Windows · Linux · Docker · Web#Multi-user access#Role-based access

An open-source toolkit for on-device AI across mobile, web and desktop. Apache 2.0 licensed; input stays local, while API metrics go to Google.

37.1KUpdated 19 hours agoApache-2.0

iOS · Android · Web

Favicon of EasyOCR

EasyOCR

2 videos
An open-source Python OCR library that extracts text from images locally, supports multilingual reading, and uses the Apache 2.0 license.

30KUpdated 10 months agoApache-2.0

Windows#Multilingual

An open-source LLM server that runs locally or in the cloud, with OpenAI-compatible APIs, a chat UI, and an Apache 2.0 license.

12.5KUpdated 4 months agoApache-2.0

Docker · Web#Hugging Face integration#OpenAI-compatible API

Open-source Python tools optimize Hugging Face models for local, edge and cloud hardware, with ONNX Runtime, OpenVINO and TensorRT-LLM integrations.

3.5KUpdated 6 days agoApache-2.0

#Hugging Face integration#ONNX#Quantization

An open-source local LLM stack manager that connects Ollama, llama.cpp and AI apps through Docker Compose. Includes a CLI and companion app.

3.2KUpdated 5 days agoApache-2.0

macOS · Linux · Docker#GGUF#llama.cpp backend#MCP

Open-source Python library for object detection and segmentation, with pretrained models and deployment exports. Licensed under Apache 2.0.

34.7KUpdated 2 days agoApache-2.0

AI model hub with an Apache 2.0 Python library for local inference, training and evaluation, plus hosted demos and cloud notebooks.

9.2KUpdated 7 days agoApache-2.0

Docker#Image-to-image#Inpainting#Multimodal input

A local LLM chat app for Android that runs GGUF models on-device through llama.cpp. Open source under the Apache 2.0 license.

894Updated 3 months agoApache-2.0

Android#GGUF#llama.cpp backend

An open-source React Native library that runs GGUF models on iOS and Android through llama.cpp, with GPU acceleration and image and audio understanding.

1KUpdated 3 days agoMIT

iOS · Android#GGUF#llama.cpp backend#Multilingual

A self-hosted model serving library for Python and LLM APIs. Run it on a laptop or cluster with request batching, streaming, and CPU or GPU resources.

44KUpdated 18 hours agoApache-2.0

#Batch processing#Hugging Face integration#ONNX

An open-source LLM training toolkit that turns documents into specialist datasets, with offline generation on macOS and Linux and optional cloud compute.

1.9KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization

An open-source AI music model and synthesis engine that runs locally on Apple Silicon, with a macOS app and AUv3 plugin for DAWs. Apache 2.0 licensed.

1.8KUpdated 2 months agoApache-2.0

macOS#MLX#Streaming inference

Self-hostable MLOps software for experiment tracking, pipelines, dataset versioning and model serving, with an Apache 2.0 Python SDK.

6.9KUpdated 2 days agoApache-2.0

Docker · Web#Code execution#Git integration#Multi-user access

Face analysis toolkit for self-hosted recognition on CPU or NVIDIA GPU, local video face redaction, and commercially licensed models.

29.9KUpdated 3 weeks ago

macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization

Self-hosted AI compute orchestration under MPL-2.0 for training and inference on GPU clouds, Kubernetes, VMs and bare-metal servers.

2.3KUpdated 1 day agoMPL-2.0

macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access

An open-source Ollama model manager for macOS and Linux, with a terminal interface for model cleanup, inspection and memory estimates. MIT licensed.

1.8KUpdated 20 hours agoMIT

macOS · Linux#Ollama integration

A local LLM inference engine for sparse models, with CPU and GPU support on Linux and Windows. Open source under MIT, with CPU-only support on Apple Silicon.

9.8KUpdated 5 months agoMIT

macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration

Open-source distributed LLM software runs inference and fine-tuning across shared GPUs, with public or private networks and support for Llama 3.1.

10.6KUpdated 2 years agoMIT

macOS · Windows · Linux · Docker#Hugging Face integration

Open-source Python face recognition library for local or self-hosted use, with TensorFlow and PyTorch support, facial analysis, and an MIT license.

23.5KUpdated 23 hours agoMIT

Docker

A Python reranking library with a shared interface for local models and cloud APIs, CPU inference through FlashRank, and an Apache 2.0 license.

1.6KUpdated 9 months agoApache-2.0

#Hugging Face integration#Multilingual#Multimodal input

An on-device AI engine runs automation models locally on phones and tiny devices, with speech, vision and optional cloud routing.

6.1KUpdated 5 days ago

macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization

An open-source PyTorch toolkit for training and testing object detection and segmentation models, with GPU operations and an Apache 2.0 license.

33KUpdated 3 years agoApache-2.0

An open-source voice AI framework that processes speech directly, with local inference on Mac and iPhone through MLX and self-hosted server backends.

11.2KUpdated 5 months agoApache-2.0

macOS · iOS · Web#Hugging Face integration#MLX#Quantization

An open-source macOS dictation app that runs Whisper and Parakeet models locally on Apple Silicon and transcribes microphone recordings or audio files.

3KUpdated 3 weeks agoMIT

macOS#Batch processing#Hugging Face integration#Multilingual

A self-hosted LLM API server that runs ExLlamaV3 models on your hardware, with OpenAI-compatible endpoints and an AGPL-3.0 license.

1.4KUpdated 2 days agoAGPL-3.0

Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

An open source local LLM runner that splits inference and memory across your computers. Runs on Linux, macOS and Windows under the MIT license.

3.1KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization

Local LLM toolkit that runs language and multimodal models on Rockchip NPUs, with model conversion, quantization and C/C++ interfaces.

1.7KUpdated 2 days ago

Linux#Multimodal input#Quantization

Favicon of llm-d

llm-d

1 video
An open-source LLM inference stack for self-hosted Kubernetes clusters, with vLLM and SGLang backends and support for GPUs, TPUs, XPUs and CPUs.

4.7KUpdated 24 hours agoApache-2.0

#Batch processing#Distributed execution#OpenAI-compatible API

A CLI and Python library for local LLMs and remote APIs, with SQLite prompt history, embeddings, and support for OpenAI-compatible endpoints.

12.6KUpdated 1 week agoApache-2.0

#LM Studio integration#Multimodal input#OpenAI-compatible API

Favicon of Willow

Willow

1 video
An open-source voice assistant for ESP32-S3-BOX hardware, with offline commands, self-hosted speech recognition, and Home Assistant integration.

3.1KUpdated 2 months agoApache-2.0

#Home Assistant integration#Voice activity detection#Wake word detection

A free local AI image generator that runs Stable Diffusion on Windows, Linux and macOS, with a browser interface for prompts, image editing and ControlNet.

10.5KUpdated 3 weeks ago

macOS · Windows · Linux · Web#Batch processing#ControlNet#Image-to-image