Tools to Run AI Models Locally

The software that loads a model and runs it on your machine: one-step apps like Ollama for a laptop, and servers like vLLM for a GPU box.

Subcategories

100+ tools
An open source Android AI chat app that runs GGUF models offline through llama.cpp and connects to remote providers, including Ollama and OpenAI.

2.7KUpdated 2 weeks agoMIT

Android#GGUF#Hugging Face integration#llama.cpp backend

PyTorch quantization library that reduces model memory use and speeds training and inference on your hardware, with CPU, GPU and mobile deployment support.

3KUpdated 5 days ago

Linux · iOS#Hugging Face integration#LoRA#Quantization

Local LLM desktop app for Windows, macOS and Linux. Run GGUF models offline or connect other apps through OpenAI- and Anthropic-compatible APIs.

47.7KUpdated 1 month agoAGPL-3.0

macOS · Windows · Linux · Docker · Web#GGUF#llama.cpp backend#LoRA

A self-hosted text-to-speech API for Kokoro-82M. Generate speech locally on CPU, NVIDIA GPU or Apple Silicon, with multi-speaker audio and captions.

5.5KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API

Self-hosted AI inference server runs TensorRT, PyTorch and ONNX models on GPUs or CPUs, with dynamic batching and a BSD-3-Clause license.

11KUpdated 1 week agoBSD-3-Clause

Windows · Linux · Docker#Batch processing#ONNX

A local LLM runner that packages a model and runtime in one file for macOS, Linux, BSD and Windows. Open source under Apache 2.0.

26.1KUpdated 5 hours ago

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

Favicon of DeepSpeed

DeepSpeed

1 video
Open-source PyTorch optimization library for distributed training and inference, with Apache 2.0 licensing and support for NVIDIA and AMD GPUs.

43.2KUpdated 1 day agoApache-2.0

Windows#Distributed execution

An open source engine for running ONNX models on Windows, macOS, Linux, mobile devices and the web, with CPU, GPU and NPU support.

22KUpdated 29 minutes agoMIT

macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX

Favicon of Kokoro

Kokoro

6 videos
Open-source text-to-speech model and library for local speech generation, with multilingual voices, Apache 2.0 licensing and Apple Silicon GPU support.

9.1KUpdated 1 year agoApache-2.0

macOS · Windows#Batch processing#Multilingual#ONNX

A desktop AI art manager for Windows, macOS, and Linux that runs Stable Diffusion Web UIs locally and shares model files across them.

8.9KUpdated 2 weeks agoAGPL-3.0

macOS · Windows · Linux#Hugging Face integration#LoRA

An open-source local LLM training framework with LoRA, multimodal support, and deployment through vLLM, SGLang or LMDeploy. Apache 2.0 licensed.

15.8KUpdated 2 days agoApache-2.0

Web#Distributed execution#Hugging Face integration#LoRA

A local AI inference library for C++ and Python that runs Transformer models on CPUs and GPUs, with quantization to reduce memory use. MIT licensed.

4.7KUpdated 5 days agoMIT

#Batch processing#Multilingual#Quantization

Python library for running GGUF models locally through llama.cpp, with a self-hosted OpenAI-compatible server and CPU or GPU support.

10.6KUpdated 1 week agoMIT

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

An open-source local LLM runner for Linux, macOS and Windows via WSL2, with Podman or Docker isolation and llama.cpp or vLLM inference.

3.1KUpdated 1 day agoMIT

macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend

Favicon of koboldcpp

koboldcpp

1 video
Local LLM runner for GGUF and GGML models on Windows, macOS and Linux, with CPU or GPU support, a browser UI and an AGPL-3.0 license.

11.9KUpdated 4 days agoAGPL-3.0

macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

An open-source JavaScript library that runs AI models in your browser using ONNX Runtime, with CPU execution and optional WebGPU acceleration.

16.3KUpdated 1 week agoApache-2.0

Web#Hugging Face integration#Image-to-image#Multilingual

A mobile AI assistant that runs GGUF models on iOS and Android. Core chat works offline after a model download and needs no account.

8.5KUpdated 2 days agoMIT

iOS · Android#GGUF#Hugging Face integration#llama.cpp backend

Favicon of Draw Things

Draw Things

4 videos
An AI image generator that runs offline on iPhone, iPad and Mac, with on-device LoRA training and optional self-hosted or managed cloud compute.

575Updated 1 day agoGPL-3.0

macOS · Linux · iOS · Docker#Image-to-image#Inpainting#LoRA

An open-source Python framework for AI model serving. Build inference APIs and multi-model pipelines locally or deploy with Docker under Apache 2.0.

8.9KUpdated 3 weeks agoApache-2.0

Docker#Batch processing#ControlNet#Distributed execution

Open-source Python toolkit for training and serving LLMs on your own hardware, with Apache 2.0 licensing, quantization and multi-GPU support.

13.7KUpdated 3 weeks agoApache-2.0

#Hugging Face integration#LoRA#Quantization

Command-line client for the Hugging Face Hub. Download models to a local cache, manage repositories, and access hosted inference. Open source under Apache 2.0.

3.9KUpdated 1 day agoApache-2.0

#Hugging Face integration

An open-source local LLM compiler and deployment engine with GPU support across desktop, browser and mobile platforms, plus an OpenAI-compatible API.

23.2KUpdated 24 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API

A local LLM inference engine with OpenAI and Anthropic-compatible APIs. Runs on macOS, Linux and Windows with CPU, CUDA or Apple Silicon support.

7.7KUpdated 5 days agoMIT

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration

A self-hosted text embedding server with a REST API, CPU and GPU support, and offline operation with downloaded model weights. Apache 2.0 licensed.

5.1KUpdated 1 week agoApache-2.0

macOS · Linux · Docker#Batch processing#Hugging Face integration#LLM tracing

An open-weight vision model for image questions, captions and object detection. Run it locally or use hosted inference and fine-tuning.

10.1KUpdated 5 months agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline

Self-hosted computer vision server for images and video, with Docker support, NVIDIA GPU acceleration, and optional Roboflow hosted compute.

2.5KUpdated 1 day ago

macOS · Windows · Linux · Docker#Batch processing#Code execution#Multimodal input

Favicon of Lemonade

Lemonade

1 video
An open source local AI server for chat, image generation, and speech on Windows, macOS, and Linux, with APIs for apps and agents.

5.8KUpdated 2 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend

Open-source Python toolkit for local AI audio generation, fine-tuning and training, with a Gradio interface and support for Stable Audio Open.

3.9KUpdated 4 months agoMIT

Web#Hugging Face integration

Open-source LLM serving toolkit for your own GPU servers, with quantization, text and vision models, and OpenAI-compatible APIs. Apache 2.0 licensed.

8.1KUpdated 3 days agoApache-2.0

#Batch processing#Distributed execution#Hugging Face integration

A self-hosted LLM inference library built on PyTorch for NVIDIA GPUs, with a Python API, OpenAI-compatible serving, and multi-node support.

14.7KUpdated 21 hours ago

Docker#Batch processing#Distributed execution#LoRA

Self-hosted speech API for transcription, translation and speech generation. Runs via Docker on CPU or GPU with faster-whisper, Kokoro and Piper.

3.7KUpdated 5 months agoMIT

Docker#OpenAI-compatible API#Streaming inference

A PyTorch quantization library that reduces LLM memory use for inference and fine-tuning with 8-bit optimizers, LLM.int8() and QLoRA. MIT licensed.

8.5KUpdated 4 weeks agoMIT

macOS · Windows · Linux#LoRA#Quantization

Self-hosted AI model serving platform for Linux, Windows and macOS. Run language, speech and image models through an OpenAI-compatible API under Apache 2.0.

9.6KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input

Favicon of llama-swap

llama-swap

1 video
A local AI proxy that switches models on demand through OpenAI and Anthropic compatible APIs. Runs on macOS, Windows, Linux and FreeBSD under MIT.

5.8KUpdated 2 days agoMIT

macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend

An open-source local LLM framework that splits work across CPUs and GPUs, with SGLang serving and LlamaFactory fine-tuning under Apache 2.0.

19.5KUpdated 1 week agoApache-2.0

Docker#LoRA#Multimodal input#Prompt caching

An open-source Android LLM chat app that runs GGUF models on-device through llama.cpp or connects to Ollama, OpenAI and Claude. Licensed under AGPL-3.0.

2.8KUpdated 1 week agoAGPL-3.0

Android#GGUF#llama.cpp backend#Ollama integration