Tools to Run AI Models Locally

The software that loads a model and runs it on your machine: one-step apps like Ollama for a laptop, and servers like vLLM for a GPU box.

Subcategories

100+ tools
A self-hostable model-sharing platform with accounts, uploads and comments. Its public website adds hosted generation that local setup does not include.

7.3KUpdated 17 hours agoApache-2.0

Linux · Docker · Web#Multi-user access

A C++ neural translation library based on Marian NMT, with native and WebAssembly builds for local processing. Open source under MPL 2.0.

550Updated 2 years agoMPL-2.0

Web#Multilingual

A local document retrieval library that matches text queries to page images without OCR. MIT-licensed Python code supports NVIDIA and Apple Silicon GPUs.

2.8KUpdated 1 month agoMIT

macOS#Batch processing#Hugging Face integration#LoRA

Local LLM quantization library for smaller model weights and inference on your hardware. MIT licensed, with CPU and GPU support; archived and unmaintained.

2.3KUpdated 1 year agoMIT

Linux#Batch processing#GGUF#Hugging Face integration

Desktop AI app powered by llama.cpp, with downloads for Windows and Intel or Apple Silicon Macs. The desktop app is deprecated and unsupported.

desktop.backyard.aiDesktop Chat Apps

macOS · Windows#llama.cpp backend

Self-hosted AI application server with an OpenAI-compatible API, local Ollama and vLLM backends, document search and agent tool calling. MIT licensed.

8.4KUpdated 21 hours agoMIT

#Agent Skills#Batch processing#Guardrails

A self-hosted ChatGPT alternative that runs Llama 2 and Code Llama locally, with an MIT license and an OpenAI-compatible API.

10.9KUpdated 3 years agoMIT

macOS · Docker · Web#GGUF#llama.cpp backend#OpenAI-compatible API

An offline AI assistant for Android and iOS that runs models on your phone, keeps chats encrypted on-device, and offers optional hosted models.

layla-network.aiAI Characters and Roleplay

iOS · Android#Code execution#GGUF#llama.cpp backend

Local LLM web interface for Windows, macOS and Linux. Use GGUF models, Ollama or cloud APIs, with local chat storage and an Apache 2.0 license.

4.8KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

Self-hosted LLM inference engine for Hugging Face models, with OpenAI-compatible APIs, multimodal support, and CPU or GPU execution under AGPL-3.0.

1.9KUpdated 3 weeks agoAGPL-3.0

macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

An on-device AI SDK that runs text, image and audio models on macOS, Windows and Linux, with GGUF, MLX and an OpenAI-compatible API.

qualcomm/GenieXInference Libraries and Bindings

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

Open-source vision language model for local image and text tasks, with Apache 2.0 licensing, Transformers support, and a small GPU memory footprint.

3.9KUpdated 1 week agoApache-2.0

#Hugging Face integration#LoRA#Multimodal input

An open source on-device AI engine for local LLMs and image models, with iOS, Android, CPU and GPU support under Apache 2.0.

16.2KUpdated 1 day agoApache-2.0

Windows · iOS · Android#Image-to-image#Multimodal input#ONNX

A self-hosted Kubernetes operator that deploys Hugging Face models with vLLM, manages GPU capacity, and runs fine-tuning and document retrieval services.

1KUpdated 1 week ago

#Distributed execution#Hugging Face integration#LoRA

A self-hosted AI serving framework that deploys models and pipelines on Kubernetes, on-premises or in the cloud, under the Business Source License.

4.8KUpdated 8 months ago

A local LLM app that runs models offline on iOS and macOS, supports text and vision models, and uses ggml and llama.cpp under the MIT license.

2.1KUpdated 8 months agoMIT

macOS · iOS#llama.cpp backend#Multimodal input#RAG

Local vision-language models for image and video understanding, with Apache 2.0 code, mobile deployment and support for Ollama and llama.cpp.

26.5KUpdated 3 weeks agoApache-2.0

macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend

An open-source Python text-to-speech library that runs locally on CPU or CUDA GPUs and controls voice style through text descriptions.

5.6KUpdated 2 years agoApache-2.0

macOS#Hugging Face integration

A Python library for running and training CLIP image-text models on your own hardware, with local checkpoints and Hugging Face model support.

14.2KUpdated 5 days ago

#Hugging Face integration#Multimodal input

Open-source machine learning library in C/C++ with CPU, GPU, NPU and browser backends, quantization support, and an MIT license.

15.4KUpdated 6 days agoMIT

Web#Quantization

Local LLM acceleration library for Intel CPUs, GPUs and NPUs. Runs on Windows and Linux, integrates with Ollama and llama.cpp, and is archived.

8.9KUpdated 8 months agoApache-2.0

Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration

A local LLM inference library for Windows and Linux with NVIDIA GPUs, MIT licensing, GPTQ and EXL2 support, and an OpenAI-compatible API through TabbyAPI.

4.6KUpdated 7 months agoMIT

Windows · Linux#Batch processing#Quantization#Speculative decoding

A self-hosted LLM inference server under Apache 2.0, with Docker deployment, multi-GPU support and an OpenAI-compatible chat API. The project is archived.

10.9KUpdated 6 months agoApache-2.0

Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

An open-source ChatGPT alternative for Chrome that uses Ollama models locally to chat across tabs, analyze files, and assist with writing.

1.1KUpdated 7 months agoAGPL-3.0

Web · Browser Extension#Multilingual#Multimodal input#Ollama integration

An open-source on-device AI framework for Android, iOS, desktop and web, with model conversion from PyTorch, TensorFlow and JAX.

3.5KUpdated 20 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input

An AI inference framework for Android, iOS, desktop and browsers, with CPU and Vulkan GPU support and PyTorch and ONNX model conversion.

23.9KUpdated 6 days ago

macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization

An open-source model quantization library for local LLMs and vision models, with Apache 2.0 licensing and Hugging Face Transformers integration.

960Updated 7 months agoApache-2.0

#Hugging Face integration#LoRA#Quantization

A local LLM app for iPhone, iPad, and Mac that works offline after model download, keeps chats on-device, and connects to Siri and Apple Shortcuts.

privatellm.appAutomation and No-Code AI

macOS · iOS#Multilingual#Quantization#Works offline

An open-source wake word library for local voice apps, with English models, custom phrase training, and ONNX support on Linux and Windows.

2.8KUpdated 9 months agoApache-2.0

Windows · Linux#Batch processing#ONNX#Voice activity detection

An open-source Python OCR library that reads PDFs and images locally with PyTorch, runs on CPU or GPU, and supports Docker deployment.

6.4KUpdated 1 day agoApache-2.0

Docker · Web

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

5.1KUpdated 21 hours ago

macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API

An open-source Python package for local text embeddings on CPU, with MIT licensing and integrations for Sentence Transformers and LangChain.

2.2KUpdated 1 day agoMIT

#Hugging Face integration#Multilingual

Open-source AI training and inference framework for your own GPU hardware, with distributed parallelism and Apache 2.0 licensing.

41.4KUpdated 3 days agoApache-2.0

#Distributed execution

A Python toolkit for local LLM compression and inference on Linux, macOS and Windows, with GPTQ, AWQ, GGUF and integrations for vLLM and SGLang.

1.3KUpdated 1 day ago

macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA

A text-to-music model with melody conditioning and local GPU inference. AudioCraft code is MIT licensed; pretrained weights have a noncommercial license.

23.7KUpdated 2 years agoMIT

#Multimodal input

Favicon of Moonshine

Moonshine

3 videos
An on-device AI voice toolkit for speech recognition, intent recognition and text to speech, with support for desktop, mobile, browsers and Raspberry Pi.

11.2KUpdated 1 month ago

macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference