Tools to Run AI Models Locally

The software that loads a model and runs it on your machine: one-step apps like Ollama for a laptop, and servers like vLLM for a GPU box.

Subcategories

100+ tools
Favicon of GPT4All

GPT4All

1 video
An open-source local AI chatbot for Windows, macOS and Linux. Run models without a GPU or cloud API, and chat privately with your documents.

77.4KUpdated 1 year agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API

Favicon of MLX LM

MLX LM

2 videos
A Python package for local LLM inference and fine-tuning on Apple Silicon, built on MLX with Hugging Face model support and an MIT license.

7.2KUpdated 1 day agoMIT

macOS#Batch processing#Distributed execution#Hugging Face integration

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

1.3KUpdated 1 day agoApache-2.0

Web#LoRA#Multimodal input#Ollama integration

Self-hosted text-to-speech server runs Chatterbox models on CPU or GPU, with voice cloning, audiobook generation and an OpenAI-compatible API.

1.5KUpdated 4 months agoMIT

macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API

Self-hosted speech-to-text API that runs in Docker on CPU or CUDA GPUs, with Whisper, Faster Whisper and WhisperX. Open source under MIT.

3.3KUpdated 2 months agoMIT

Docker · Web#Multilingual#Speaker diarization#Voice activity detection

Local LLM library runs GGUF models through llama.cpp in Node.js, Bun and Electron. MIT licensed, with GPU support and JSON schema enforcement.

2.2KUpdated 3 days agoMIT

macOS · Windows · Linux#Batch processing#GGUF#Guardrails

An MIT-licensed translation app for Windows, macOS and Linux that runs Marian and Bergamot models locally and connects to Chrome and Firefox.

631Updated 2 years agoMIT

macOS · Windows · Linux · Web · Browser Extension#Batch processing#Multilingual#Quantization

Favicon of Pinokio

Pinokio

5 videos
A local AI app launcher for Windows, macOS, and Linux that helps you find and run community projects. It's open source under the MIT license.

8.2KUpdated 4 weeks agoMIT

macOS · Windows · Linux#Code execution

Open-source speaker diarization toolkit with local PyTorch models, CUDA GPU support, and an optional hosted service that processes audio on pyannoteAI servers.

10.6KUpdated 3 months agoMIT

#Hugging Face integration#Speaker diarization#Voice activity detection

Open-source text-to-speech toolkit for local speech generation, voice cloning and model training on Linux, macOS and Windows, licensed under MPL-2.0.

2.3KUpdated 4 months agoMPL-2.0

macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning

A self-hosted inference framework that coordinates NVIDIA GPU clusters with vLLM, SGLang or TensorRT-LLM and exposes an OpenAI-compatible API.

8.2KUpdated 20 hours ago

#Distributed execution#Multimodal input#OpenAI-compatible API

Open-source Python toolkit for semantic search and RAG, with BGE embedding models, multilingual rerankers, evaluation and fine-tuning under MIT.

12.2KUpdated 1 month agoMIT

#Multilingual#Semantic search

Favicon of OpenVINO

OpenVINO

2 videos
Open-source AI inference toolkit for local or self-hosted deployment on Linux, Windows and macOS, with CPU, Intel GPU and NPU support.

10.9KUpdated 24 hours agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX

A Python library for local embeddings and reranking, using ONNX Runtime with CPU or GPU support. Open source under Apache 2.0.

3.2KUpdated 24 hours agoApache-2.0

#Batch processing#Multilingual#ONNX

A PyTorch library for training and inference on local machines or clusters, with CPU, GPU and TPU support. Open source under Apache 2.0.

9.9KUpdated 1 day agoApache-2.0

#Distributed execution

Open-source computer vision library for local detection, segmentation and tracking, with AGPL-3.0 licensing and exports to ONNX, TensorRT and CoreML.

62.1KUpdated 24 hours agoAGPL-3.0

#ONNX

Favicon of KServe

KServe

2 videos
Self-hosted AI model serving platform for Kubernetes. Serve LLMs and predictive models with vLLM, Hugging Face support and an OpenAI-compatible API.

6KUpdated 23 hours agoApache-2.0

#Hugging Face integration#ONNX#OpenAI-compatible API

Favicon of WebLLM

WebLLM

2 videos
A local LLM engine that runs models in the browser with WebGPU acceleration, OpenAI API compatibility, and an Apache 2.0 license.

19.2KUpdated 2 weeks agoApache-2.0

Web · Browser Extension#OpenAI-compatible API#Streaming inference#Structured output

Python library for running and training embedding and reranker models locally, with Apache 2.0 licensing and pretrained models on Hugging Face.

19.1KUpdated 1 week agoApache-2.0

#Hugging Face integration#Multilingual#Multimodal input

Open-source Rust ML framework for running models locally on CPUs, NVIDIA GPUs or in browsers, with Apache 2.0 licensing and quantized LLM support.

21.1KUpdated 2 days agoApache-2.0

macOS · Web#GGUF#Hugging Face integration#Multilingual

Self-hosted embedding and reranking API with MIT licensing, Hugging Face models, and CPU, NVIDIA, AMD and Apple MPS support.

2.9KUpdated 6 months agoMIT

macOS · Docker#Batch processing#Hugging Face integration#Multimodal input

Favicon of Msty

Msty

3 videos
Chat with local or hosted models, compare answers side by side and ask questions over your documents in Msty Studio's AI workspace.

msty.appAI Notes and Knowledge Bases

Web#llama.cpp backend#MLX#Ollama integration

Run AI models through Docker Desktop, Docker Engine or a standalone binary, with local inference and OpenAI and Ollama compatible APIs.

655Updated 2 days agoApache-2.0

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

An open-source Python library for local speech transcription with Whisper models. It runs on CPUs or NVIDIA GPUs and uses CTranslate2.

25.6KUpdated 3 hours agoMIT

#Batch processing#Hugging Face integration#Quantization

Local AI app for Android and iOS. Run Gemma 4 on-device, ask questions about photos, transcribe audio and compare model performance.

24.8KUpdated 21 hours agoApache-2.0

iOS · Android#Agent Skills#Hugging Face integration#Multilingual

Chat with models on your iPhone or iPad through Liquid LEAP, or connect to your own Ollama or LM Studio server. Cloud providers are optional.

liquid.aiMobile Chat Apps

macOS · iOS

A local AI desktop workspace for Apple Silicon Macs, Windows and Linux x64, with GGUF and MLX models, document search and optional cloud providers.

46Updated 5 days agoAGPL-3.0

macOS · Windows · Linux · Browser Extension#Code execution#GGUF#llama.cpp backend

An AI dictation app for iPhone and iPad that uses Gemma models locally. It supports offline transcription and optional cloud text features.

apps.apple.comDictation and Voice Typing

iOS#Multilingual#Works offline