
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
The desktop app runs on Windows, macOS and Linux, with a separate Windows ARM build. You don't need a GPU or a cloud API to run models. CPU inference works on supported computers, while NVIDIA and AMD GPUs can accelerate inference through Vulkan. Apple Silicon M-series processors are recommended for the macOS app; the Linux build supports x86-64 hardware.
GPT4All supports GGUF models, including Mistral and DeepSeek R1 distillations, and lets you customize your local assistant. LocalDocs brings document questions into the same chat application, so you can use a model with your own material as well as for general conversation. Model inference and document chat run locally without requiring a cloud connection.
For developers, a Python client built around llama.cpp provides access to local models from code. A Docker-based API server exposes an OpenAI-compatible HTTP endpoint, which lets applications use GPT4All as a model backend on hardware you control. Integrations include LangChain and the Weaviate vector database, with OpenLIT support for monitoring. The GPT4All software uses the MIT license and permits commercial use.
Claim this page with an email at nomic.ai. GPT4All gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find GPT4All?Promote it
Something wrong or outdated on this page?
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
46Updated 5 days agoAGPL-3.0
macOS · Windows · Linux · Browser Extension#Code execution#GGUF#llama.cpp backend
9.8KUpdated 5 months agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration
23.2KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
407Updated 2 days agoMIT
macOS#MLX#Multimodal input#OpenAI-compatible API
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
Vyact is an open source desktop workspace for people who want to use a local LLM with their documents, email and code. It runs on Apple Silicon Macs, Windows and Linux x64 under the AGPL-3.0 license. Intel Macs aren't supported.
PowerInfer is a local LLM inference engine for developers and researchers who want to run large models on a PC with a consumer GPU. It splits work between the CPU and GPU to reduce GPU memory demands and data transfers. The code is open source under the MIT license.
Slotstream runs Qwen3.8-Flash-Next on Apple Silicon Macs that don't have enough RAM to hold the whole model. It's aimed at people with 16 to 64 GB of memory who want local chat, image questions or a model backend for coding agents. Most model weights stay on the SSD, while frequently used expert networks stay in memory. The full model remains available.