Local & Self-Hosted AI Tools

Browse local and self-hosted AI software, from chat apps and model servers to coding, image, video and voice tools.

800+ tools
Favicon of DeepEval

DeepEval

1 video
Open-source LLM evaluation framework in Python with local testing, explainable scores, and support for any LLM judge. Licensed under Apache 2.0.

18.5KUpdated 1 day agoApache-2.0

#LLM tracing#Multimodal input

An open-source Python RAG system under the MIT license that uses knowledge graphs and community summaries to answer questions about private datasets.

36.2KUpdated 7 days agoMIT

#Knowledge graphs#RAG#Semantic search

An open-source container toolkit that gives Docker workloads access to NVIDIA GPUs on Linux. Requires the NVIDIA driver, but not the host CUDA Toolkit.

4.6KUpdated 1 week agoApache-2.0

Linux

Open source data labeling and AI evaluation platform that runs locally or on your server, with custom annotation interfaces and model-assisted labeling.

28.4KUpdated 1 day agoApache-2.0

macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input

Local AI audiobook converter that turns EPUB books into M4B audio with Kokoro voices. Runs on Windows, macOS and Linux under the MIT license.

8.7KUpdated 7 months agoMIT

macOS · Windows · Linux#Multilingual

Open-source LLM serving toolkit for your own GPU servers, with quantization, text and vision models, and OpenAI-compatible APIs. Apache 2.0 licensed.

8.1KUpdated 3 days agoApache-2.0

#Batch processing#Distributed execution#Hugging Face integration

A self-hosted AI chatbot platform with a visual builder, local LLM support and human handoff. Runs on Docker or Kubernetes, with a hosted option.

322Updated 2 weeks agoMIT

iOS · Android · Docker · Web#Batch processing#Code execution#Guardrails

Favicon of FastMCP

FastMCP

2 videos
An open-source Python MCP framework for building servers, connecting to local or remote tools, and adding interactive interfaces to conversations. Apache 2.0.

27.9KUpdated 1 day agoApache-2.0

#MCP#Structured output#Tool calling

An open-source desktop AI agent for Windows, macOS and browsers, with visual recognition, mouse and keyboard control, and fully local processing.

39.2KUpdated 7 days agoApache-2.0

macOS · Windows · Web#Multimodal input

A self-hosted LLM inference library built on PyTorch for NVIDIA GPUs, with a Python API, OpenAI-compatible serving, and multi-node support.

14.7KUpdated 22 hours ago

Docker#Batch processing#Distributed execution#LoRA

DeepSeek-V3 and R1 are downloadable language-model families for local text generation and reasoning.

91.9KUpdated 1 year agoMIT

#Hugging Face integration

Self-hosted web extraction API converts pages and PDFs to Markdown for LLMs. Apache 2.0 service code runs in Docker; a hosted API is also available.

12.1KUpdated 4 months agoApache-2.0

Docker#Multimodal input#Structured output

Self-hosted speech API for transcription, translation and speech generation. Runs via Docker on CPU or GPU with faster-whisper, Kokoro and Piper.

3.7KUpdated 5 months agoMIT

Docker#OpenAI-compatible API#Streaming inference

Open source search and analytics suite built on Apache Lucene, with vector search for AI applications and Apache 2.0 licensing across its components.

13.8KUpdated 23 hours agoApache-2.0

#Semantic search

An open-source browser automation server for AI agents that runs locally on macOS, Windows or Linux and reads page structure without a vision model.

37.7KUpdated 2 days agoApache-2.0

macOS · Windows · Linux · Docker#Code execution#LM Studio integration#MCP

GPU management and monitoring software for NVIDIA data-center GPUs on Linux, with Kubernetes telemetry and an Apache 2.0 open-source core.

798Updated 1 month agoApache-2.0

Linux

Local diffusion model training software for Windows, macOS and Linux, with full fine-tuning, LoRA, dataset preparation and an AGPL-3.0 license.

3.2KUpdated 1 month agoAGPL-3.0

macOS · Windows · Linux#Inpainting#LoRA

Open-source vision-language models for visual chat, document questions and image retrieval, with downloadable weights and Hugging Face Transformers support.

10.2KUpdated 1 year agoMIT

#Hugging Face integration#Multimodal input

A PyTorch quantization library that reduces LLM memory use for inference and fine-tuning with 8-bit optimizers, LLM.int8() and QLoRA. MIT licensed.

8.5KUpdated 4 weeks agoMIT

macOS · Windows · Linux#LoRA#Quantization

Self-hosted AI model serving platform for Linux, Windows and macOS. Run language, speech and image models through an OpenAI-compatible API under Apache 2.0.

9.6KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input

Generate textures and concept art with a Blender add-on that runs Stable Diffusion locally on CUDA or Apple Silicon GPUs, with optional DreamStudio cloud use.

8.2KUpdated 2 weeks agoGPL-3.0

macOS#Image-to-image#Inpainting#Visual workflows

Favicon of llama-swap

llama-swap

1 video
A local AI proxy that switches models on demand through OpenAI and Anthropic compatible APIs. Runs on macOS, Windows, Linux and FreeBSD under MIT.

5.8KUpdated 2 days agoMIT

macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend

A server-based home video platform that connects cameras to HomeKit, Google Home and Alexa, with an NVR plugin for continuous recording and smart detections.

6KUpdated 2 days ago

#Home Assistant integration

An open-source local LLM framework that splits work across CPUs and GPUs, with SGLang serving and LlamaFactory fine-tuning under Apache 2.0.

19.5KUpdated 1 week agoApache-2.0

Docker#LoRA#Multimodal input#Prompt caching

Open-source neural network model viewer for macOS, Windows, Linux and browsers, with support for ONNX, PyTorch, TensorFlow Lite and Core ML.

33.5KUpdated 2 days agoMIT

macOS · Windows · Linux · Web#GGUF#ONNX

A Python library for asking questions about SQL, CSV and parquet data, with chart generation and a separate hosted business intelligence app.

23.8KUpdated 11 months ago

Docker#Code execution#RAG

LLM training framework with ready-made research scripts, NVIDIA GPU parallelism, and Hugging Face checkpoint conversion through Megatron Bridge.

18KUpdated 1 day ago

Docker#Distributed execution#Hugging Face integration#Quantization

An open-source Android LLM chat app that runs GGUF models on-device through llama.cpp or connects to Ollama, OpenAI and Claude. Licensed under AGPL-3.0.

2.8KUpdated 1 week agoAGPL-3.0

Android#GGUF#llama.cpp backend#Ollama integration

Favicon of GPT4All

GPT4All

1 video
An open-source local AI chatbot for Windows, macOS and Linux. Run models without a GPU or cloud API, and chat privately with your documents.

77.4KUpdated 1 year agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API

Self-hosted subtitle generator runs Whisper locally on CPU or NVIDIA GPU and connects to Bazarr, Plex, Jellyfin, Emby and Tautulli. MIT licensed.

1.5KUpdated 2 months agoMIT

Docker#Batch processing#Multilingual#OpenAI-compatible API

A self-hosted search engine that combines full-text and semantic search, stores vectors for RAG, and integrates with LangChain and MCP.

59.4KUpdated 1 day ago

#Hybrid search#MCP#Multilingual

Local document converter turns PDFs and Office files into Markdown, JSON or HTML, with OCR on CPU, NVIDIA GPUs or Apple Silicon and optional LLM support.

40.1KUpdated 3 weeks agoApache-2.0

macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual

Favicon of smolagents

smolagents

1 video
An open-source Python AI agent library with local models through Transformers or Ollama, cloud API support, and Docker sandboxing.

29.6KUpdated 1 week agoApache-2.0

#Code execution#Hugging Face integration#MCP

A vector search SQLite extension written in C with no dependencies. Runs on Linux, macOS, Windows and in browsers via WebAssembly under Apache 2.0.

8.2KUpdated 4 months agoApache-2.0

macOS · Windows · Linux · Web#Semantic search

Favicon of CrewAI

CrewAI

3 videos
Open-source Python AI agent framework under MIT. Run workflows on your own infrastructure with Ollama models or connect to the OpenAI API.

59.2KUpdated 1 day agoMIT

#MCP#Multi-agent workflows#Ollama integration

Favicon of LlamaIndex

LlamaIndex

2 videos
An MIT-licensed Python framework for document-based AI apps. It works with Ollama, while its separate document services run locally or in the cloud.

52.4KUpdated 2 days agoMIT

#Ollama integration#RAG#Reranking

An open-source LLM evaluation toolkit for macOS and Linux. Test local models on CPU or GPUs, or evaluate hosted APIs, under the MIT license.

2.5KUpdated 2 days agoMIT

macOS · Linux#Hugging Face integration#Multilingual

Open-source Python computer vision library for annotating images, tracking objects and managing datasets. MIT licensed, with RF-DETR and Ultralytics support.

51.1KUpdated 1 day agoMIT

A self-hosted AI agent builder with visual workflows and document retrieval. Runs through Docker, with hosted and commercial editions also available.

29.8KUpdated 1 day ago

Docker · Web#LLM tracing#MCP#Multi-user access

Favicon of Ragas

Ragas

2 videos
An open-source Python library for evaluating LLM apps and RAG systems, with custom metrics, test data generation, and LangChain and LlamaIndex integrations.

15.9KUpdated 7 months agoApache-2.0

Favicon of Applio

Applio

2 videos
Open-source AI voice conversion software for Windows, macOS and Linux. Convert audio, change your voice live and train models locally.

3.8KUpdated 2 days agoMIT

macOS · Windows · Linux · Web#Batch processing#Voice conversion

Favicon of Weaviate

Weaviate

1 video
A self-hosted vector database that combines semantic and keyword search with RAG. Run it locally with Docker, on Kubernetes, or in a managed cloud service.

16.9KUpdated 1 day ago

Docker#Hybrid search#RAG#Reranking

An open-source Python framework for AI chat apps you can self-host, with customizable interfaces, OAuth and integrations including LangGraph and HuggingFace.

12.5KUpdated 1 month agoApache-2.0

Web#LLM tracing#Multi-user access#Streaming inference

Favicon of MLX LM

MLX LM

2 videos
A Python package for local LLM inference and fine-tuning on Apple Silicon, built on MLX with Hugging Face model support and an MIT license.

7.2KUpdated 1 day agoMIT

macOS#Batch processing#Distributed execution#Hugging Face integration

A local NVIDIA Jetson monitoring tool with a terminal interface, Python API and Docker support. Open source under AGPL-3.0.

2.6KUpdated 1 week agoAGPL-3.0

Linux · Docker

Favicon of Vibe

Vibe

1 video
Local transcription app for macOS, Windows and Linux. Uses Whisper, Nemotron and Parakeet, with Ollama analysis and optional Claude API summaries.

7.7KUpdated 4 weeks agoMIT

macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration

Favicon of Cognee

Cognee

2 videos
An open-source AI agent memory platform that runs locally on CPU, connects to Claude Code and Codex, and supports self-hosting or managed cloud hosting.

31.2KUpdated 22 hours agoApache-2.0

Docker · Web#Knowledge graphs#MCP#Multi-user access

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

1.3KUpdated 2 days agoApache-2.0

Web#LoRA#Multimodal input#Ollama integration