407Updated 2 days agoMIT
macOS#MLX#Multimodal input#OpenAI-compatible API
Slotstream runs Qwen3.8-Flash-Next on Apple Silicon Macs that don't have enough RAM to hold the whole model. It's aimed at people with 16 to 64 GB of memory who want local chat, image questions or a model backend for coding agents. Most model weights stay on the SSD, while frequently used expert networks stay in memory. The full model remains available.
7.4KUpdated 12 hours agoApache-2.0
Linux#Agent Skills#OpenAI-compatible API#Prompt versioning
Reef is self-hosted infrastructure for developers who want AI agents to improve through feedback on actual interactions. It connects inference and learning with versioned deployment, so an agent can update its model weights or its prompts, rules, and skills while continuing to serve requests. It's open source under Apache 2.0.
544Updated 14 hours agoMIT
macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
31.3KUpdated 3 weeks agoMIT
macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference
Meetily is a local AI meeting assistant for people who want meeting notes while keeping recordings on their own device. It captures calls from Zoom, Google Meet, Microsoft Teams and other meeting software without placing a bot in the meeting. You can watch the transcript appear during the call, then generate a summary.
13.2KUpdated 2 days agoApache-2.0
#ControlNet#Image-to-image#Inpainting
DiffSynth-Studio is a Python diffusion model engine for developers and researchers who want to generate media and train models on their own hardware. It supports large models on consumer GPUs through memory offloading and quantization, with inference and training in the same framework. It's open source under Apache 2.0.
47.7KUpdated 1 month agoApache-2.0
macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX
exo is a local LLM runner that combines your devices into a cluster, letting you use models too large for one machine's memory. It's for people who want to run large models on their own hardware and developers connecting existing AI clients to local inference. It runs on macOS and Linux under the Apache 2.0 license.
44.7KUpdated 5 hours ago
macOS · Windows · Linux#Hugging Face integration#MCP#OpenAI-compatible API
Jan gives people a ChatGPT-style chat interface for AI models running on their own computer. It's free and available for Windows, macOS and Linux. Local chats work offline, with the model and conversation kept on your machine. You can also use cloud models when you want access to a hosted provider.
29.2KUpdated 1 day agoAGPL-3.0
macOS · Windows · Linux · iOS · Android · Web#Multi-user access#Role-based access#Semantic search
Ente Photos is a Google Photos alternative for people who want encrypted photo backups and AI search without giving a storage provider access to their library. You can use its hosted service or self-host the server. Face detection, grouping and natural-language search run on your device; the hosted service stores your encrypted backups.
93KUpdated 1 hour agoApache-2.0
macOS · Docker#Batch processing#Distributed execution#GGUF
vLLM is an open source engine for serving large language models on hardware you control. It suits developers and teams that need to handle many requests through an API while making efficient use of memory and compute. It's licensed under Apache 2.0 and can run with GPUs or on a CPU.
34.6KUpdated 1 day agoApache-2.0
macOS#ControlNet#Hugging Face integration#Image-to-image
Diffusers is an open-source Python library for developers and researchers who want to run diffusion models on their own hardware or build generation features into an application. It uses PyTorch and supports image, video and audio generation. The library is licensed under Apache 2.0 and supports Apple Silicon.
noemaai.comChat With Your Documents
macOS · iOS#GGUF#MCP#MLX
Noema is a private local AI assistant for iPhone, iPad, Mac, and Vision Pro. It's for people who want to chat with models and work with their own files on Apple hardware without depending on a cloud service. Local chats can stay on your device, and file retrieval runs there too.
lmstudio.aiComputer and Browser Agents
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
LM Studio is a desktop application for downloading and running language models on macOS, Windows and Linux. You can search for models, manage downloads and chat with them through the app. Downloaded models can run offline, including document chat that uses files on your computer.
130KUpdated 40 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
26.6KUpdated 2 months agoMIT
Linux#Multilingual#Voice cloning#Voice conversion
Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.
28.6KUpdated 24 hours agoMIT
macOS · Linux#Distributed execution#LoRA
MLX is a machine learning array framework for researchers and developers building models on their own hardware. Its distinctive feature on Apple silicon is shared CPU and GPU memory: both processors can work on the same arrays without copying data between them. It's open source under the MIT license.
54KUpdated 2 days agoMIT
macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference
whisper.cpp runs OpenAI's Whisper speech recognition models on your own hardware, with fully offline transcription once you've downloaded a model. It's for developers building speech-to-text into applications and people who want to transcribe audio locally. Audio can stay on-device rather than going to a cloud transcription service. The project is open source under the MIT license.
182KUpdated 16 hours agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input
Ollama runs language models on your own computer or server. It provides a command-line runner and a local API for people building AI applications or connecting existing tools to models they host themselves. The software is distributed under the MIT license.
36.7KUpdated 1 hour agoApache-2.0
#Batch processing#Distributed execution#LoRA
SGLang is a self-hosted inference framework for teams that need to serve language and multimodal models on their own hardware. It runs on a single GPU or across distributed clusters and exposes an OpenAI-compatible API. The project is open source under the Apache 2.0 license.
166.8KUpdated 1 day agoApache-2.0
#Hugging Face integration#Multimodal input
Transformers is a Python library for developers and researchers who want to run pretrained AI models or train their own on hardware they control. It covers language, images, audio, video and multimodal work through a shared way of defining models. The library runs in a local Python environment; pretrained checkpoints are available from the separate Hugging Face Hub.
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
77KUpdated 21 hours agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Image-to-image
Unsloth brings model training and everyday AI use into a desktop app for people who want to run models on their own hardware. Its no-code interface covers chat, fine-tuning and media generation on macOS, Windows and Linux. The Unsloth software is open source under Apache 2.0.
5.7KUpdated 2 days agoGPL-3.0
#ONNX
Piper turns text into spoken audio on local hardware. It's an open-source, GPL-3.0 text-to-speech engine for developers building voice features, accessibility tools and self-hosted AI projects. Speech generation runs locally, giving people a way to add a voice to software they control.
3.2KUpdated 4 days agoMIT
macOS · Windows · Linux · iOS · Android#GGUF#Human approval#LM Studio integration
Off Grid AI runs language models on iOS, Android, macOS, Windows and Linux. You can chat, analyze documents and generate images on your own hardware. The mobile app uses the MIT license, while the desktop app uses AGPL.
locallyai.appDesktop Chat Apps
macOS · iOS#MLX#Multilingual#Multimodal input
Locally AI is a native app for running language and vision models on recent iPhones, iPads, and Macs. It's for people who want a private AI assistant on their own device, with text, image processing, and voice conversations available without cloud processing. Once a model is downloaded, it works offline and doesn't require an account.
recurse.chatChat With Your Documents
macOS#GGUF#Hugging Face integration#OpenAI-compatible API
RecurseChat is a paid Mac app for people who want to chat with AI and ask questions about their files on their own computer. It runs local LLMs without an internet connection or a server, keeping local conversations on the device. The same app also connects to Claude and ChatGPT, so you can choose local processing or a cloud provider.
sindresorhus.comOn-Device and In-Browser AI
macOS · iOS#Batch processing#Multilingual
Aiko is a paid, native transcription app for macOS, iOS and visionOS that processes speech on your device with OpenAI's Whisper model. It's for people turning meetings, lectures or other recordings into text while keeping the audio local, including sensitive recordings.
1.1KUpdated 4 weeks agoBSD-3-Clause
Android · Browser Extension#Multilingual#Ollama integration#Works offline
Linguist is a free browser translation extension for people who read across languages and want control over where their text goes. Its built-in Bergamot translator processes text on your device without an internet connection. You can also choose an external translation provider or connect a local backend such as Ollama or LibreTranslate.
318Updated 3 weeks agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Quantization
picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
592Updated 5 days agoMIT
Docker#Ollama integration
Ollama Helm Chart packages Ollama for teams that want to run a local LLM service on their own Kubernetes cluster. It's a community-maintained, open source chart under the MIT license, aimed at developers and infrastructure teams managing AI alongside other cluster services.
818Updated 4 years agoApache-2.0
macOS · Windows · Linux · Docker#Home Assistant integration#Works offline
DeepStack is a self-hosted computer vision API for developers adding image analysis to camera systems, home automation or other applications. It runs prebuilt and custom models on your own hardware and works fully offline. Image processing stays on the device or server where you host it, with no cloud service required.
982Updated 1 year ago
macOS · Windows · Linux · Docker#Home Assistant integration#Image-to-image#Multimodal input
CodeProject.AI Server gives developers a shared API for AI tasks that run on their own hardware. It's a self-hosted service for adding image analysis, text processing and generation to applications. Processing stays on the machine running the server, without cloud calls or sending data outside your device or network.
12KUpdated 12 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access
h2oGPT is a self-hosted ChatGPT alternative for people who want to chat with local models and ask questions about their own documents. The project is archived and no longer maintained. It's open source under Apache 2.0, with support for Linux, macOS, Windows and Docker.
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.
857Updated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API
OpenedAI Speech is a self-hosted text-to-speech server for developers who want local speech generation in apps built around OpenAI's speech API. The project is archived and no longer maintained. It's open source under AGPL-3.0, and it generates audio on your own hardware without an OpenAI API key.
3.9KUpdated 2 years agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#Multimodal input
Riffusion is a Python library for generating music and audio on your own hardware using Stable Diffusion. It's for developers and musicians who want to experiment with text-driven sound generation or build it into an app. The hobby project is no longer actively maintained.