407Updated 2 days agoMIT
macOS#MLX#Multimodal input#OpenAI-compatible API
Slotstream runs Qwen3.8-Flash-Next on Apple Silicon Macs that don't have enough RAM to hold the whole model. It's aimed at people with 16 to 64 GB of memory who want local chat, image questions or a model backend for coding agents. Most model weights stay on the SSD, while frequently used expert networks stay in memory. The full model remains available.
7.4KUpdated 12 hours agoApache-2.0
Linux#Agent Skills#OpenAI-compatible API#Prompt versioning
Reef is self-hosted infrastructure for developers who want AI agents to improve through feedback on actual interactions. It connects inference and learning with versioned deployment, so an agent can update its model weights or its prompts, rules, and skills while continuing to serve requests. It's open source under Apache 2.0.
544Updated 14 hours agoMIT
macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
47.7KUpdated 1 month agoApache-2.0
macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX
exo is a local LLM runner that combines your devices into a cluster, letting you use models too large for one machine's memory. It's for people who want to run large models on their own hardware and developers connecting existing AI clients to local inference. It runs on macOS and Linux under the Apache 2.0 license.
93KUpdated 1 hour agoApache-2.0
macOS · Docker#Batch processing#Distributed execution#GGUF
vLLM is an open source engine for serving large language models on hardware you control. It suits developers and teams that need to handle many requests through an API while making efficient use of memory and compute. It's licensed under Apache 2.0 and can run with GPUs or on a CPU.
130KUpdated 40 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
182KUpdated 16 hours agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input
Ollama runs language models on your own computer or server. It provides a command-line runner and a local API for people building AI applications or connecting existing tools to models they host themselves. The software is distributed under the MIT license.
36.7KUpdated 1 hour agoApache-2.0
#Batch processing#Distributed execution#LoRA
SGLang is a self-hosted inference framework for teams that need to serve language and multimodal models on their own hardware. It runs on a single GPU or across distributed clusters and exposes an OpenAI-compatible API. The project is open source under the Apache 2.0 license.
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
592Updated 5 days agoMIT
Docker#Ollama integration
Ollama Helm Chart packages Ollama for teams that want to run a local LLM service on their own Kubernetes cluster. It's a community-maintained, open source chart under the MIT license, aimed at developers and infrastructure teams managing AI alongside other cluster services.
818Updated 4 years agoApache-2.0
macOS · Windows · Linux · Docker#Home Assistant integration#Works offline
DeepStack is a self-hosted computer vision API for developers adding image analysis to camera systems, home automation or other applications. It runs prebuilt and custom models on your own hardware and works fully offline. Image processing stays on the device or server where you host it, with no cloud service required.
982Updated 1 year ago
macOS · Windows · Linux · Docker#Home Assistant integration#Image-to-image#Multimodal input
CodeProject.AI Server gives developers a shared API for AI tasks that run on their own hardware. It's a self-hosted service for adding image analysis, text processing and generation to applications. Processing stays on the machine running the server, without cloud calls or sending data outside your device or network.
857Updated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API
OpenedAI Speech is a self-hosted text-to-speech server for developers who want local speech generation in apps built around OpenAI's speech API. The project is archived and no longer maintained. It's open source under AGPL-3.0, and it generates audio on your own hardware without an OpenAI API key.
8.4KUpdated 20 hours agoMIT
#Agent Skills#Batch processing#Guardrails
OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
4.8KUpdated 8 months ago
Seldon Core 2 is an AI model serving framework for teams running production machine learning and LLM applications on Kubernetes. It can run on your own infrastructure or in a cloud environment. Its focus is managing individual models and connected applications within the same deployment system.
10.9KUpdated 6 months agoApache-2.0
Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.
2.6KUpdated 20 hours agoApache-2.0
Web#OpenAI-compatible API#Prompt caching
vLLM Production Stack is an open source inference stack for teams serving LLMs on their own Kubernetes GPU clusters. It brings request routing and monitoring around vLLM, so applications can move from one serving instance to a distributed deployment without changing their code. It requires a GPU-enabled Kubernetes environment.
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
8.3KUpdated 3 years agoApache-2.0
macOS · Windows · Linux · Docker · Web#Multi-user access#Role-based access
CompreFace is a self-hosted face recognition service for developers who want to add facial identification to an application without building or training their own machine learning system. It runs as a Docker-based server on your hardware or in a cloud deployment you manage. It's free and open source under the Apache 2.0 license.
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.
44KUpdated 18 hours agoApache-2.0
#Batch processing#Hugging Face integration#ONNX
Ray Serve is a self-hosted Python library for developers building inference APIs that combine models with application logic. It runs on a laptop, on-premise servers, Kubernetes, or cloud infrastructure you choose. It's open source under Apache 2.0.
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.
29.9KUpdated 3 weeks ago
macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization
InsightFace is a face analysis toolkit for developers and teams building identity verification, access control, or face editing software. The code uses the MIT license. Its Python tools and self-hosted recognition server run inference on your own hardware. It also offers commercial models and API access for face swapping and deepfake detection.
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
9.8KUpdated 5 months agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration
PowerInfer is a local LLM inference engine for developers and researchers who want to run large models on a PC with a consumer GPU. It splits work between the CPU and GPU to reduce GPU memory demands and data transfers. The code is open source under the MIT license.
10.6KUpdated 2 years agoMIT
macOS · Windows · Linux · Docker#Hugging Face integration
Petals lets developers and researchers use large language models that won't fit on a single consumer GPU by sharing the work across a network of machines. It supports text generation and fine-tuning from a desktop computer or Google Colab. Each participant holds part of the model, while other computers handle the remaining parts.
11.2KUpdated 5 months agoApache-2.0
macOS · iOS · Web#Hugging Face integration#MLX#Quantization
Moshi is a voice AI model and dialogue framework that can listen while it speaks. It processes speech directly, retaining information such as emotion and non-verbal cues that a text transcription can miss. It's aimed at researchers and developers building spoken AI applications, with local inference and self-hosted server options.
1.4KUpdated 2 days agoAGPL-3.0
Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.
3.1KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization
Distributed Llama runs a local LLM across several computers, sharing both the computation and the model's memory use. It's for people who want to use their own networked hardware for inference rather than keep the entire workload on one machine. The C++ project is open source under the MIT license.
4.7KUpdated 23 hours agoApache-2.0
#Batch processing#Distributed execution#OpenAI-compatible API
llm-d is an open-source stack for teams serving large language models on their own Kubernetes clusters. It coordinates model servers such as vLLM and SGLang across multiple machines, with routing and resource management for production traffic. It uses the Apache 2.0 license.
5.5KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API
Kokoro-FastAPI runs the Kokoro-82M speech model on your own machine or server and exposes an OpenAI-compatible speech API. It's for developers adding local text-to-speech to assistants, reading apps or audiobook workflows. Speech generation runs locally, and the API doesn't require an OpenAI account.