4.7KUpdated 1 day agoApache-2.0
#Batch processing#Distributed execution#OpenAI-compatible API
llm-d is an open-source stack for teams serving large language models on their own Kubernetes clusters. It coordinates model servers such as vLLM and SGLang across multiple machines, with routing and resource management for production traffic. It uses the Apache 2.0 license.
3.1KUpdated 2 months agoApache-2.0
#Home Assistant integration#Voice activity detection#Wake word detection
Willow is a self-hosted voice assistant platform for people who want home automation voice control on their own hardware. It runs on Espressif's ESP32-S3-BOX family and connects to Home Assistant, openHAB, or other services that accept speech results over HTTP. The project is open source under Apache 2.0.
11KUpdated 1 week agoBSD-3-Clause
Windows · Linux · Docker#Batch processing#ONNX
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.
1.9KUpdated 2 months agoMIT
macOS
macmon is a local system monitor for Apple Silicon Macs that reads hardware performance data without administrator privileges. It's for people checking resource use during local AI workloads or other demanding tasks, and developers who need those measurements in their own tools. It runs on macOS and supports M1 through M5 chips.
43.2KUpdated 18 hours agoApache-2.0
Windows#Distributed execution
DeepSpeed is an open-source library for developers and researchers training or running large AI models on their own hardware or compute clusters. It works with PyTorch and focuses on memory use, training speed, and distributing work across GPUs. It's licensed under Apache 2.0.
8.9KUpdated 3 weeks agoApache-2.0
Docker#Batch processing#ControlNet#Distributed execution
BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
2.5KUpdated 1 day ago
macOS · Windows · Linux · Docker#Batch processing#Code execution#Multimodal input
Roboflow Inference is a self-hosted computer vision server for teams building camera and image analysis systems. It runs on your own computer, server, or edge device and combines model predictions with workflows for tracking, counting, measuring, and responding to events. Roboflow also offers hosted servers and a Serverless Cloud API, where processing runs on its infrastructure.
4.6KUpdated 1 week agoApache-2.0
Linux
NVIDIA Container Toolkit lets Docker containers use NVIDIA GPUs on a Linux machine or server. It's for developers and server operators who need GPU acceleration for containerized workloads, including self-hosted AI software. The project is open source under the Apache 2.0 license.
14.8KUpdated 3 hours ago
Docker#Batch processing#Distributed execution#LoRA
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.
798Updated 1 month agoApache-2.0
Linux
NVIDIA DCGM monitors and manages NVIDIA data-center GPUs on your own Linux servers. It's for infrastructure teams running GPU clusters, including those hosting AI workloads, who need to track hardware health, investigate slow jobs and control power use. It supports x86_64 and aarch64 (SBSA) systems.
9.6KUpdated 7 hours agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
2.6KUpdated 1 week agoAGPL-3.0
Linux · Docker
jetson-stats monitors NVIDIA Jetson hardware and gives you control over its power and cooling settings. It runs locally on the board, with a terminal interface called jtop for people developing or running workloads on Jetson devices, including local AI applications.
1.3KUpdated 2 days agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
15.9KUpdated 1 month agoApache-2.0
Web#MLX#Multi-user access
Kubeflow is a self-hosted AI platform for teams that run machine learning workloads on Kubernetes. It brings model development, training and production workflows into a modular stack that can run on a local laptop, on-premises infrastructure or a cloud Kubernetes cluster. It's open source under Apache 2.0.
7.2KUpdated 2 days agoApache-2.0
Windows · Linux · Docker · Web
nvitop is an interactive terminal monitor for NVIDIA GPUs and the processes using them. It runs locally on Linux and Windows and suits people running AI workloads who need to see device usage alongside host process information. The project is open source under Apache 2.0.
8.2KUpdated 1 hour ago
#Distributed execution#Multimodal input#OpenAI-compatible API
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
10.9KUpdated 1 day agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.
9.9KUpdated 13 hours agoApache-2.0
#Distributed execution
Accelerate is a Python library for developers and researchers who write their own PyTorch training loops and want to use the same code on a local machine or a distributed cluster. It handles the hardware-specific work while leaving the training logic under your control.
62.1KUpdated 1 day agoAGPL-3.0
#ONNX
Ultralytics YOLO is an open-source Python computer vision library for developers building applications that analyze images and video on their own hardware. It supports local and edge deployment, including NVIDIA Jetson, Raspberry Pi and mobile phones. A separate hosted platform provides browser-based annotation, cloud GPU training and managed prediction endpoints.
6.1KUpdated 6 hours agoApache-2.0
#Hugging Face integration#ONNX#OpenAI-compatible API
KServe is an open source platform for teams serving LLMs and predictive machine learning models on their own Kubernetes infrastructure. It puts both kinds of workloads under a common serving API, so teams can manage different model frameworks through the same platform. It uses the Apache 2.0 license.
11KUpdated 3 days ago
macOS · Windows · Linux · Docker
nvtop is a terminal monitor that shows activity across multiple GPUs and accelerators in an interface familiar to htop users. It's useful for people running local LLM workloads or managing compute servers who want to see which processes are using their hardware and how much GPU memory they consume. It runs locally on Linux and is open source under GPLv3 or later.
30KUpdated 1 month agoMIT
macOS · Windows · Linux · Android · Docker · Web#LM Studio integration#MCP#Multimodal input
PicoClaw is a self-hosted personal AI assistant with messaging channels, a browser interface and tools for file operations, code execution and web research. It can handle scheduled reminders and recurring tasks through its cron tool. The Go application runs on your own hardware.