Infrastructure for Running AI on Your Hardware

Monitor GPUs, serve models on Kubernetes and run small models on Jetson boards, Raspberry Pi and phones.

Subcategories

59 tools
Favicon of llm-d

llm-d

1 video
An open-source LLM inference stack for self-hosted Kubernetes clusters, with vLLM and SGLang backends and support for GPUs, TPUs, XPUs and CPUs.

4.7KUpdated 1 day agoApache-2.0

#Batch processing#Distributed execution#OpenAI-compatible API

Favicon of Willow

Willow

1 video
An open-source voice assistant for ESP32-S3-BOX hardware, with offline commands, self-hosted speech recognition, and Home Assistant integration.

3.1KUpdated 2 months agoApache-2.0

#Home Assistant integration#Voice activity detection#Wake word detection

Self-hosted AI inference server runs TensorRT, PyTorch and ONNX models on GPUs or CPUs, with dynamic batching and a BSD-3-Clause license.

11KUpdated 1 week agoBSD-3-Clause

Windows · Linux · Docker#Batch processing#ONNX

A local system monitor for Apple Silicon Macs that tracks power, temperature and memory without root access, with JSON and Prometheus output. It's MIT licensed.

1.9KUpdated 2 months agoMIT

macOS

Favicon of DeepSpeed

DeepSpeed

1 video
Open-source PyTorch optimization library for distributed training and inference, with Apache 2.0 licensing and support for NVIDIA and AMD GPUs.

43.2KUpdated 18 hours agoApache-2.0

Windows#Distributed execution

An open-source Python framework for AI model serving. Build inference APIs and multi-model pipelines locally or deploy with Docker under Apache 2.0.

8.9KUpdated 3 weeks agoApache-2.0

Docker#Batch processing#ControlNet#Distributed execution

An open-weight vision model for image questions, captions and object detection. Run it locally or use hosted inference and fine-tuning.

10.1KUpdated 5 months agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline

Self-hosted computer vision server for images and video, with Docker support, NVIDIA GPU acceleration, and optional Roboflow hosted compute.

2.5KUpdated 1 day ago

macOS · Windows · Linux · Docker#Batch processing#Code execution#Multimodal input

An open-source container toolkit that gives Docker workloads access to NVIDIA GPUs on Linux. Requires the NVIDIA driver, but not the host CUDA Toolkit.

4.6KUpdated 1 week agoApache-2.0

Linux

A self-hosted LLM inference library built on PyTorch for NVIDIA GPUs, with a Python API, OpenAI-compatible serving, and multi-node support.

14.8KUpdated 3 hours ago

Docker#Batch processing#Distributed execution#LoRA

GPU management and monitoring software for NVIDIA data-center GPUs on Linux, with Kubernetes telemetry and an Apache 2.0 open-source core.

798Updated 1 month agoApache-2.0

Linux

Self-hosted AI model serving platform for Linux, Windows and macOS. Run language, speech and image models through an OpenAI-compatible API under Apache 2.0.

9.6KUpdated 7 hours agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input

A local NVIDIA Jetson monitoring tool with a terminal interface, Python API and Docker support. Open source under AGPL-3.0.

2.6KUpdated 1 week agoAGPL-3.0

Linux · Docker

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

1.3KUpdated 2 days agoApache-2.0

Web#LoRA#Multimodal input#Ollama integration

Favicon of Kubeflow

Kubeflow

1 video
An open source AI platform for Kubernetes that supports self-hosted ML workflows, distributed training and model management under Apache 2.0.

15.9KUpdated 1 month agoApache-2.0

Web#MLX#Multi-user access

An open-source NVIDIA GPU process monitor for Linux and Windows, with interactive terminal views, Python APIs and Grafana dashboard support.

7.2KUpdated 2 days agoApache-2.0

Windows · Linux · Docker · Web

A self-hosted inference framework that coordinates NVIDIA GPU clusters with vLLM, SGLang or TensorRT-LLM and exposes an OpenAI-compatible API.

8.2KUpdated 1 hour ago

#Distributed execution#Multimodal input#OpenAI-compatible API

Favicon of OpenVINO

OpenVINO

2 videos
Open-source AI inference toolkit for local or self-hosted deployment on Linux, Windows and macOS, with CPU, Intel GPU and NPU support.

10.9KUpdated 1 day agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX

A PyTorch library for training and inference on local machines or clusters, with CPU, GPU and TPU support. Open source under Apache 2.0.

9.9KUpdated 13 hours agoApache-2.0

#Distributed execution

Open-source computer vision library for local detection, segmentation and tracking, with AGPL-3.0 licensing and exports to ONNX, TensorRT and CoreML.

62.1KUpdated 1 day agoAGPL-3.0

#ONNX

Favicon of KServe

KServe

3 videos
Self-hosted AI model serving platform for Kubernetes. Serve LLMs and predictive models with vLLM, Hugging Face support and an OpenAI-compatible API.

6.1KUpdated 6 hours agoApache-2.0

#Hugging Face integration#ONNX#OpenAI-compatible API

A local GPU and accelerator monitor for Linux, with process lists, usage charts and support for NVIDIA, AMD, Intel and dedicated AI hardware.

11KUpdated 3 days ago

macOS · Windows · Linux · Docker

Run a personal assistant with Ollama, messaging channels, a browser interface and scheduled tasks on your computer, server or Android device.

30KUpdated 1 month agoMIT

macOS · Windows · Linux · Android · Docker · Web#LM Studio integration#MCP#Multimodal input