Infrastructure for Running AI on Your Hardware

Monitor GPUs, serve models on Kubernetes and run small models on Jetson boards, Raspberry Pi and phones.

Subcategories

58 tools
Favicon of exo

exo

1 video
An open-source local LLM runner for macOS and Linux that splits models across devices and works offline with downloaded models. Apache 2.0 licensed.

47.7KUpdated 1 month agoApache-2.0

macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX

Favicon of vLLM

vLLM

9 videos
An open source LLM serving engine that runs on your hardware, supports NVIDIA and AMD GPUs, and provides an OpenAI-compatible API.

93KUpdated 1 hour agoApache-2.0

macOS · Docker#Batch processing#Distributed execution#GGUF

An open-weight local LLM family from Google DeepMind for phones, PCs and servers, with Ollama and LM Studio support and an Apache 2.0 JAX library.

5.8KUpdated 2 days agoApache-2.0

Android#LM Studio integration#LoRA#Multilingual

Favicon of SGLang

SGLang

3 videos
An open-source inference framework for serving language and multimodal models on your own hardware, with an OpenAI-compatible API.

36.7KUpdated 1 hour agoApache-2.0

#Batch processing#Distributed execution#LoRA

A self-hosted AI runtime with an OpenAI-compatible API. It runs models on CPUs or GPUs and keeps inference on your own hardware.

49.3KUpdated 2 hours agoMIT

macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend

Self-hosted AI gateway routes OpenAI-compatible requests to cloud providers or vLLM, with centralized credentials and an Apache 2.0 license.

2.2KUpdated 20 hours agoApache-2.0

Docker#LLM tracing#MCP#Multi-user access

Self-hosted LLM inference for Kubernetes with NVIDIA, AMD and Apple Silicon support, OpenAI-compatible APIs, and an Apache 2.0 license.

223Updated 16 hours agoApache-2.0

macOS · Linux#GGUF#Git integration#Guardrails

A community Helm chart that deploys Ollama on Kubernetes, with CPU or NVIDIA and AMD GPU support. Open source under the MIT license.

592Updated 5 days agoMIT

Docker#Ollama integration

A self-hosted voice satellite for Home Assistant with local wake word detection. Runs on Raspberry Pi hardware under the MIT license; archived and unmaintained.

1.2KUpdated 8 months agoMIT

Linux · Docker#Code execution#Home Assistant integration#Voice activity detection

An open-source GPU computing platform for AI training and inference on AMD hardware, with Linux and Windows support and PyTorch, JAX, vLLM and SGLang.

6.8KUpdated 4 weeks agoMIT

Windows · Linux

A terminal performance monitor for Apple Silicon Macs that tracks CPU, GPU, memory and power use. Runs locally on macOS and uses the MIT license.

4.6KUpdated 4 years agoMIT

macOS

An open-source computer vision API that runs offline on your hardware, with face recognition, object detection and support for custom models.

818Updated 4 years agoApache-2.0

macOS · Windows · Linux · Docker#Home Assistant integration#Works offline

An open source on-device AI engine for local LLMs and image models, with iOS, Android, CPU and GPU support under Apache 2.0.

16.2KUpdated 1 day agoApache-2.0

Windows · iOS · Android#Image-to-image#Multimodal input#ONNX

A self-hosted Kubernetes operator that deploys Hugging Face models with vLLM, manages GPU capacity, and runs fine-tuning and document retrieval services.

1KUpdated 7 days ago

#Distributed execution#Hugging Face integration#LoRA

A self-hosted AI serving framework that deploys models and pipelines on Kubernetes, on-premises or in the cloud, under the Business Source License.

4.8KUpdated 8 months ago

An open-source on-device AI framework for Android, iOS, desktop and web, with model conversion from PyTorch, TensorFlow and JAX.

3.5KUpdated 19 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input

An AI inference framework for Android, iOS, desktop and browsers, with CPU and Vulkan GPU support and PyTorch and ONNX model conversion.

23.9KUpdated 6 days ago

macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

5.1KUpdated 20 hours ago

macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API

Self-hosted API and AI gateway with Apache 2.0 licensing, Kubernetes support, and routing across OpenAI, Anthropic, Gemini and other LLM providers.

44.2KUpdated 2 days agoApache-2.0

#MCP

Open-source AI training and inference framework for your own GPU hardware, with distributed parallelism and Apache 2.0 licensing.

41.4KUpdated 3 days agoApache-2.0

#Distributed execution

Open-source AI compute management software that runs jobs across your Kubernetes and Slurm clusters or cloud accounts, using GPUs, TPUs and CPUs.

10.7KUpdated 21 hours agoApache-2.0

#Code execution#Distributed execution#Multi-user access

A self-hosted LLM serving stack built on vLLM, with an OpenAI-compatible API, request routing and GPU cluster monitoring. Licensed under Apache 2.0.

2.6KUpdated 20 hours agoApache-2.0

Web#OpenAI-compatible API#Prompt caching

Self-hosted LLM observability and evaluation platform under Apache 2.0. Monitor Ollama, vLLM and coding agents with portable OpenTelemetry traces.

2.8KUpdated 1 day agoApache-2.0

Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning

A self-hosted AI gateway under Apache 2.0 that runs in Docker, manages LLM APIs and MCP servers, and supports Kubernetes inference routing.

9.5KUpdated 1 day agoApache-2.0

Docker · Web#Guardrails#MCP#Tool calling

Open-source LLM serving infrastructure for Kubernetes with multi-node inference, demand-based autoscaling, LoRA management and vLLM integration.

5.1KUpdated 20 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

An open-source toolkit for on-device AI across mobile, web and desktop. Apache 2.0 licensed; input stays local, while API metrics go to Google.

37.1KUpdated 19 hours agoApache-2.0

iOS · Android · Web

An open-source LLM server that runs locally or in the cloud, with OpenAI-compatible APIs, a chat UI, and an Apache 2.0 license.

12.5KUpdated 4 months agoApache-2.0

Docker · Web#Hugging Face integration#OpenAI-compatible API

A self-hosted model serving library for Python and LLM APIs. Run it on a laptop or cluster with request batching, streaming, and CPU or GPU resources.

44KUpdated 18 hours agoApache-2.0

#Batch processing#Hugging Face integration#ONNX

Self-hostable MLOps software for experiment tracking, pipelines, dataset versioning and model serving, with an Apache 2.0 Python SDK.

6.9KUpdated 2 days agoApache-2.0

Docker · Web#Code execution#Git integration#Multi-user access

Self-hosted AI compute orchestration under MPL-2.0 for training and inference on GPU clouds, Kubernetes, VMs and bare-metal servers.

2.3KUpdated 1 day agoMPL-2.0

macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access

A Docker container build system for local AI on NVIDIA Jetson, with CUDA support and packages for Ollama, llama.cpp, PyTorch and ROS.

4.9KUpdated 3 months ago

Linux · Docker#llama.cpp backend#Multimodal input#Ollama integration

An open-source voice assistant platform for Linux, Raspberry Pi and Docker, with offline speech plugins and support for local OpenAI-compatible LLM servers.

295Updated 2 days agoApache-2.0

Linux · Docker#OpenAI-compatible API#Wake word detection#Works offline

A command-line GPU monitor that runs locally on NVIDIA hardware, shows memory use and running processes, and provides JSON output under the MIT license.

4.4KUpdated 2 weeks agoMIT

An on-device AI engine runs automation models locally on phones and tiny devices, with speech, vision and optional cloud routing.

6.1KUpdated 5 days ago

macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization

An open source local LLM runner that splits inference and memory across your computers. Runs on Linux, macOS and Windows under the MIT license.

3.1KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization

Local LLM toolkit that runs language and multimodal models on Rockchip NPUs, with model conversion, quantization and C/C++ interfaces.

1.7KUpdated 2 days ago

Linux#Multimodal input#Quantization