223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
6.8KUpdated 4 weeks agoMIT
Windows · Linux
ROCm is AMD's open-source GPU computing platform for developers running AI training, inference and scientific workloads on their own hardware or servers. It supports selected Linux and Windows configurations on AMD Instinct, Radeon and Ryzen AI devices. Check the version-specific GPU, operating-system, driver and firmware compatibility matrix before installing. It's the software foundation for applications that need AMD GPU acceleration, including local LLM workloads.
4.6KUpdated 4 years agoMIT
macOS
asitop is a terminal hardware monitor for Apple Silicon Macs, with separate views of CPU, GPU and Apple Neural Engine activity. Its README targets Apple Silicon on macOS Monterey; hardware-counter availability can differ on newer macOS releases. It runs locally and suits people who want to watch hardware use during demanding workloads, including local AI inference. It's free and open source under the MIT license.
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
10.7KUpdated 21 hours agoApache-2.0
#Code execution#Distributed execution#Multi-user access
SkyPilot is an open-source system for AI teams that need to run training, inference and development workloads across their own clusters and cloud accounts. It brings Kubernetes, Slurm and cloud compute under one interface, so teams can move jobs between providers without rewriting their workload code.
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
4.4KUpdated 2 weeks agoMIT
gpustat is a local command-line GPU monitor for people running AI models or other workloads on NVIDIA hardware. It puts GPU activity and the processes using each device into a compact terminal view, useful when you need to check memory use or see who’s occupying a shared GPU. It's open source under the MIT license.
1.9KUpdated 2 months agoMIT
macOS
macmon is a local system monitor for Apple Silicon Macs that reads hardware performance data without administrator privileges. It's for people checking resource use during local AI workloads or other demanding tasks, and developers who need those measurements in their own tools. It runs on macOS and supports M1 through M5 chips.
4.6KUpdated 1 week agoApache-2.0
Linux
NVIDIA Container Toolkit lets Docker containers use NVIDIA GPUs on a Linux machine or server. It's for developers and server operators who need GPU acceleration for containerized workloads, including self-hosted AI software. The project is open source under the Apache 2.0 license.
798Updated 1 month agoApache-2.0
Linux
NVIDIA DCGM monitors and manages NVIDIA data-center GPUs on your own Linux servers. It's for infrastructure teams running GPU clusters, including those hosting AI workloads, who need to track hardware health, investigate slow jobs and control power use. It supports x86_64 and aarch64 (SBSA) systems.
2.6KUpdated 1 week agoAGPL-3.0
Linux · Docker
jetson-stats monitors NVIDIA Jetson hardware and gives you control over its power and cooling settings. It runs locally on the board, with a terminal interface called jtop for people developing or running workloads on Jetson devices, including local AI applications.
7.2KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web
nvitop is an interactive terminal monitor for NVIDIA GPUs and the processes using them. It runs locally on Linux and Windows and suits people running AI workloads who need to see device usage alongside host process information. The project is open source under Apache 2.0.
8.2KUpdated 20 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
11KUpdated 3 days ago
macOS · Windows · Linux · Docker
nvtop is a terminal monitor that shows activity across multiple GPUs and accelerators in an interface familiar to htop users. It's useful for people running local LLM workloads or managing compute servers who want to see which processes are using their hardware and how much GPU memory they consume. It runs locally on Linux and is open source under GPLv3 or later.