544Updated 14 hours agoMIT
macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
47.7KUpdated 1 month agoApache-2.0
macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX
exo is a local LLM runner that combines your devices into a cluster, letting you use models too large for one machine's memory. It's for people who want to run large models on their own hardware and developers connecting existing AI clients to local inference. It runs on macOS and Linux under the Apache 2.0 license.
93KUpdated 2 hours agoApache-2.0
macOS · Docker#Batch processing#Distributed execution#GGUF
vLLM is an open source engine for serving large language models on hardware you control. It suits developers and teams that need to handle many requests through an API while making efficient use of memory and compute. It's licensed under Apache 2.0 and can run with GPUs or on a CPU.
28.6KUpdated 1 day agoMIT
macOS · Linux#Distributed execution#LoRA
MLX is a machine learning array framework for researchers and developers building models on their own hardware. Its distinctive feature on Apple silicon is shared CPU and GPU memory: both processors can work on the same arrays without copying data between them. It's open source under the MIT license.
36.7KUpdated 2 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
SGLang is a self-hosted inference framework for teams that need to serve language and multimodal models on their own hardware. It runs on a single GPU or across distributed clusters and exposes an OpenAI-compatible API. The project is open source under the Apache 2.0 license.
3.4KUpdated 1 day agoApache-2.0
#Batch processing#Distributed execution#Multilingual
DataTrove is an open-source Python library for teams preparing large text datasets, including LLM training corpora. It runs on your own machine or on Slurm and Ray clusters, with processing steps that carry across those environments. It uses the Apache 2.0 license.
11.8KUpdated 4 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
Ludwig is an open-source Python framework for developers and researchers who want to train custom AI models on their own hardware. A YAML file describes the model and training pipeline, while Ludwig handles preprocessing, training and evaluation. It uses the Apache 2.0 license. Install the Python package with the optional LLM dependencies for fine-tuning; current source requires Python 3.12 or later.
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
1KUpdated 1 week ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
8.9KUpdated 8 months agoApache-2.0
Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration
Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.
10.9KUpdated 6 months agoApache-2.0
Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.
31.4KUpdated 1 week agoApache-2.0
macOS#Distributed execution#ONNX
PyTorch Lightning is a Python framework for researchers and developers who want to pretrain or fine-tune models on their own hardware or in the cloud. It handles repetitive training code while leaving model logic under your control. The framework is open source under Apache 2.0.
2.8KUpdated 1 week agoApache-2.0
Linux#Distributed execution#Hugging Face integration
Nanotron is a Python library for researchers and developers who want to pretrain language models on their own datasets and GPU infrastructure. Built on PyTorch, it supports NVIDIA CUDA GPUs and training across multiple servers with Slurm. It's open source under Apache 2.0.
7.5KUpdated 2 days agoApache-2.0
#Distributed execution#Hugging Face integration#OpenAI-compatible API
OpenCompass is an open-source LLM evaluation platform for researchers, model developers and teams comparing models for their applications. Its Python framework evaluates local models and cloud APIs within the same experiment, so teams can compare candidates on shared benchmarks. It uses the Apache 2.0 license.
41.4KUpdated 3 days agoApache-2.0
#Distributed execution
Colossal-AI is a Python framework for developers and researchers training or serving large AI models on their own GPU hardware. It addresses the memory and computing demands of models that are difficult to fit on a single GPU, with tools for distributing work across a cluster. It's open source under Apache 2.0.
2KUpdated 2 days agoGPL-3.0
Windows · Linux#Distributed execution#LoRA#Quantization
diffusion-pipe is a local diffusion model training tool for people fine-tuning image and video models on their own GPU hardware. Its main distinction is that it can divide a model across several GPUs when it won't fit on one, while also distributing training work across GPUs. The Python project is open source under GPL-3.0 and uses DeepSpeed.
10.7KUpdated 21 hours agoApache-2.0
#Code execution#Distributed execution#Multi-user access
SkyPilot is an open-source system for AI teams that need to run training, inference and development workloads across their own clusters and cloud accounts. It brings Kubernetes, Slurm and cloud compute under one interface, so teams can move jobs between providers without rewriting their workload code.
5.1KUpdated 21 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
2.9KUpdated 20 hours agoAGPL-3.0
macOS · Docker · Web#ControlNet#Distributed execution#Human approval
SimpleTuner is an open-source toolkit for fine-tuning image, video and audio generation models on your own hardware or GPU servers. It's for creators and researchers adapting models to their datasets, and teams sharing training infrastructure. A web dashboard manages training jobs.
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.
1.8KUpdated 19 hours agoApache-2.0
Linux · Docker#Distributed execution#Hugging Face integration#Multilingual
NeMo Curator is an open source Python toolkit for ML engineers and data teams preparing AI training datasets on their own hardware. It handles text, images, video and audio, with reusable pipelines that can run on a laptop or scale across a multi-node Ray cluster. NVIDIA uses it to prepare data for Nemotron models.
5.8KUpdated 19 hours agoBSD-3-Clause
Linux#Distributed execution#Hugging Face integration
torchtitan is an open-source training platform for researchers and developers building generative AI models on their own GPU machines or server clusters. It uses PyTorch's distributed training tools and keeps the model code relatively simple when spreading work across GPUs. The Python codebase has extension points and replaceable components for experiments with model architectures and training infrastructure.
1.4KUpdated 2 days agoAGPL-3.0
Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
3.1KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization
Distributed Llama runs a local LLM across several computers, sharing both the computation and the model's memory use. It's for people who want to use their own networked hardware for inference rather than keep the entire workload on one machine. The C++ project is open source under the MIT license.
4.7KUpdated 24 hours agoApache-2.0
#Batch processing#Distributed execution#OpenAI-compatible API
llm-d is an open-source stack for teams serving large language models on their own Kubernetes clusters. It coordinates model servers such as vLLM and SGLang across multiple machines, with routing and resource management for production traffic. It uses the Apache 2.0 license.
43.2KUpdated 1 day agoApache-2.0
Windows#Distributed execution
DeepSpeed is an open-source library for developers and researchers training or running large AI models on their own hardware or compute clusters. It works with PyTorch and focuses on memory use, training speed, and distributing work across GPUs. It's licensed under Apache 2.0.
22KUpdated 20 minutes agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
19.4KUpdated 1 day agoApache-2.0
#Distributed execution#LoRA#Quantization
TRL is a Python library for developers and researchers who want to adapt foundation models on their own hardware. It builds on Hugging Face Transformers and covers supervised fine-tuning, reinforcement learning and training from preference feedback. It's open source under Apache 2.0.
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
8.9KUpdated 3 weeks agoApache-2.0
Docker#Batch processing#ControlNet#Distributed execution
BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.
10.1KUpdated 2 weeks agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
OpenRLHF is a self-hosted Python framework for researchers and teams training language models with human feedback or custom rewards. It runs on your own NVIDIA GPU hardware, with Docker support and distributed training across servers. It's open source under Apache 2.0.
8.1KUpdated 3 days agoApache-2.0
#Batch processing#Distributed execution#Hugging Face integration
LMDeploy is an open-source toolkit for developers serving language and vision-language models on their own hardware. It combines model compression with inference and self-hosted APIs, so teams can use it for batch processing or as the model backend for an application. It uses the Apache 2.0 license.
14.7KUpdated 21 hours ago
Docker#Batch processing#Distributed execution#LoRA
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.
18KUpdated 1 day ago
Docker#Distributed execution#Hugging Face integration#Quantization
Megatron-LM is a Python framework for research teams training large language models on NVIDIA GPU infrastructure. It pairs ready-made training scripts with Megatron Core, a library developers can use to build their own training systems. Its focus is distributed training, with benchmarks on H100 clusters spanning thousands of GPUs.