Tools tagged with "Distributed execution"

43 tools
A browser-based local LLM tool that pools laptop, desktop and phone GPUs for chat and coding. Open source under MIT, with no account required.

544Updated 14 hours agoMIT

macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration

Favicon of exo

exo

1 video
An open-source local LLM runner for macOS and Linux that splits models across devices and works offline with downloaded models. Apache 2.0 licensed.

47.7KUpdated 1 month agoApache-2.0

macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX

Favicon of vLLM

vLLM

9 videos
An open source LLM serving engine that runs on your hardware, supports NVIDIA and AMD GPUs, and provides an OpenAI-compatible API.

93KUpdated 2 hours agoApache-2.0

macOS · Docker#Batch processing#Distributed execution#GGUF

Favicon of MLX

MLX

2 videos
An open-source machine learning array framework with NumPy-style APIs, shared CPU and GPU memory on Apple silicon, and Linux CPU and CUDA backends.

28.6KUpdated 1 day agoMIT

macOS · Linux#Distributed execution#LoRA

Favicon of SGLang

SGLang

3 videos
An open-source inference framework for serving language and multimodal models on your own hardware, with an OpenAI-compatible API.

36.7KUpdated 2 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

An open-source Python library for preparing LLM training data locally or on Slurm and Ray clusters, with filtering, deduplication and generation.

3.4KUpdated 1 day agoApache-2.0

#Batch processing#Distributed execution#Multilingual

Open-source AI training framework built on PyTorch. Train on local CPUs or GPUs, fine-tune HuggingFace models, and serve models on your own server.

11.8KUpdated 4 days agoApache-2.0

Docker#Distributed execution#Hugging Face integration#LoRA

Self-hosted LLM inference engine for Hugging Face models, with OpenAI-compatible APIs, multimodal support, and CPU or GPU execution under AGPL-3.0.

1.9KUpdated 3 weeks agoAGPL-3.0

macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

A self-hosted Kubernetes operator that deploys Hugging Face models with vLLM, manages GPU capacity, and runs fine-tuning and document retrieval services.

1KUpdated 1 week ago

#Distributed execution#Hugging Face integration#LoRA

Local LLM acceleration library for Intel CPUs, GPUs and NPUs. Runs on Windows and Linux, integrates with Ollama and llama.cpp, and is archived.

8.9KUpdated 8 months agoApache-2.0

Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration

A self-hosted LLM inference server under Apache 2.0, with Docker deployment, multi-GPU support and an OpenAI-compatible chat API. The project is archived.

10.9KUpdated 6 months agoApache-2.0

Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

Open-source PyTorch training framework for pretraining and fine-tuning models on your own CPUs or GPUs, with Apache 2.0 licensing.

31.4KUpdated 1 week agoApache-2.0

macOS#Distributed execution#ONNX

Open-source LLM pretraining library built on PyTorch for custom datasets, with distributed NVIDIA GPU training. Licensed under Apache 2.0.

2.8KUpdated 1 week agoApache-2.0

Linux#Distributed execution#Hugging Face integration

An open-source LLM evaluation platform under Apache 2.0 that compares Hugging Face models and cloud APIs across reasoning, coding, safety and other tasks.

7.5KUpdated 2 days agoApache-2.0

#Distributed execution#Hugging Face integration#OpenAI-compatible API

Open-source AI training and inference framework for your own GPU hardware, with distributed parallelism and Apache 2.0 licensing.

41.4KUpdated 3 days agoApache-2.0

#Distributed execution

An open-source diffusion model trainer that splits large models across GPUs. Uses DeepSpeed and supports Windows through WSL 2 under GPL-3.0.

2KUpdated 2 days agoGPL-3.0

Windows · Linux#Distributed execution#LoRA#Quantization

Open-source AI compute management software that runs jobs across your Kubernetes and Slurm clusters or cloud accounts, using GPUs, TPUs and CPUs.

10.7KUpdated 21 hours agoApache-2.0

#Code execution#Distributed execution#Multi-user access

Open-source LLM serving infrastructure for Kubernetes with multi-node inference, demand-based autoscaling, LoRA management and vLLM integration.

5.1KUpdated 21 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

A self-hosted diffusion model trainer with a web UI, LoRA and full fine-tuning, NVIDIA, AMD and Apple Silicon support, and an AGPL-3.0 license.

2.9KUpdated 20 hours agoAGPL-3.0

macOS · Docker · Web#ControlNet#Distributed execution#Human approval

An open-source LLM training and deployment platform under Apache 2.0. Build specialized models on your own infrastructure or use its hosted service.

9.4KUpdated 2 days agoApache-2.0

Docker#Distributed execution#LoRA#Multimodal input

An open-source LLM training toolkit that turns documents into specialist datasets, with offline generation on macOS and Linux and optional cloud compute.

1.9KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization

A self-hosted Python toolkit for curating AI training data across text, images, video and audio, with NVIDIA GPU support and an Apache 2.0 license.

1.8KUpdated 19 hours agoApache-2.0

Linux · Docker#Distributed execution#Hugging Face integration#Multilingual

Open-source LLM training software for your own GPU servers, with PyTorch distributed training, NVIDIA and AMD support, and a BSD-3-Clause license.

5.8KUpdated 19 hours agoBSD-3-Clause

Linux#Distributed execution#Hugging Face integration

A self-hosted LLM API server that runs ExLlamaV3 models on your hardware, with OpenAI-compatible endpoints and an AGPL-3.0 license.

1.4KUpdated 2 days agoAGPL-3.0

Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

Open-source AI data processing framework for local machines and Ray clusters, with multimodal cleaning, deduplication and Apache 2.0 licensing.

7.1KUpdated 2 days agoApache-2.0

Docker#Batch processing#Distributed execution#Multimodal input

An open source local LLM runner that splits inference and memory across your computers. Runs on Linux, macOS and Windows under the MIT license.

3.1KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization

Favicon of llm-d

llm-d

1 video
An open-source LLM inference stack for self-hosted Kubernetes clusters, with vLLM and SGLang backends and support for GPUs, TPUs, XPUs and CPUs.

4.7KUpdated 24 hours agoApache-2.0

#Batch processing#Distributed execution#OpenAI-compatible API

Favicon of DeepSpeed

DeepSpeed

1 video
Open-source PyTorch optimization library for distributed training and inference, with Apache 2.0 licensing and support for NVIDIA and AMD GPUs.

43.2KUpdated 1 day agoApache-2.0

Windows#Distributed execution

An open source engine for running ONNX models on Windows, macOS, Linux, mobile devices and the web, with CPU, GPU and NPU support.

22KUpdated 20 minutes agoMIT

macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX

Favicon of TRL

TRL

1 video
Open-source Python library for LLM fine-tuning on your own GPUs, with Transformers support, preference training and Apache 2.0 licensing.

19.4KUpdated 1 day agoApache-2.0

#Distributed execution#LoRA#Quantization

An open-source local LLM training framework with LoRA, multimodal support, and deployment through vLLM, SGLang or LMDeploy. Apache 2.0 licensed.

15.8KUpdated 2 days agoApache-2.0

Web#Distributed execution#Hugging Face integration#LoRA

An open-source Python framework for AI model serving. Build inference APIs and multi-model pipelines locally or deploy with Docker under Apache 2.0.

8.9KUpdated 3 weeks agoApache-2.0

Docker#Batch processing#ControlNet#Distributed execution

An open source RLHF framework for training models on your own NVIDIA GPUs, with HuggingFace model support and Ray, vLLM and DeepSpeed backends.

10.1KUpdated 2 weeks agoApache-2.0

Docker#Distributed execution#Hugging Face integration#LoRA

Open-source LLM serving toolkit for your own GPU servers, with quantization, text and vision models, and OpenAI-compatible APIs. Apache 2.0 licensed.

8.1KUpdated 3 days agoApache-2.0

#Batch processing#Distributed execution#Hugging Face integration

A self-hosted LLM inference library built on PyTorch for NVIDIA GPUs, with a Python API, OpenAI-compatible serving, and multi-node support.

14.7KUpdated 21 hours ago

Docker#Batch processing#Distributed execution#LoRA

LLM training framework with ready-made research scripts, NVIDIA GPU parallelism, and Hugging Face checkpoint conversion through Megatron Bridge.

18KUpdated 1 day ago

Docker#Distributed execution#Hugging Face integration#Quantization