9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.
1.7KUpdated 2 days agoApache-2.0
#Batch processing#Code execution#Multimodal input
Curator is a Python library for developers preparing LLM training datasets or extracting structured records from existing data. It supports local inference through Ollama and vLLM alongside cloud model APIs, so the same data pipeline can use models on your hardware or a hosted provider. It's open source under Apache 2.0.
3.5KUpdated 6 days agoApache-2.0
#Hugging Face integration#ONNX#Quantization
Optimum is a collection of Python packages for developers who want to train or run Hugging Face models more efficiently on specific hardware. It extends Transformers, Diffusers, TIMM and Sentence Transformers, with integrations for local machines, mobile and edge devices, and cloud accelerators. It's open source under Apache 2.0.
38.6KUpdated 2 months agoMIT
Windows · Linux · Web#Hugging Face integration#ONNX#Voice conversion
RVC WebUI is a local AI voice conversion tool for people who want to train a custom voice, change the voice in a recording, or use a live voice changer. It runs on Windows and Linux, including Ubuntu servers, with a browser interface for training and conversion and a separate interface for live use. It's free and open source under the MIT license.
34.7KUpdated 2 days agoApache-2.0
Detectron2 is an open-source Python library for developers and researchers building computer vision applications. It provides algorithms for locating objects in images and segmenting image regions, with support for training models and building research projects on top of the library. Facebook AI Research developed it as the successor to Detectron and maskrcnn-benchmark.
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.
9.2KUpdated 7 days agoApache-2.0
Docker#Image-to-image#Inpainting#Multimodal input
ModelScope combines a hosted model and dataset hub with a Python library you can run locally. It's for developers and researchers who want to use AI models in their own applications, fine-tune them on their own data, or compare their performance. The library is open source under Apache 2.0.
58.3KUpdated 3 months agoMIT
macOS#Code execution#Tool calling
nanochat is an MIT-licensed toolkit for training your own LLM and chatting with it on hardware you control. It's aimed at researchers and developers who want to study or modify the full training pipeline, with a small Python codebase built on PyTorch.
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.
1.8KUpdated 20 hours agoApache-2.0
Linux · Docker#Distributed execution#Hugging Face integration#Multilingual
NeMo Curator is an open source Python toolkit for ML engineers and data teams preparing AI training datasets on their own hardware. It handles text, images, video and audio, with reusable pipelines that can run on a laptop or scale across a multi-node Ray cluster. NVIDIA uses it to prepare data for Nemotron models.
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
16.2KUpdated 19 hours agoGPL-3.0
macOS · Windows · Linux#Multimodal input
LabelMe is a desktop image annotation app for people preparing computer vision datasets. It combines manual drawing with AI assistance for outlining objects and creating labels from text. It runs on 64-bit macOS, Windows and Linux..
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
17.8KUpdated 2 weeks agoApache-2.0
#Code execution#Hugging Face integration#Human approval
CAMEL is an open-source Python framework for developers and researchers building systems where AI agents work together. Its focus is on agent roles, communication, and behavior across extended tasks, with applications in synthetic training data, task automation, and simulated societies. It uses the Apache 2.0 license.
10.6KUpdated 2 years agoMIT
macOS · Windows · Linux · Docker#Hugging Face integration
Petals lets developers and researchers use large language models that won't fit on a single consumer GPU by sharing the work across a network of machines. It supports text generation and fine-tuning from a desktop computer or Google Colab. Each participant holds part of the model, while other computers handle the remaining parts.
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
7.2KUpdated 1 month agoApache-2.0
macOS · Web#Works offline
TensorBoard is a browser-based toolkit for inspecting TensorFlow experiments on your own machine or server. It's for researchers and ML developers who need to understand training behavior, compare runs, and investigate model performance. It works entirely offline, including behind a corporate firewall or in a datacenter, so experiment data can stay within your own environment.
3.4KUpdated 10 months agoApache-2.0
#Structured output
Distilabel is an open-source Python framework for engineers building datasets to train or evaluate AI models. It pairs synthetic data generation with LLM feedback, so a pipeline can create examples and judge their quality. It uses the Apache 2.0 license.
33KUpdated 3 years agoApache-2.0
MMDetection is a Python toolkit for researchers and developers building object detection and image segmentation systems. Part of OpenMMLab, it combines ready-made model architectures with interchangeable components for custom models. It's open source under Apache 2.0.
5.8KUpdated 19 hours agoBSD-3-Clause
Linux#Distributed execution#Hugging Face integration
torchtitan is an open-source training platform for researchers and developers building generative AI models on their own GPU machines or server clusters. It uses PyTorch's distributed training tools and keeps the model code relatively simple when spreading work across GPUs. The Python codebase has extension points and replaceable components for experiments with model architectures and training infrastructure.
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
1.7KUpdated 2 days ago
Linux#Multimodal input#Quantization
RKLLM is a software stack for developers building local AI applications on Rockchip hardware. It uses the chip's neural processing unit (NPU) to run language and multimodal models on development boards, with support for the RK3588, RK3576, RK3562 and RV1126B series.
5.8KUpdated 4 years agoApache-2.0
LayoutParser is an open-source Python library for developers and researchers who need to detect page structure in document images and turn OCR output into structured data. Its pretrained deep learning models share a common interface, so you can work with models trained on different document datasets without rewriting the surrounding pipeline.
16.8KUpdated 1 day agoMIT
Docker · Web#Hugging Face integration#Multi-user access#ONNX
CVAT is a browser-based data annotation platform for teams building computer vision datasets. Its open-source Community edition runs on your own infrastructure with Docker and uses the MIT license. CVAT Online is hosted by CVAT, while the Enterprise offering runs in an organization's own cloud or internal environment.
3KUpdated 5 days ago
Linux · iOS#Hugging Face integration#LoRA#Quantization
TorchAO is a PyTorch library for developers who want to train or run models on their own hardware with less memory and faster computation. It reduces the precision of model weights and activations, with options for language models and image or video generation. Its PyTorch integration works with torch.compile and FSDP2 across most Hugging Face PyTorch models.
47.7KUpdated 1 month agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#GGUF#llama.cpp backend#LoRA
text-generation-webui, also called TextGen, runs language models on your own hardware through a desktop app or a self-hosted browser interface. It's for people who want private chat and writing tools, and developers who need a local model API. It works offline without telemetry; web search and page fetching use the internet.
43.2KUpdated 1 day agoApache-2.0
Windows#Distributed execution
DeepSpeed is an open-source library for developers and researchers training or running large AI models on their own hardware or compute clusters. It works with PyTorch and focuses on memory use, training speed, and distributing work across GPUs. It's licensed under Apache 2.0.
22KUpdated 28 minutes agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
19.4KUpdated 1 day agoApache-2.0
#Distributed execution#LoRA#Quantization
TRL is a Python library for developers and researchers who want to adapt foundation models on their own hardware. It builds on Hugging Face Transformers and covers supervised fine-tuning, reinforcement learning and training from preference feedback. It's open source under Apache 2.0.
165.2KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Web#Batch processing#Code execution#Image-to-image
Stable Diffusion web UI (AUTOMATIC1111) is a browser interface for generating and editing images with models running on your own hardware. It's for artists and anyone who wants control over prompts, models and image variations. The software is open source under AGPL-3.0.
12.6KUpdated 3 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#ControlNet#Hugging Face integration#LoRA
Kohya's GUI lets you train and fine-tune image generation models on your own GPU-equipped computer through a browser interface. It's for artists and model makers who want to teach a model a particular style or subject while controlling the training settings. The interface builds on Kohya's Stable Diffusion training scripts, with a command-line interface available too.
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
10.8KUpdated 8 months agoMIT
Windows · Docker · Web#Multi-user access#Multilingual
Doccano is a self-hosted text annotation tool for machine learning practitioners who need labeled training or evaluation data. It runs on your own machine or server, with a browser interface and Docker support. The software is open source under the MIT license.
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
4.7KUpdated 5 days agoMIT
#Batch processing#Multilingual#Quantization
CTranslate2 is an open-source C++ and Python library for developers running Transformer models on their own hardware or servers. It handles translation, text generation, text encoding and speech recognition. Its custom runtime focuses on reducing inference time and memory use compared with general-purpose deep learning frameworks.