Tools tagged with "Quantization"

100+ tools
A browser-based local LLM tool that pools laptop, desktop and phone GPUs for chat and coding. Open source under MIT, with no account required.

544Updated 14 hours agoMIT

macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration

Favicon of Wan2GP

Wan2GP

5 videos
Local AI media generator with a browser interface, support for NVIDIA and AMD GPUs, and select models that run with 6 GB of VRAM.

9.7KUpdated 1 day ago

macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF

Open-source diffusion engine in Python for local image, video and music generation, with low-VRAM inference and model training under Apache 2.0.

13.2KUpdated 2 days agoApache-2.0

#ControlNet#Image-to-image#Inpainting

Favicon of exo

exo

1 video
An open-source local LLM runner for macOS and Linux that splits models across devices and works offline with downloaded models. Apache 2.0 licensed.

47.7KUpdated 1 month agoApache-2.0

macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX

Favicon of vLLM

vLLM

9 videos
An open source LLM serving engine that runs on your hardware, supports NVIDIA and AMD GPUs, and provides an OpenAI-compatible API.

93KUpdated 2 hours agoApache-2.0

macOS · Docker#Batch processing#Distributed execution#GGUF

A Python library for running diffusion models locally with PyTorch, including Stable Diffusion, LoRA adapters and Apple Silicon support. Apache 2.0 licensed.

34.6KUpdated 1 day agoApache-2.0

macOS#ControlNet#Hugging Face integration#Image-to-image

Self-hosted AI image and video generation WebUI with Stable Diffusion, CPU and GPU support, and an Apache 2.0 license for Windows, Linux and macOS.

7.3KUpdated 1 week agoApache-2.0

macOS · Windows · Linux · Docker · Web#ControlNet#Image-to-image#Inpainting

Favicon of llama.cpp

llama.cpp

11 videos
An open source local LLM engine for GGUF models, with CPU and GPU support, a built-in web UI, and an OpenAI-compatible server.

130KUpdated 1 hour agoMIT

Web#Code execution#GGUF#Hugging Face integration

Favicon of ComfyUI

ComfyUI

15 videos
Open source visual AI software for building image, video and audio workflows locally on Windows, macOS and Linux, with optional cloud access.

135.6KUpdated 1 hour agoGPL-3.0

macOS · Windows · Linux · Web#ControlNet#Inpainting#LoRA

Open-source LLM fine-tuning framework under Apache 2.0, with LoRA, QLoRA, multimodal training and inference through vLLM or SGLang.

75.2KUpdated 2 days agoApache-2.0

Web#LoRA#Multimodal input#OpenAI-compatible API

An open-source speech-to-text engine that runs Whisper models locally on desktop and mobile, with CPU-only inference and GPU acceleration. MIT licensed.

54KUpdated 2 days agoMIT

macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference

Favicon of SGLang

SGLang

3 videos
An open-source inference framework for serving language and multimodal models on your own hardware, with an OpenAI-compatible API.

36.7KUpdated 2 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

Favicon of Qwen3

Qwen3

1 video
A language model family with public weights for local CPU or GPU use through Ollama, llama.cpp, and LM Studio, plus server deployment.

27.7KUpdated 9 months ago

#Batch processing#GGUF#Hugging Face integration

Favicon of VibeVoice

VibeVoice

2 videos
Open-source voice AI models for local transcription and speech generation, with MIT licensing, CPU inference, and streaming audio support.

54.5KUpdated 4 weeks agoMIT

#Hugging Face integration#Multilingual#Quantization

Favicon of Unsloth

Unsloth

6 videos
An open-source local LLM app for macOS, Windows and Linux. Run and train models, generate media, and connect coding agents to your hardware.

77KUpdated 22 hours agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Image-to-image

An on-device document search engine that combines keyword and semantic search with local GGUF models, plus MCP access for AI agents. MIT licensed.

30.1KUpdated 3 weeks agoMIT

macOS#GGUF#Hugging Face integration#Hybrid search

A local LLM inference SDK for desktop, mobile and browsers. Runs on CPU or GPU, with offline inference and online Picovoice account validation.

318Updated 3 weeks agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Quantization

Open-source Python library for synthetic data generation and LLM training, with local models, API-based models, caching, and resumable workflows.

1.1KUpdated 2 years agoMIT

#Hugging Face integration#LoRA#Quantization

Local AI music separation software splits songs into stems on Windows, macOS and Linux. MIT licensed, with CPU or CUDA GPU processing.

10.4KUpdated 3 years agoMIT

macOS · Windows · Linux · Docker#Batch processing#Quantization

A local AI coding assistant for VS Code that uses Ollama on your computer or a server you control. Open source under MIT, with no telemetry.

2.1KUpdated 2 years agoMIT

macOS · Windows · VS Code#Multilingual#Ollama integration#Quantization

A multilingual text embedding model that runs offline on phones, laptops and tablets, with open weights and a quantized memory footprint under 200MB.

5.8KUpdated 11 hours agoApache-2.0

#Hugging Face integration#Multilingual#Quantization

Jina’s embedding models encode multilingual text and media for retrieval, with local weights, noncommercial licenses and commercial deployment options.

jina.aiEmbedding and Reranker Models

Docker#GGUF#LoRA#MLX

Open-source model optimization library under Apache 2.0. Compress Hugging Face, PyTorch and ONNX models for TensorRT-LLM, vLLM and SGLang.

5.1KUpdated 22 hours agoApache-2.0

Windows · Docker#Agent Skills#Hugging Face integration#ONNX

Open-source AI training framework built on PyTorch. Train on local CPUs or GPUs, fine-tune HuggingFace models, and serve models on your own server.

11.8KUpdated 4 days agoApache-2.0

Docker#Distributed execution#Hugging Face integration#LoRA

An open-source Python library for model quantization on your own hardware, with PyTorch, TensorFlow and JAX support under Apache 2.0.

2.7KUpdated 1 week agoApache-2.0

Linux · Docker#Hugging Face integration#Quantization

Local LLM quantization library for smaller model weights and inference on your hardware. MIT licensed, with CPU and GPU support; archived and unmaintained.

2.3KUpdated 1 year agoMIT

Linux#Batch processing#GGUF#Hugging Face integration

An open-source LLM fine-tuning tool with a browser interface. Runs on Ubuntu with NVIDIA GPUs or in Docker, under the Apache 2.0 license.

5.2KUpdated 4 days agoApache-2.0

Linux · Docker · Web#Hugging Face integration#LoRA#Quantization

Local AI video generator with downloadable models, NVIDIA GPU support and LoRA fine-tuning. Code and the 2B model use Apache 2.0.

13KUpdated 11 months agoApache-2.0

Windows · Web#Hugging Face integration#LoRA#Multimodal input

A local text-to-image model for Chinese and English prompts, with Diffusers support, a community ComfyUI wrapper, and Apache 2.0 code.

1.1KUpdated 2 years agoApache-2.0

Web#Batch processing#Hugging Face integration#LoRA

A self-hosted ChatGPT alternative that runs Llama 2 and Code Llama locally, with an MIT license and an OpenAI-compatible API.

10.9KUpdated 3 years agoMIT

macOS · Docker · Web#GGUF#llama.cpp backend#OpenAI-compatible API

Local LLM web interface for Windows, macOS and Linux. Use GGUF models, Ollama or cloud APIs, with local chat storage and an Apache 2.0 license.

4.8KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

Self-hosted LLM inference engine for Hugging Face models, with OpenAI-compatible APIs, multimodal support, and CPU or GPU execution under AGPL-3.0.

1.9KUpdated 3 weeks agoAGPL-3.0

macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

An on-device AI SDK that runs text, image and audio models on macOS, Windows and Linux, with GGUF, MLX and an OpenAI-compatible API.

qualcomm/GenieXInference Libraries and Bindings

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

Open-source vision language model for local image and text tasks, with Apache 2.0 licensing, Transformers support, and a small GPU memory footprint.

3.9KUpdated 1 week agoApache-2.0

#Hugging Face integration#LoRA#Multimodal input

An open source on-device AI engine for local LLMs and image models, with iOS, Android, CPU and GPU support under Apache 2.0.

16.2KUpdated 1 day agoApache-2.0

Windows · iOS · Android#Image-to-image#Multimodal input#ONNX

A self-hosted LLM with a 256K context window, vLLM and Transformers support, and research and commercial use under the Jamba Open Model License.

huggingface.coOpen-Weight LLMs

#Hugging Face integration#LoRA#Multilingual