Tools tagged with "Quantization"

100+ tools
Local OCR and document parsing software that converts PDFs and images to Markdown, recognizes tables and formulas, and runs on NVIDIA GPUs.

6.7KUpdated 2 months agoApache-2.0

Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual

An open-source voice AI framework that processes speech directly, with local inference on Mac and iPhone through MLX and self-hosted server backends.

11.2KUpdated 5 months agoApache-2.0

macOS · iOS · Web#Hugging Face integration#MLX#Quantization

An open source local LLM runner that splits inference and memory across your computers. Runs on Linux, macOS and Windows under the MIT license.

3.1KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization

Local LLM toolkit that runs language and multimodal models on Rockchip NPUs, with model conversion, quantization and C/C++ interfaces.

1.7KUpdated 2 days ago

Linux#Multimodal input#Quantization

A free local AI image generator that runs Stable Diffusion on Windows, Linux and macOS, with a browser interface for prompts, image editing and ControlNet.

10.5KUpdated 3 weeks ago

macOS · Windows · Linux · Web#Batch processing#ControlNet#Image-to-image

PyTorch quantization library that reduces model memory use and speeds training and inference on your hardware, with CPU, GPU and mobile deployment support.

3KUpdated 5 days ago

Linux · iOS#Hugging Face integration#LoRA#Quantization

Local LLM desktop app for Windows, macOS and Linux. Run GGUF models offline or connect other apps through OpenAI- and Anthropic-compatible APIs.

47.7KUpdated 1 month agoAGPL-3.0

macOS · Windows · Linux · Docker · Web#GGUF#llama.cpp backend#LoRA

Local AI image generation UI built on Stable Diffusion WebUI, with Flux, GGUF and LoRA support. Open source under AGPL-3.0.

13KUpdated 1 year agoAGPL-3.0

Windows · Web#ControlNet#GGUF#Image-to-image

Favicon of TRL

TRL

1 video
Open-source Python library for LLM fine-tuning on your own GPUs, with Transformers support, preference training and Apache 2.0 licensing.

19.4KUpdated 1 day agoApache-2.0

#Distributed execution#LoRA#Quantization

JavaScript client for Ollama in Node.js and browsers. Connect to local models or Ollama's cloud with an open-source, MIT-licensed library.

4.4KUpdated 2 days agoMIT

Web#LoRA#Multimodal input#Ollama integration

An open-source local LLM training framework with LoRA, multimodal support, and deployment through vLLM, SGLang or LMDeploy. Apache 2.0 licensed.

15.8KUpdated 2 days agoApache-2.0

Web#Distributed execution#Hugging Face integration#LoRA

A local AI inference library for C++ and Python that runs Transformer models on CPUs and GPUs, with quantization to reduce memory use. MIT licensed.

4.7KUpdated 5 days agoMIT

#Batch processing#Multilingual#Quantization

Favicon of WhisperX

WhisperX

1 video
Open source speech-to-text software that runs locally, aligns transcripts word by word, and can label speakers.

24.3KUpdated 4 days agoBSD-2-Clause

macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual

An open-source local LLM runner for Linux, macOS and Windows via WSL2, with Podman or Docker isolation and llama.cpp or vLLM inference.

3.1KUpdated 1 day agoMIT

macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend

Favicon of koboldcpp

koboldcpp

1 video
Local LLM runner for GGUF and GGML models on Windows, macOS and Linux, with CPU or GPU support, a browser UI and an AGPL-3.0 license.

11.9KUpdated 4 days agoAGPL-3.0

macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

An open-source JavaScript library that runs AI models in your browser using ONNX Runtime, with CPU execution and optional WebGPU acceleration.

16.3KUpdated 1 week agoApache-2.0

Web#Hugging Face integration#Image-to-image#Multilingual

Self-hosted speech-to-text server for Home Assistant with Whisper and other backends. Runs on CPU or NVIDIA GPUs and works offline with downloaded models.

390Updated 1 day agoMIT

Docker#Home Assistant integration#Hugging Face integration#Multilingual

A mobile AI assistant that runs GGUF models on iOS and Android. Core chat works offline after a model download and needs no account.

8.5KUpdated 2 days agoMIT

iOS · Android#GGUF#Hugging Face integration#llama.cpp backend

Open-source Python toolkit for training and serving LLMs on your own hardware, with Apache 2.0 licensing, quantization and multi-GPU support.

13.7KUpdated 3 weeks agoApache-2.0

#Hugging Face integration#LoRA#Quantization

An open source RLHF framework for training models on your own NVIDIA GPUs, with HuggingFace model support and Ray, vLLM and DeepSpeed backends.

10.1KUpdated 2 weeks agoApache-2.0

Docker#Distributed execution#Hugging Face integration#LoRA

A local LLM inference engine with OpenAI and Anthropic-compatible APIs. Runs on macOS, Linux and Windows with CPU, CUDA or Apple Silicon support.

7.7KUpdated 5 days agoMIT

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration

An open-source LLM fine-tuning library for single-GPU and distributed training, with LoRA and QLoRA support. Development is no longer active.

5.8KUpdated 5 months agoBSD-3-Clause

#Hugging Face integration#LoRA#Quantization

A local LLM family for developers and researchers, with downloadable weights, text and vision models, and custom licensing for research and commercial use.

7.7KUpdated 12 months ago

#Hugging Face integration#Multimodal input#Quantization

Open-source LLM serving toolkit for your own GPU servers, with quantization, text and vision models, and OpenAI-compatible APIs. Apache 2.0 licensed.

8.1KUpdated 3 days agoApache-2.0

#Batch processing#Distributed execution#Hugging Face integration

A self-hosted LLM inference library built on PyTorch for NVIDIA GPUs, with a Python API, OpenAI-compatible serving, and multi-node support.

14.7KUpdated 22 hours ago

Docker#Batch processing#Distributed execution#LoRA

A PyTorch quantization library that reduces LLM memory use for inference and fine-tuning with 8-bit optimizers, LLM.int8() and QLoRA. MIT licensed.

8.5KUpdated 4 weeks agoMIT

macOS · Windows · Linux#LoRA#Quantization

An open-source local LLM framework that splits work across CPUs and GPUs, with SGLang serving and LlamaFactory fine-tuning under Apache 2.0.

19.5KUpdated 1 week agoApache-2.0

Docker#LoRA#Multimodal input#Prompt caching

LLM training framework with ready-made research scripts, NVIDIA GPU parallelism, and Hugging Face checkpoint conversion through Megatron Bridge.

18KUpdated 1 day ago

Docker#Distributed execution#Hugging Face integration#Quantization

An open-source Android LLM chat app that runs GGUF models on-device through llama.cpp or connects to Ollama, OpenAI and Claude. Licensed under AGPL-3.0.

2.8KUpdated 1 week agoAGPL-3.0

Android#GGUF#llama.cpp backend#Ollama integration

Favicon of GPT4All

GPT4All

1 video
An open-source local AI chatbot for Windows, macOS and Linux. Run models without a GPU or cloud API, and chat privately with your documents.

77.4KUpdated 1 year agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API

Self-hosted subtitle generator runs Whisper locally on CPU or NVIDIA GPU and connects to Bazarr, Plex, Jellyfin, Emby and Tautulli. MIT licensed.

1.5KUpdated 2 months agoMIT

Docker#Batch processing#Multilingual#OpenAI-compatible API

Favicon of MLX LM

MLX LM

2 videos
A Python package for local LLM inference and fine-tuning on Apple Silicon, built on MLX with Hugging Face model support and an MIT license.

7.2KUpdated 1 day agoMIT

macOS#Batch processing#Distributed execution#Hugging Face integration

Favicon of PEFT

PEFT

1 video
An open-source Python library for adapting LLMs and diffusion models with less memory and storage, under Apache 2.0.

21.7KUpdated 1 day agoApache-2.0

macOS#Distributed execution#Hugging Face integration#LoRA

An MIT-licensed translation app for Windows, macOS and Linux that runs Marian and Bergamot models locally and connects to Chrome and Firefox.

631Updated 2 years agoMIT

macOS · Windows · Linux · Web · Browser Extension#Batch processing#Multilingual#Quantization

Favicon of Axolotl

Axolotl

1 video
Apache 2.0 LLM fine-tuning framework for local or cloud GPUs, supporting NVIDIA, AMD, LoRA, QLoRA and multimodal training.

12.5KUpdated 2 days agoApache-2.0

Docker#Distributed execution#Hugging Face integration#LoRA

A Python library for local embeddings and reranking, using ONNX Runtime with CPU or GPU support. Open source under Apache 2.0.

3.2KUpdated 1 day agoApache-2.0

#Batch processing#Multilingual#ONNX