Tools tagged with "Quantization"

100+ tools
A family of AI models you can run offline with Ollama, llama.cpp or LM Studio, with open weights and training data for building specialized agents.

2.1KUpdated 3 weeks agoApache-2.0

Linux#GGUF#Guardrails#Hugging Face integration

Open-source LLM for self-hosted AI agents, with thinking and direct-response modes, MIT licensing, and support for vLLM and SGLang.

huggingface.coCoding Models

#Hugging Face integration#LoRA#Multilingual

An open source LLM evaluation framework for local models and hosted APIs, with GGUF, Hugging Face transformers and llama.cpp support. MIT licensed.

14.1KUpdated 2 weeks agoMIT

macOS#Batch processing#GGUF#Hugging Face integration

A local LLM family for chat, coding and multilingual tasks, with GGUF and Hugging Face formats for CPU or GPU use and support for llama.cpp and MLX.

127Updated 12 months ago

macOS#GGUF#Hugging Face integration#llama.cpp backend

Local vision-language models for image and video understanding, with Apache 2.0 code, mobile deployment and support for Ollama and llama.cpp.

26.5KUpdated 3 weeks agoApache-2.0

macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend

A self-hosted language model series for coding and tool use, with Base and Instruct variants and a hosted OpenAI/Anthropic-compatible API.

11.1KUpdated 11 months ago

#Hugging Face integration#Quantization#Tool calling

A self-hosted vision-language model project for image chat, with a local Gradio interface, GPU inference and Apache 2.0 code.

25KUpdated 2 years agoApache-2.0

macOS · Web#LoRA#Multimodal input#Quantization

A self-hosted language model for reasoning and agent tasks, with fast and slow thinking modes, a 256K context window, and Transformers and vLLM support.

820Updated 1 year ago

#Quantization#Tool calling

Language models for English, Korean and Spanish, with reasoning and tool use. Includes an on-device model and GGUF, GPTQ and AWQ formats.

107Updated 1 year ago

#GGUF#Hugging Face integration#Multilingual

Code completion models that run locally on CPU or GPUs, with Hugging Face Transformers support and LoRA fine-tuning.

2.1KUpdated 3 years agoApache-2.0

#Hugging Face integration#LoRA#Quantization

A bilingual local LLM family for English and Chinese, with chat and base models, version-specific licensing, and quantized variants for consumer GPUs.

7.8KUpdated 2 years agoApache-2.0

Docker#Hugging Face integration#llama.cpp backend#Multilingual

An open-source image and video generation framework with 4K text-to-image models, laptop GPU support, and ComfyUI and Diffusers integrations.

9.2KUpdated 2 weeks agoApache-2.0

#ControlNet#LoRA#Multimodal input

A local image captioning model with open weights, Apache 2.0 code, SFW and NSFW coverage, and support for ComfyUI and vLLM.

1.3KUpdated 7 months agoApache-2.0

Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API

Open-source machine learning library in C/C++ with CPU, GPU, NPU and browser backends, quantization support, and an MIT license.

15.4KUpdated 6 days agoMIT

Web#Quantization

Local LLM acceleration library for Intel CPUs, GPUs and NPUs. Runs on Windows and Linux, integrates with Ollama and llama.cpp, and is archived.

8.9KUpdated 8 months agoApache-2.0

Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration

A local LLM inference library for Windows and Linux with NVIDIA GPUs, MIT licensing, GPTQ and EXL2 support, and an OpenAI-compatible API through TabbyAPI.

4.6KUpdated 7 months agoMIT

Windows · Linux#Batch processing#Quantization#Speculative decoding

A self-hosted LLM inference server under Apache 2.0, with Docker deployment, multi-GPU support and an OpenAI-compatible chat API. The project is archived.

10.9KUpdated 6 months agoApache-2.0

Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

An open-source on-device AI framework for Android, iOS, desktop and web, with model conversion from PyTorch, TensorFlow and JAX.

3.5KUpdated 20 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input

An AI inference framework for Android, iOS, desktop and browsers, with CPU and Vulkan GPU support and PyTorch and ONNX model conversion.

23.9KUpdated 6 days ago

macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization

An open-source model quantization library for local LLMs and vision models, with Apache 2.0 licensing and Hugging Face Transformers integration.

960Updated 7 months agoApache-2.0

#Hugging Face integration#LoRA#Quantization

A local LLM app for iPhone, iPad, and Mac that works offline after model download, keeps chats on-device, and connects to Siri and Apple Shortcuts.

privatellm.appAutomation and No-Code AI

macOS · iOS#Multilingual#Quantization#Works offline

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

5.1KUpdated 22 hours ago

macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API

Local LoRA training scripts for image and video models, with Windows and Linux support, NVIDIA GPU memory savings, and multi-GPU training.

2.1KUpdated 3 days ago

Windows · Linux#LoRA#Quantization

A free, open-source meme search engine that runs locally in Docker, with AI image descriptions, semantic search and optional OpenAI-compatible vision APIs.

752Updated 3 weeks agoApache-2.0

Linux · Docker · Web · Browser Extension#Batch processing#OpenAI-compatible API#Quantization

An open-source diffusion model trainer that splits large models across GPUs. Uses DeepSpeed and supports Windows through WSL 2 under GPL-3.0.

2KUpdated 2 days agoGPL-3.0

Windows · Linux#Distributed execution#LoRA#Quantization

A Python toolkit for local LLM compression and inference on Linux, macOS and Windows, with GPTQ, AWQ, GGUF and integrations for vLLM and SGLang.

1.3KUpdated 1 day ago

macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA

A local LLM chat app for iOS, iPadOS, macOS and visionOS. Runs on Apple silicon, works offline and saves chat history locally. Licensed under MIT.

2.3KUpdated 1 year agoMIT

macOS · iOS#MLX#Quantization#Works offline

An open-source LLM training and deployment platform under Apache 2.0. Build specialized models on your own infrastructure or use its hosted service.

9.4KUpdated 2 days agoApache-2.0

Docker#Distributed execution#LoRA#Multimodal input

Open-source Python tools optimize Hugging Face models for local, edge and cloud hardware, with ONNX Runtime, OpenVINO and TensorRT-LLM integrations.

3.5KUpdated 6 days agoApache-2.0

#Hugging Face integration#ONNX#Quantization

An open-source dictation app that types speech into your active window using local Whisper models, CPU or NVIDIA processing, or OpenAI's API.

1.1KUpdated 2 years agoGPL-3.0

macOS · Windows · Linux#Multilingual#OpenAI-compatible API#Quantization

An open-source React Native library that runs GGUF models on iOS and Android through llama.cpp, with GPU acceleration and image and audio understanding.

1KUpdated 3 days agoMIT

iOS · Android#GGUF#llama.cpp backend#Multilingual

A self-hosted text-to-speech server that connects Piper to Home Assistant through Wyoming, with custom ONNX voices and optional NVIDIA GPU support.

214Updated 3 weeks agoMIT

Linux · Docker · Web#Home Assistant integration#Hugging Face integration#Multilingual

An open-source LLM training toolkit that turns documents into specialist datasets, with offline generation on macOS and Linux and optional cloud compute.

1.9KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization

Face analysis toolkit for self-hosted recognition on CPU or NVIDIA GPU, local video face redaction, and commercially licensed models.

29.9KUpdated 3 weeks ago

macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization

A local LLM inference engine for sparse models, with CPU and GPU support on Linux and Windows. Open source under MIT, with CPU-only support on Apple Silicon.

9.8KUpdated 5 months agoMIT

macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration

An on-device AI engine runs automation models locally on phones and tiny devices, with speech, vision and optional cloud routing.

6.1KUpdated 5 days ago

macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization