Tools tagged with "Quantization"

100+ tools
AI video generation model you can run on your own GPUs, with text-to-video and image-to-video support, ComfyUI integration and downloadable weights.

12.6KUpdated 3 months ago

Web#Multimodal input#Quantization

Open-source Python library for compressing local or Hugging Face LLM checkpoints, with Apache 2.0 licensing and vLLM-compatible output.

3.8KUpdated 1 day agoApache-2.0

#Hugging Face integration#Quantization

A self-hosted AI video generation model with ComfyUI and Diffusers support, image animation, video editing, and fine-tuning tools.

11KUpdated 9 months agoApache-2.0

#Hugging Face integration#LoRA#Multimodal input

Open-source Rust ML framework for running models locally on CPUs, NVIDIA GPUs or in browsers, with Apache 2.0 licensing and quantized LLM support.

21.1KUpdated 2 days agoApache-2.0

macOS · Web#GGUF#Hugging Face integration#Multilingual

Self-hosted embedding and reranking API with MIT licensing, Hugging Face models, and CPU, NVIDIA, AMD and Apple MPS support.

2.9KUpdated 6 months agoMIT

macOS · Docker#Batch processing#Hugging Face integration#Multimodal input

Run AI models through Docker Desktop, Docker Engine or a standalone binary, with local inference and OpenAI and Ollama compatible APIs.

655Updated 2 days agoApache-2.0

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

An open-source Python library for local speech transcription with Whisper models. It runs on CPUs or NVIDIA GPUs and uses CTranslate2.

25.6KUpdated 3 hours agoMIT

#Batch processing#Hugging Face integration#Quantization

Open-weight local LLMs under Apache 2.0, with 20B and 120B models that work with Ollama, LM Studio and vLLM.

20.4KUpdated 2 months agoApache-2.0

macOS · Linux#Hugging Face integration#LM Studio integration#Ollama integration