Tools tagged with "GGUF"

54 tools
Favicon of Handy

Handy

7 videos
Free, open source desktop dictation for Windows, macOS and Linux. It uses local Whisper or Parakeet models and keeps your voice off the cloud.

32.5KUpdated 3 days agoMIT

macOS · Windows · Linux#GGUF#Hugging Face integration#Voice activity detection

Favicon of Wan2GP

Wan2GP

5 videos
Local AI media generator with a browser interface, support for NVIDIA and AMD GPUs, and select models that run with 6 GB of VRAM.

9.7KUpdated 1 day ago

macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF

Favicon of vLLM

vLLM

9 videos
An open source LLM serving engine that runs on your hardware, supports NVIDIA and AMD GPUs, and provides an OpenAI-compatible API.

93KUpdated 2 hours agoApache-2.0

macOS · Docker#Batch processing#Distributed execution#GGUF

A local AI assistant for iPhone, iPad, Mac, and Vision Pro that runs models on device, searches your files, and can block external web connections.

noemaai.comChat With Your Documents

macOS · iOS#GGUF#MCP#MLX

Favicon of llama.cpp

llama.cpp

11 videos
An open source local LLM engine for GGUF models, with CPU and GPU support, a built-in web UI, and an OpenAI-compatible server.

130KUpdated 1 hour agoMIT

Web#Code execution#GGUF#Hugging Face integration

Favicon of InvokeAI

InvokeAI

2 videos
A self-hosted AI image generator that runs on Windows, macOS, Linux and Docker, with canvas editing, visual workflows and an Apache 2.0 license.

28.3KUpdated 3 days agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF

Favicon of Ollama

Ollama

31 videos
Open-source local LLM runner for macOS, Windows, Linux and Docker, with optional cloud models and coding agent integrations.

182KUpdated 17 hours agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input

Favicon of Qwen3

Qwen3

1 video
A language model family with public weights for local CPU or GPU use through Ollama, llama.cpp, and LM Studio, plus server deployment.

27.7KUpdated 9 months ago

#Batch processing#GGUF#Hugging Face integration

Favicon of Unsloth

Unsloth

6 videos
An open-source local LLM app for macOS, Windows and Linux. Run and train models, generate media, and connect coding agents to your hardware.

77KUpdated 22 hours agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Image-to-image

Run local models on your phone or computer for chat, document analysis, image generation and speech transcription. Connect to your own model server.

3.2KUpdated 4 days agoMIT

macOS · Windows · Linux · iOS · Android#GGUF#Human approval#LM Studio integration

An on-device document search engine that combines keyword and semantic search with local GGUF models, plus MCP access for AI agents. MIT licensed.

30.1KUpdated 3 weeks agoMIT

macOS#GGUF#Hugging Face integration#Hybrid search

Local AI chat app for macOS that runs GGUF models offline, searches files on-device, and also connects to Claude, ChatGPT, and OpenAI-compatible endpoints.

recurse.chatChat With Your Documents

macOS#GGUF#Hugging Face integration#OpenAI-compatible API

Self-hosted LLM inference for Kubernetes with NVIDIA, AMD and Apple Silicon support, OpenAI-compatible APIs, and an Apache 2.0 license.

223Updated 17 hours agoApache-2.0

macOS · Linux#GGUF#Git integration#Guardrails

Jina’s embedding models encode multilingual text and media for retrieval, with local weights, noncommercial licenses and commercial deployment options.

jina.aiEmbedding and Reranker Models

Docker#GGUF#LoRA#MLX

Local LLM quantization library for smaller model weights and inference on your hardware. MIT licensed, with CPU and GPU support; archived and unmaintained.

2.3KUpdated 1 year agoMIT

Linux#Batch processing#GGUF#Hugging Face integration

A local AI companion for Windows, macOS and Linux with voice chat, Live2D avatars and camera input. Runs offline with local models or connects to cloud APIs.

14KUpdated 5 months ago

macOS · Windows · Linux · Web#GGUF#LM Studio integration#MCP

A self-hosted ChatGPT alternative that runs Llama 2 and Code Llama locally, with an MIT license and an OpenAI-compatible API.

10.9KUpdated 3 years agoMIT

macOS · Docker · Web#GGUF#llama.cpp backend#OpenAI-compatible API

An offline AI assistant for Android and iOS that runs models on your phone, keeps chats encrypted on-device, and offers optional hosted models.

layla-network.aiAI Characters and Roleplay

iOS · Android#Code execution#GGUF#llama.cpp backend

Local LLM web interface for Windows, macOS and Linux. Use GGUF models, Ollama or cloud APIs, with local chat storage and an Apache 2.0 license.

4.8KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

An on-device AI SDK that runs text, image and audio models on macOS, Windows and Linux, with GGUF, MLX and an OpenAI-compatible API.

qualcomm/GenieXInference Libraries and Bindings

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

A family of AI models you can run offline with Ollama, llama.cpp or LM Studio, with open weights and training data for building specialized agents.

2.1KUpdated 3 weeks agoApache-2.0

Linux#GGUF#Guardrails#Hugging Face integration

An open source LLM evaluation framework for local models and hosted APIs, with GGUF, Hugging Face transformers and llama.cpp support. MIT licensed.

14.1KUpdated 2 weeks agoMIT

macOS#Batch processing#GGUF#Hugging Face integration

Favicon of Amuse

Amuse

1 video
Local AI creation app for Windows with image, video, audio and text models, GGUF and ONNX support, and acceleration for NVIDIA, AMD and Intel GPUs.

606Updated 4 days agoApache-2.0

Windows#ControlNet#GGUF#Inpainting

A local LLM family for chat, coding and multilingual tasks, with GGUF and Hugging Face formats for CPU or GPU use and support for llama.cpp and MLX.

127Updated 12 months ago

macOS#GGUF#Hugging Face integration#llama.cpp backend

Local vision-language models for image and video understanding, with Apache 2.0 code, mobile deployment and support for Ollama and llama.cpp.

26.5KUpdated 3 weeks agoApache-2.0

macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend

Language models for English, Korean and Spanish, with reasoning and tool use. Includes an on-device model and GGUF, GPTQ and AWQ formats.

107Updated 1 year ago

#GGUF#Hugging Face integration#Multilingual

Local LLM acceleration library for Intel CPUs, GPUs and NPUs. Runs on Windows and Linux, integrates with Ollama and llama.cpp, and is archived.

8.9KUpdated 8 months agoApache-2.0

Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration

A Python toolkit for local LLM compression and inference on Linux, macOS and Windows, with GPTQ, AWQ, GGUF and integrations for vLLM and SGLang.

1.3KUpdated 1 day ago

macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA

Open-source speech recognition model for Mandarin, Cantonese, English, Japanese and Korean. Runs locally on CPU or GPU under the MIT license.

9.4KUpdated 3 weeks agoMIT

Docker#Batch processing#GGUF#Hugging Face integration

An open-source local LLM stack manager that connects Ollama, llama.cpp and AI apps through Docker Compose. Includes a CLI and companion app.

3.2KUpdated 5 days agoApache-2.0

macOS · Linux · Docker#GGUF#llama.cpp backend#MCP

A local LLM chat app for Android that runs GGUF models on-device through llama.cpp. Open source under the Apache 2.0 license.

894Updated 3 months agoApache-2.0

Android#GGUF#llama.cpp backend

An open-source React Native library that runs GGUF models on iOS and Android through llama.cpp, with GPU acceleration and image and audio understanding.

1KUpdated 3 days agoMIT

iOS · Android#GGUF#llama.cpp backend#Multilingual

A local LLM integration for Home Assistant with voice, chat and AI automations. Runs on Raspberry Pi without a GPU and supports Ollama and llama.cpp.

1.4KUpdated 3 days ago

#GGUF#Home Assistant integration#Hugging Face integration

A local LLM inference engine for sparse models, with CPU and GPU support on Linux and Windows. Open source under MIT, with CPU-only support on Apple Silicon.

9.8KUpdated 5 months agoMIT

macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration

Local OCR model that extracts plain or formatted text from images, supports multi-page recognition, and works with Hugging Face Transformers.

8.2KUpdated 2 years ago

#Batch processing#GGUF#Hugging Face integration

An open source Android AI chat app that runs GGUF models offline through llama.cpp and connects to remote providers, including Ollama and OpenAI.

2.7KUpdated 2 weeks agoMIT

Android#GGUF#Hugging Face integration#llama.cpp backend