LLM Inference Engines and Libraries

The engines and bindings that run models inside your own code: llama.cpp, MLX, ONNX Runtime and Hugging Face Transformers.

77 tools
Favicon of Slotstream

Slotstream

1 video
Local LLM runner for Apple Silicon Macs that runs Qwen3.8-Flash-Next from SSD. Works offline after download and connects to coding agents and chat apps.

407Updated 2 days agoMIT

macOS#MLX#Multimodal input#OpenAI-compatible API

Open-source diffusion engine in Python for local image, video and music generation, with low-VRAM inference and model training under Apache 2.0.

13.2KUpdated 2 days agoApache-2.0

#ControlNet#Image-to-image#Inpainting

A Python library for running diffusion models locally with PyTorch, including Stable Diffusion, LoRA adapters and Apple Silicon support. Apache 2.0 licensed.

34.6KUpdated 1 day agoApache-2.0

macOS#ControlNet#Hugging Face integration#Image-to-image

Favicon of llama.cpp

llama.cpp

11 videos
An open source local LLM engine for GGUF models, with CPU and GPU support, a built-in web UI, and an OpenAI-compatible server.

130KUpdated 39 minutes agoMIT

Web#Code execution#GGUF#Hugging Face integration

Favicon of MLX

MLX

2 videos
An open-source machine learning array framework with NumPy-style APIs, shared CPU and GPU memory on Apple silicon, and Linux CPU and CUDA backends.

28.6KUpdated 24 hours agoMIT

macOS · Linux#Distributed execution#LoRA

An open-source speech-to-text engine that runs Whisper models locally on desktop and mobile, with CPU-only inference and GPU acceleration. MIT licensed.

54KUpdated 2 days agoMIT

macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference

An open-source Python library for running and training text, vision, audio and multimodal models locally, with Apache 2.0 licensing and PyTorch support.

166.8KUpdated 1 day agoApache-2.0

#Hugging Face integration#Multimodal input

A local LLM inference SDK for desktop, mobile and browsers. Runs on CPU or GPU, with offline inference and online Picovoice account validation.

318Updated 3 weeks agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Quantization

Python toolkit for semantic search and RAG with BGE embedding models, multilingual rerankers, fine-tuning and evaluation. MIT licensed.

12.2KUpdated 1 month agoMIT

#Multilingual#Multimodal input#Semantic search

Local AI music generation library built on Stable Diffusion. Run it on your own hardware with CUDA, Apple Silicon or CPU. MIT licensed and no longer maintained.

3.9KUpdated 2 years agoMIT

macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#Multimodal input

A C++ neural translation library based on Marian NMT, with native and WebAssembly builds for local processing. Open source under MPL 2.0.

550Updated 2 years agoMPL-2.0

Web#Multilingual

A local document retrieval library that matches text queries to page images without OCR. MIT-licensed Python code supports NVIDIA and Apple Silicon GPUs.

2.8KUpdated 1 month agoMIT

macOS#Batch processing#Hugging Face integration#LoRA

Local LLM quantization library for smaller model weights and inference on your hardware. MIT licensed, with CPU and GPU support; archived and unmaintained.

2.3KUpdated 1 year agoMIT

Linux#Batch processing#GGUF#Hugging Face integration

Self-hosted LLM inference engine for Hugging Face models, with OpenAI-compatible APIs, multimodal support, and CPU or GPU execution under AGPL-3.0.

1.9KUpdated 3 weeks agoAGPL-3.0

macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

An on-device AI SDK that runs text, image and audio models on macOS, Windows and Linux, with GGUF, MLX and an OpenAI-compatible API.

qualcomm/GenieXInference Libraries and Bindings

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

An open source on-device AI engine for local LLMs and image models, with iOS, Android, CPU and GPU support under Apache 2.0.

16.2KUpdated 1 day agoApache-2.0

Windows · iOS · Android#Image-to-image#Multimodal input#ONNX

An open-source Python text-to-speech library that runs locally on CPU or CUDA GPUs and controls voice style through text descriptions.

5.6KUpdated 2 years agoApache-2.0

macOS#Hugging Face integration

A Python library for running and training CLIP image-text models on your own hardware, with local checkpoints and Hugging Face model support.

14.2KUpdated 5 days ago

#Hugging Face integration#Multimodal input

Open-source machine learning library in C/C++ with CPU, GPU, NPU and browser backends, quantization support, and an MIT license.

15.4KUpdated 6 days agoMIT

Web#Quantization

Local LLM acceleration library for Intel CPUs, GPUs and NPUs. Runs on Windows and Linux, integrates with Ollama and llama.cpp, and is archived.

8.9KUpdated 8 months agoApache-2.0

Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration

A local LLM inference library for Windows and Linux with NVIDIA GPUs, MIT licensing, GPTQ and EXL2 support, and an OpenAI-compatible API through TabbyAPI.

4.6KUpdated 7 months agoMIT

Windows · Linux#Batch processing#Quantization#Speculative decoding

An open-source on-device AI framework for Android, iOS, desktop and web, with model conversion from PyTorch, TensorFlow and JAX.

3.5KUpdated 19 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input

An AI inference framework for Android, iOS, desktop and browsers, with CPU and Vulkan GPU support and PyTorch and ONNX model conversion.

23.9KUpdated 6 days ago

macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization

An open-source model quantization library for local LLMs and vision models, with Apache 2.0 licensing and Hugging Face Transformers integration.

960Updated 7 months agoApache-2.0

#Hugging Face integration#LoRA#Quantization

An open-source wake word library for local voice apps, with English models, custom phrase training, and ONNX support on Linux and Windows.

2.8KUpdated 9 months agoApache-2.0

Windows · Linux#Batch processing#ONNX#Voice activity detection

An open-source Python OCR library that reads PDFs and images locally with PyTorch, runs on CPU or GPU, and supports Docker deployment.

6.4KUpdated 1 day agoApache-2.0

Docker · Web

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

5.1KUpdated 20 hours ago

macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API

An open-source Python package for local text embeddings on CPU, with MIT licensing and integrations for Sentence Transformers and LangChain.

2.2KUpdated 1 day agoMIT

#Hugging Face integration#Multilingual

Open-source AI training and inference framework for your own GPU hardware, with distributed parallelism and Apache 2.0 licensing.

41.4KUpdated 3 days agoApache-2.0

#Distributed execution

A Python toolkit for local LLM compression and inference on Linux, macOS and Windows, with GPTQ, AWQ, GGUF and integrations for vLLM and SGLang.

1.3KUpdated 1 day ago

macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA

A text-to-music model with melody conditioning and local GPU inference. AudioCraft code is MIT licensed; pretrained weights have a noncommercial license.

23.7KUpdated 2 years agoMIT

#Multimodal input

Favicon of MeloTTS

MeloTTS

1 video
A local text-to-speech library with real-time CPU inference, multilingual voices and custom dataset training. Free under the MIT license.

7.7KUpdated 2 years agoMIT

Docker · Web#Multilingual

An open-source toolkit for on-device AI across mobile, web and desktop. Apache 2.0 licensed; input stays local, while API metrics go to Google.

37.1KUpdated 19 hours agoApache-2.0

iOS · Android · Web

Favicon of EasyOCR

EasyOCR

2 videos
An open-source Python OCR library that extracts text from images locally, supports multilingual reading, and uses the Apache 2.0 license.

30KUpdated 10 months agoApache-2.0

Windows#Multilingual

Open-source Python tools optimize Hugging Face models for local, edge and cloud hardware, with ONNX Runtime, OpenVINO and TensorRT-LLM integrations.

3.5KUpdated 6 days agoApache-2.0

#Hugging Face integration#ONNX#Quantization

Open-source Python library for object detection and segmentation, with pretrained models and deployment exports. Licensed under Apache 2.0.

34.7KUpdated 2 days agoApache-2.0

More in Run Models Locally