LLM Inference Engines and Libraries

The engines and bindings that run models inside your own code: llama.cpp, MLX, ONNX Runtime and Hugging Face Transformers.

77 tools
AI model hub with an Apache 2.0 Python library for local inference, training and evaluation, plus hosted demos and cloud notebooks.

9.2KUpdated 7 days agoApache-2.0

Docker#Image-to-image#Inpainting#Multimodal input

An open-source React Native library that runs GGUF models on iOS and Android through llama.cpp, with GPU acceleration and image and audio understanding.

1KUpdated 3 days agoMIT

iOS · Android#GGUF#llama.cpp backend#Multilingual

An open-source AI music model and synthesis engine that runs locally on Apple Silicon, with a macOS app and AUv3 plugin for DAWs. Apache 2.0 licensed.

1.8KUpdated 2 months agoApache-2.0

macOS#MLX#Streaming inference

Face analysis toolkit for self-hosted recognition on CPU or NVIDIA GPU, local video face redaction, and commercially licensed models.

29.9KUpdated 3 weeks ago

macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization

A local LLM inference engine for sparse models, with CPU and GPU support on Linux and Windows. Open source under MIT, with CPU-only support on Apple Silicon.

9.8KUpdated 5 months agoMIT

macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration

Open-source distributed LLM software runs inference and fine-tuning across shared GPUs, with public or private networks and support for Llama 3.1.

10.6KUpdated 2 years agoMIT

macOS · Windows · Linux · Docker#Hugging Face integration

Open-source Python face recognition library for local or self-hosted use, with TensorFlow and PyTorch support, facial analysis, and an MIT license.

23.5KUpdated 23 hours agoMIT

Docker

A Python reranking library with a shared interface for local models and cloud APIs, CPU inference through FlashRank, and an Apache 2.0 license.

1.6KUpdated 9 months agoApache-2.0

#Hugging Face integration#Multilingual#Multimodal input

An on-device AI engine runs automation models locally on phones and tiny devices, with speech, vision and optional cloud routing.

6.1KUpdated 5 days ago

macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization

An open-source PyTorch toolkit for training and testing object detection and segmentation models, with GPU operations and an Apache 2.0 license.

33KUpdated 3 years agoApache-2.0

Local LLM toolkit that runs language and multimodal models on Rockchip NPUs, with model conversion, quantization and C/C++ interfaces.

1.7KUpdated 2 days ago

Linux#Multimodal input#Quantization

PyTorch quantization library that reduces model memory use and speeds training and inference on your hardware, with CPU, GPU and mobile deployment support.

3KUpdated 5 days ago

Linux · iOS#Hugging Face integration#LoRA#Quantization

Favicon of DeepSpeed

DeepSpeed

1 video
Open-source PyTorch optimization library for distributed training and inference, with Apache 2.0 licensing and support for NVIDIA and AMD GPUs.

43.2KUpdated 1 day agoApache-2.0

Windows#Distributed execution

An open source engine for running ONNX models on Windows, macOS, Linux, mobile devices and the web, with CPU, GPU and NPU support.

22KUpdated 27 minutes agoMIT

macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX

Favicon of Kokoro

Kokoro

6 videos
Open-source text-to-speech model and library for local speech generation, with multilingual voices, Apache 2.0 licensing and Apple Silicon GPU support.

9.1KUpdated 1 year agoApache-2.0

macOS · Windows#Batch processing#Multilingual#ONNX

A local AI inference library for C++ and Python that runs Transformer models on CPUs and GPUs, with quantization to reduce memory use. MIT licensed.

4.7KUpdated 5 days agoMIT

#Batch processing#Multilingual#Quantization

Python library for running GGUF models locally through llama.cpp, with a self-hosted OpenAI-compatible server and CPU or GPU support.

10.6KUpdated 1 week agoMIT

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

An open-source JavaScript library that runs AI models in your browser using ONNX Runtime, with CPU execution and optional WebGPU acceleration.

16.3KUpdated 1 week agoApache-2.0

Web#Hugging Face integration#Image-to-image#Multilingual

Open-source Python toolkit for training and serving LLMs on your own hardware, with Apache 2.0 licensing, quantization and multi-GPU support.

13.7KUpdated 3 weeks agoApache-2.0

#Hugging Face integration#LoRA#Quantization

An open-source local LLM compiler and deployment engine with GPU support across desktop, browser and mobile platforms, plus an OpenAI-compatible API.

23.2KUpdated 24 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API

A local LLM inference engine with OpenAI and Anthropic-compatible APIs. Runs on macOS, Linux and Windows with CPU, CUDA or Apple Silicon support.

7.7KUpdated 5 days agoMIT

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration

An open-weight vision model for image questions, captions and object detection. Run it locally or use hosted inference and fine-tuning.

10.1KUpdated 5 months agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline

Open-source Python toolkit for local AI audio generation, fine-tuning and training, with a Gradio interface and support for Stable Audio Open.

3.9KUpdated 4 months agoMIT

Web#Hugging Face integration

Open-source LLM serving toolkit for your own GPU servers, with quantization, text and vision models, and OpenAI-compatible APIs. Apache 2.0 licensed.

8.1KUpdated 3 days agoApache-2.0

#Batch processing#Distributed execution#Hugging Face integration

A self-hosted LLM inference library built on PyTorch for NVIDIA GPUs, with a Python API, OpenAI-compatible serving, and multi-node support.

14.7KUpdated 21 hours ago

Docker#Batch processing#Distributed execution#LoRA

A PyTorch quantization library that reduces LLM memory use for inference and fine-tuning with 8-bit optimizers, LLM.int8() and QLoRA. MIT licensed.

8.5KUpdated 4 weeks agoMIT

macOS · Windows · Linux#LoRA#Quantization

An open-source local LLM framework that splits work across CPUs and GPUs, with SGLang serving and LlamaFactory fine-tuning under Apache 2.0.

19.5KUpdated 1 week agoApache-2.0

Docker#LoRA#Multimodal input#Prompt caching

Favicon of GPT4All

GPT4All

1 video
An open-source local AI chatbot for Windows, macOS and Linux. Run models without a GPU or cloud API, and chat privately with your documents.

77.4KUpdated 1 year agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API

Favicon of MLX LM

MLX LM

2 videos
A Python package for local LLM inference and fine-tuning on Apple Silicon, built on MLX with Hugging Face model support and an MIT license.

7.2KUpdated 1 day agoMIT

macOS#Batch processing#Distributed execution#Hugging Face integration

Local LLM library runs GGUF models through llama.cpp in Node.js, Bun and Electron. MIT licensed, with GPU support and JSON schema enforcement.

2.2KUpdated 3 days agoMIT

macOS · Windows · Linux#Batch processing#GGUF#Guardrails

Open-source speaker diarization toolkit with local PyTorch models, CUDA GPU support, and an optional hosted service that processes audio on pyannoteAI servers.

10.6KUpdated 3 months agoMIT

#Hugging Face integration#Speaker diarization#Voice activity detection

Open-source text-to-speech toolkit for local speech generation, voice cloning and model training on Linux, macOS and Windows, licensed under MPL-2.0.

2.3KUpdated 4 months agoMPL-2.0

macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning

Open-source Python toolkit for semantic search and RAG, with BGE embedding models, multilingual rerankers, evaluation and fine-tuning under MIT.

12.2KUpdated 1 month agoMIT

#Multilingual#Semantic search

Favicon of OpenVINO

OpenVINO

2 videos
Open-source AI inference toolkit for local or self-hosted deployment on Linux, Windows and macOS, with CPU, Intel GPU and NPU support.

10.9KUpdated 24 hours agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX

A Python library for local embeddings and reranking, using ONNX Runtime with CPU or GPU support. Open source under Apache 2.0.

3.2KUpdated 24 hours agoApache-2.0

#Batch processing#Multilingual#ONNX

A PyTorch library for training and inference on local machines or clusters, with CPU, GPU and TPU support. Open source under Apache 2.0.

9.9KUpdated 1 day agoApache-2.0

#Distributed execution

More in Run Models Locally