LLM Inference Engines and Libraries

The engines and bindings that run models inside your own code: llama.cpp, MLX, ONNX Runtime and Hugging Face Transformers.

77 tools
Open-source computer vision library for local detection, segmentation and tracking, with AGPL-3.0 licensing and exports to ONNX, TensorRT and CoreML.

62.1KUpdated 24 hours agoAGPL-3.0

#ONNX

Favicon of WebLLM

WebLLM

2 videos
A local LLM engine that runs models in the browser with WebGPU acceleration, OpenAI API compatibility, and an Apache 2.0 license.

19.2KUpdated 2 weeks agoApache-2.0

Web · Browser Extension#OpenAI-compatible API#Streaming inference#Structured output

Python library for running and training embedding and reranker models locally, with Apache 2.0 licensing and pretrained models on Hugging Face.

19.1KUpdated 1 week agoApache-2.0

#Hugging Face integration#Multilingual#Multimodal input

Open-source Rust ML framework for running models locally on CPUs, NVIDIA GPUs or in browsers, with Apache 2.0 licensing and quantized LLM support.

21.1KUpdated 2 days agoApache-2.0

macOS · Web#GGUF#Hugging Face integration#Multilingual

An open-source Python library for local speech transcription with Whisper models. It runs on CPUs or NVIDIA GPUs and uses CTranslate2.

25.6KUpdated 3 hours agoMIT

#Batch processing#Hugging Face integration#Quantization

More in Run Models Locally