ggml is an open-source library for developers building software that runs machine learning models on their own hardware. It provides tensor computation in C/C++ without dependencies, with backends for CPUs, GPUs, NPUs and browsers. The project also documents the GGUF model file format.
Its main appeal is portability. The same library supports x86 and ARM alongside RISC-V, LoongArch, PowerPC and s390x, with WebAssembly support for browser applications. Developers can use it as a foundation for local AI software across different types of hardware.
Quantization is a central capability: ggml supports integer formats from 2 to 8 bits, plus MXFP4 and NVFP4 microscaling formats. These are relevant when choosing how to represent model weights within a device's memory constraints. Optimized computation kernels use SIMD instructions on x86, ARM and RISC-V, and the library makes no memory allocations during runtime.
This is a developer library, so its capabilities matter most to people building model runtimes or integrating machine learning into an application. The project includes a small matrix multiplication example and documentation on tensor computation and GGUF. The MIT license permits use and modification in both open-source and proprietary software.
Claim this page and we'll verify you by hand. GGUF and ggml gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find GGUF and ggml?Promote it
Something wrong or outdated on this page?
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
29.9KUpdated 3 weeks ago
macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization
InsightFace is a face analysis toolkit for developers and teams building identity verification, access control, or face editing software. The code uses the MIT license. Its Python tools and self-hosted recognition server run inference on your own hardware. It also offers commercial models and API access for face swapping and deepfake detection.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
130KUpdated 39 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.