Favicon of GGUF and ggml

GGUF and ggml

Open-source machine learning library in C/C++ with CPU, GPU, NPU and browser backends, quantization support, and an MIT license.

ggml is an open-source library for developers building software that runs machine learning models on their own hardware. It provides tensor computation in C/C++ without dependencies, with backends for CPUs, GPUs, NPUs and browsers. The project also documents the GGUF model file format.

Its main appeal is portability. The same library supports x86 and ARM alongside RISC-V, LoongArch, PowerPC and s390x, with WebAssembly support for browser applications. Developers can use it as a foundation for local AI software across different types of hardware.

Quantization is a central capability: ggml supports integer formats from 2 to 8 bits, plus MXFP4 and NVFP4 microscaling formats. These are relevant when choosing how to represent model weights within a device's memory constraints. Optimized computation kernels use SIMD instructions on x86, ARM and RISC-V, and the library makes no memory allocations during runtime.

This is a developer library, so its capabilities matter most to people building model runtimes or integrating machine learning into an application. The project includes a small matrix multiplication example and documentation on tensor computation and GGUF. The MIT license permits use and modification in both open-source and proprietary software.

Similar to GGUF and ggml