Favicon of picoLLM

picoLLM

A local LLM inference SDK for desktop, mobile and browsers. Runs on CPU or GPU, with offline inference and online Picovoice account validation.

Screenshot of picoLLM website

picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.

After setup and license validation, model inference runs locally. The SDK still requires a Picovoice account and an AccessKey, and it needs an internet connection to validate that key with Picovoice's license servers. This matters for apps intended for disconnected environments: local processing doesn't remove the online licensing requirement. Model downloads also come through Picovoice Console.

The engine runs on CPU and GPU across Linux, macOS and Windows, including Apple Silicon Macs and Windows on ARM. It also supports Android, iOS, Raspberry Pi 4 and 5, and local execution in Chrome, Safari, Edge and Firefox. Developers can integrate it through SDKs for Python, C, .NET and Node.js, as well as mobile and web apps.

Supported open-weight model families include Gemma, Llama 2, Llama 3 and Llama 3.2, alongside Mistral, Mixtral and Phi. Both base models and instruction or chat variants are available. The repository carries the Apache 2.0 license and includes SDKs and prebuilt engine libraries; it does not supply the complete engine source. Picovoice AccessKey validation and the model licenses still apply.

Similar to picoLLM