
picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.
After setup and license validation, model inference runs locally. The SDK still requires a Picovoice account and an AccessKey, and it needs an internet connection to validate that key with Picovoice's license servers. This matters for apps intended for disconnected environments: local processing doesn't remove the online licensing requirement. Model downloads also come through Picovoice Console.
The engine runs on CPU and GPU across Linux, macOS and Windows, including Apple Silicon Macs and Windows on ARM. It also supports Android, iOS, Raspberry Pi 4 and 5, and local execution in Chrome, Safari, Edge and Firefox. Developers can integrate it through SDKs for Python, C, .NET and Node.js, as well as mobile and web apps.
Supported open-weight model families include Gemma, Llama 2, Llama 3 and Llama 3.2, alongside Mistral, Mixtral and Phi. Both base models and instruction or chat variants are available. The repository carries the Apache 2.0 license and includes SDKs and prebuilt engine libraries; it does not supply the complete engine source. Picovoice AccessKey validation and the model licenses still apply.
Claim this page with an email at picovoice.ai. picoLLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find picoLLM?Promote it
Something wrong or outdated on this page?
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
16.3KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Image-to-image#Multilingual
19.2KUpdated 2 weeks agoApache-2.0
Web · Browser Extension#OpenAI-compatible API#Streaming inference#Structured output
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
WebLLM runs language models directly in a user's browser, using WebGPU for GPU acceleration. It's an open-source engine for developers building web-based AI assistants and Chrome extensions that process prompts on the user's device rather than an inference server. The project uses the Apache 2.0 license.