llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
The library supports chat and text generation, including streamed responses. Developers can also generate embeddings and rerank documents, with support for models such as jina-reranker-v2-base-multilingual-GGUF and bge-reranker-v2-m3-GGUF. Those capabilities let an app use local models for search as well as conversation.
For apps that need predictable output, llama.rn supports JSON schema and GBNF constraints. Tool calling uses Jinja templates to connect model responses with app functions. It can process concurrent requests and manage their queue, rather than limiting an app to one request at a time.
Compatible multimodal models can interpret images and audio through matching mmproj projector files. Media inputs can come from local files or base64 data; audio support includes WAV and MP3. Multimodal processing needs more memory than text alone.
GPU acceleration uses Metal on iOS and OpenCL on supported Android hardware, including tested Qualcomm Adreno 700-series and newer devices. On Apple Silicon in the iOS runtime, Metal is the supported inference path because the CPU-only path has a known crash.
Experimental text-to-speech uses codec.cpp and includes reference-audio voice cloning. Model families and hardware backends vary in reliability; use the tested-model table when selecting a speech model.
Claim this page and we'll verify you by hand. llama.rn gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find llama.rn?Promote it
Something wrong or outdated on this page?
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
318Updated 3 weeks agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Quantization
picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
16.3KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Image-to-image#Multilingual
26.6KUpdated 2 months agoMIT
Linux#Multilingual#Voice cloning#Voice conversion
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).