Favicon of llama.rn

llama.rn

An open-source React Native library that runs GGUF models on iOS and Android through llama.cpp, with GPU acceleration and image and audio understanding.

llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.

The library supports chat and text generation, including streamed responses. Developers can also generate embeddings and rerank documents, with support for models such as jina-reranker-v2-base-multilingual-GGUF and bge-reranker-v2-m3-GGUF. Those capabilities let an app use local models for search as well as conversation.

For apps that need predictable output, llama.rn supports JSON schema and GBNF constraints. Tool calling uses Jinja templates to connect model responses with app functions. It can process concurrent requests and manage their queue, rather than limiting an app to one request at a time.

Compatible multimodal models can interpret images and audio through matching mmproj projector files. Media inputs can come from local files or base64 data; audio support includes WAV and MP3. Multimodal processing needs more memory than text alone.

GPU acceleration uses Metal on iOS and OpenCL on supported Android hardware, including tested Qualcomm Adreno 700-series and newer devices. On Apple Silicon in the iOS runtime, Metal is the supported inference path because the CPU-only path has a known crash.

Experimental text-to-speech uses codec.cpp and includes reference-audio voice cloning. Model families and hardware backends vary in reliability; use the tested-model table when selecting a speech model.

Similar to llama.rn