
node-llama-cpp is an open source library for developers adding local LLM inference to JavaScript and TypeScript applications. It connects Node.js, Bun and Electron to llama.cpp, running GGUF models on your own machine. Its MIT license allows use in commercial projects.
A notable strength is control over model output. The library can enforce a JSON schema during generation, so applications can request data in a defined format. Function calling lets models request information or trigger actions through functions supplied by the application. It also supports structured decisions with any model.
Beyond chat, it handles text completion and fill-in-the-middle generation, along with embeddings and reranking for applications that search or compare documents. LoRA support lets applications use model adapters. TypeScript types help developers work with these capabilities in an existing codebase.
It runs on macOS, Linux and Windows, including Apple Silicon and Windows on Arm. Metal, CUDA and Vulkan provide GPU acceleration, and the library adapts automatically to available hardware. Prebuilt native binaries reduce the need to compile the underlying runtime.
For evaluating models before building an application, it includes a terminal chat interface and an example Electron app. The library also handles chat templates, supports Jinja templates, and manages context as conversations grow. Prompt preloading and automatic batching support repeated requests, while input protections guard against special token injection.
Claim this page with an email at node-llama-cpp.withcat.ai. node-llama-cpp gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find node-llama-cpp?Promote it
Something wrong or outdated on this page?
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
318Updated 3 weeks agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Quantization
picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.
7.2KUpdated 1 day agoMIT
macOS#Batch processing#Distributed execution#Hugging Face integration
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.