Favicon of node-llama-cpp

node-llama-cpp

Local LLM library runs GGUF models through llama.cpp in Node.js, Bun and Electron. MIT licensed, with GPU support and JSON schema enforcement.

Screenshot of node-llama-cpp website

node-llama-cpp is an open source library for developers adding local LLM inference to JavaScript and TypeScript applications. It connects Node.js, Bun and Electron to llama.cpp, running GGUF models on your own machine. Its MIT license allows use in commercial projects.

A notable strength is control over model output. The library can enforce a JSON schema during generation, so applications can request data in a defined format. Function calling lets models request information or trigger actions through functions supplied by the application. It also supports structured decisions with any model.

Beyond chat, it handles text completion and fill-in-the-middle generation, along with embeddings and reranking for applications that search or compare documents. LoRA support lets applications use model adapters. TypeScript types help developers work with these capabilities in an existing codebase.

It runs on macOS, Linux and Windows, including Apple Silicon and Windows on Arm. Metal, CUDA and Vulkan provide GPU acceleration, and the library adapts automatically to available hardware. Prebuilt native binaries reduce the need to compile the underlying runtime.

For evaluating models before building an application, it includes a terminal chat interface and an example Electron app. The library also handles chat templates, supports Jinja templates, and manages context as conversations grow. Prompt preloading and automatic batching support repeated requests, while input protections guard against special token injection.

Similar to node-llama-cpp