
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
Its scope extends beyond text generation. It supports summarization, translation, sentiment analysis and questions answered from supplied text, along with embeddings for comparing content. Image tasks include object detection, background removal, segmentation and depth estimation. Audio support covers speech transcription, audio classification and text-to-speech. It can also extract text from images and answer questions about document images.
The API follows Hugging Face's Python Transformers library closely, which makes it a practical choice for developers bringing existing model-based work to JavaScript. Its pipeline interface combines a pretrained model with input preparation and output processing. Compatible architectures include BERT, DistilBERT, BART, CLIP and CodeLlama, with compatible models available through the Hugging Face Hub.
ONNX Runtime handles model execution. The browser uses the CPU through WebAssembly by default, with WebGPU available for GPU acceleration. Quantized models reduce download size and resource demands in browser environments. Developers can also bring their own pretrained PyTorch, TensorFlow or JAX models by converting them to ONNX with Hugging Face Optimum.
Claim this page and we'll verify you by hand. Transformers.js gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Transformers.js?Promote it
Something wrong or outdated on this page?
19.2KUpdated 2 weeks agoApache-2.0
Web · Browser Extension#OpenAI-compatible API#Streaming inference#Structured output
WebLLM runs language models directly in a user's browser, using WebGPU for GPU acceleration. It's an open-source engine for developers building web-based AI assistants and Chrome extensions that process prompts on the user's device rather than an inference server. The project uses the Apache 2.0 license.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
318Updated 3 weeks agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Quantization
picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.