Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
The SDK runs on macOS, Windows and Linux. It supports GGUF, MLX and Nexa's .nexa model format, including compatible models from Hugging Face. GGUF works across all three desktop platforms; MLX requires Apple Silicon macOS. Hardware backends include CUDA, Metal and Vulkan, plus Qualcomm, Intel and AMD NPUs. Qualcomm NPU inference requires a Snapdragon X Elite laptop, while Apple Neural Engine support covers speech recognition with Parakeet.
Supported models include Qwen3-VL for image understanding, Gemma-3n for multimodal inference and IBM Granite for language tasks. Parakeet and Kokoro cover speech recognition and synthesis, while SDXL handles image generation. Multimodal interactions can include multiple images or audio clips in the same conversation.
An OpenAI-compatible API server lets applications use local inference through a familiar API. It supports streamed responses and function calling defined with JSON schemas. The SDK also includes a command-line interface for chatting with models and managing downloaded models. Its Nexa ML Turbo engine targets NPU performance, and plugin isolation separates the inference backends.
This listing covers the pinned v0.2.50 release, before the repository became Qualcomm's GenieX. Some models require a Nexa account and license token.
Claim this page and we'll verify you by hand. Nexa SDK gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Nexa SDK?Promote it
Something wrong or outdated on this page?
23.2KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
407Updated 2 days agoMIT
macOS#MLX#Multimodal input#OpenAI-compatible API
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
Slotstream runs Qwen3.8-Flash-Next on Apple Silicon Macs that don't have enough RAM to hold the whole model. It's aimed at people with 16 to 64 GB of memory who want local chat, image questions or a model backend for coding agents. Most model weights stay on the SSD, while frequently used expert networks stay in memory. The full model remains available.
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.