Favicon of Nexa SDK

Nexa SDK

An on-device AI SDK that runs text, image and audio models on macOS, Windows and Linux, with GGUF, MLX and an OpenAI-compatible API.

Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.

The SDK runs on macOS, Windows and Linux. It supports GGUF, MLX and Nexa's .nexa model format, including compatible models from Hugging Face. GGUF works across all three desktop platforms; MLX requires Apple Silicon macOS. Hardware backends include CUDA, Metal and Vulkan, plus Qualcomm, Intel and AMD NPUs. Qualcomm NPU inference requires a Snapdragon X Elite laptop, while Apple Neural Engine support covers speech recognition with Parakeet.

Supported models include Qwen3-VL for image understanding, Gemma-3n for multimodal inference and IBM Granite for language tasks. Parakeet and Kokoro cover speech recognition and synthesis, while SDXL handles image generation. Multimodal interactions can include multiple images or audio clips in the same conversation.

An OpenAI-compatible API server lets applications use local inference through a familiar API. It supports streamed responses and function calling defined with JSON schemas. The SDK also includes a command-line interface for chatting with models and managing downloaded models. Its Nexa ML Turbo engine targets NPU performance, and plugin isolation separates the inference backends.

This listing covers the pinned v0.2.50 release, before the repository became Qualcomm's GenieX. Some models require a Nexa account and license token.

Similar to Nexa SDK