
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
You can use it as an API server for AI applications or through a Python API for local and batched inference. Its OpenAI-compatible endpoints support streaming, embeddings, tool calls, and reasoning output. It also provides Anthropic and Kobold APIs, along with scoring, reranking, and transcription interfaces.
For busy servers, continuous batching lets Sonar process requests together, while prefix caching reuses work on shared prompt text. Quantized weights and cache management help control memory use; speculative decoding can speed up generation. Deployments can span a single GPU, multiple GPUs, or multiple nodes.
Sonar supports image, audio, and video models, structured output, and LoRA adapter serving. Compatibility depends on the model, device, and quantization method, so these capabilities don't apply to every deployment.
Hardware support includes NVIDIA CUDA, AMD ROCm, Intel XPU, Google TPU, and CPUs. It runs on Linux and Apple silicon macOS with Metal, with Docker and WSL 2 deployment options. Production servers can expose Prometheus metrics and health endpoints.
Claim this page with an email at sonar.dphn.ai. Sonar (formerly Aphrodite Engine) gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Sonar (formerly Aphrodite Engine)?Promote it
Something wrong or outdated on this page?
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
9.8KUpdated 5 months agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
PowerInfer is a local LLM inference engine for developers and researchers who want to run large models on a PC with a consumer GPU. It splits work between the CPU and GPU to reduce GPU memory demands and data transfers. The code is open source under the MIT license.
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.