Favicon of Sonar (formerly Aphrodite Engine)

Sonar (formerly Aphrodite Engine)

Self-hosted LLM inference engine for Hugging Face models, with OpenAI-compatible APIs, multimodal support, and CPU or GPU execution under AGPL-3.0.

Screenshot of Sonar (formerly Aphrodite Engine) website

Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.

You can use it as an API server for AI applications or through a Python API for local and batched inference. Its OpenAI-compatible endpoints support streaming, embeddings, tool calls, and reasoning output. It also provides Anthropic and Kobold APIs, along with scoring, reranking, and transcription interfaces.

For busy servers, continuous batching lets Sonar process requests together, while prefix caching reuses work on shared prompt text. Quantized weights and cache management help control memory use; speculative decoding can speed up generation. Deployments can span a single GPU, multiple GPUs, or multiple nodes.

Sonar supports image, audio, and video models, structured output, and LoRA adapter serving. Compatibility depends on the model, device, and quantization method, so these capabilities don't apply to every deployment.

Hardware support includes NVIDIA CUDA, AMD ROCm, Intel XPU, Google TPU, and CPUs. It runs on Linux and Apple silicon macOS with Metal, with Docker and WSL 2 deployment options. Production servers can expose Prometheus metrics and health endpoints.

Similar to Sonar (formerly Aphrodite Engine)