Favicon of Candle

Candle

Open-source Rust ML framework for running models locally on CPUs, NVIDIA GPUs or in browsers, with Apache 2.0 licensing and quantized LLM support.

Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.

It supports model training as well as inference. Developers can extend it with custom operations and GPU kernels, including FlashAttention. The framework is open source under the Apache 2.0 license.

Hardware and deployment choices include CPU execution, NVIDIA GPUs through CUDA, and browser execution through WebAssembly. The CPU backend can use MKL on x86 or Apple's Accelerate on Macs. NCCL supports distributing work across multiple GPUs, while browser demos perform inference within the browser.

Its model implementations cover more than text generation:

  • LLaMA, Gemma, Mistral, Mixtral and Qwen3 MoE for language tasks, plus StarCoder and StarCoder2 for code generation.
  • Whisper for speech recognition and MetaVoice or Parler-TTS for speech synthesis.
  • Stable Diffusion for image generation, YOLO for object detection, and Segment Anything for image segmentation.
  • BLIP for image captions and TrOCR for printed or handwritten text recognition.

Candle loads weights from safetensors, npz, ggml and PyTorch files. It supports llama.cpp quantization types and GGUF quantized Qwen3 MoE models. Access to gated LLaMA 2 weights requires a Hugging Face account and acceptance of the model's terms.

Similar to Candle