Favicon of CTranslate2

CTranslate2

A local AI inference library for C++ and Python that runs Transformer models on CPUs and GPUs, with quantization to reduce memory use. MIT licensed.

Screenshot of CTranslate2 website

CTranslate2 is an open-source C++ and Python library for developers running Transformer models on their own hardware or servers. It handles translation, text generation, text encoding and speech recognition. Its custom runtime focuses on reducing inference time and memory use compared with general-purpose deep learning frameworks.

Model support spans different tasks. It runs translation models such as NLLB and M2M-100, Whisper for speech recognition, and text generation models including Llama, Mistral, Gemma and Qwen2. BERT, DistilBERT and XLM-RoBERTa cover text encoding. Compatible models use CTranslate2's optimized format, with converters for Transformers, OpenNMT, Fairseq, Marian and OPUS-MT.

CPU-only inference is supported on x86-64 and ARM64 processors. GPU execution supports NVIDIA CUDA, and the project provides Python packages for AMD ROCm GPUs. Quantization and reduced precision lower model storage and memory requirements, including INT8 and AWQ INT4 support. The runtime automatically selects compatible CPU optimizations and can process batches asynchronously across CPU cores or multiple GPUs. Large models can also split across GPUs through tensor parallelism.

It's MIT licensed. The Python and C++ APIs suit applications that need an inference engine, while interactive decoding can complete partial text and return alternatives at a chosen position. A separate project, ctranslate2-web-server, exposes an OpenAI-compatible REST API for applications that use that interface.

Similar to CTranslate2