
CTranslate2 is an open-source C++ and Python library for developers running Transformer models on their own hardware or servers. It handles translation, text generation, text encoding and speech recognition. Its custom runtime focuses on reducing inference time and memory use compared with general-purpose deep learning frameworks.
Model support spans different tasks. It runs translation models such as NLLB and M2M-100, Whisper for speech recognition, and text generation models including Llama, Mistral, Gemma and Qwen2. BERT, DistilBERT and XLM-RoBERTa cover text encoding. Compatible models use CTranslate2's optimized format, with converters for Transformers, OpenNMT, Fairseq, Marian and OPUS-MT.
CPU-only inference is supported on x86-64 and ARM64 processors. GPU execution supports NVIDIA CUDA, and the project provides Python packages for AMD ROCm GPUs. Quantization and reduced precision lower model storage and memory requirements, including INT8 and AWQ INT4 support. The runtime automatically selects compatible CPU optimizations and can process batches asynchronously across CPU cores or multiple GPUs. Large models can also split across GPUs through tensor parallelism.
It's MIT licensed. The Python and C++ APIs suit applications that need an inference engine, while interactive decoding can complete partial text and return alternatives at a chosen position. A separate project, ctranslate2-web-server, exposes an OpenAI-compatible REST API for applications that use that interface.
Claim this page with an email at opennmt.net. CTranslate2 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find CTranslate2?Promote it
Something wrong or outdated on this page?
1.5KUpdated 2 months agoMIT
Docker#Batch processing#Multilingual#OpenAI-compatible API
subgen generates subtitles on your own hardware for personal media libraries, including films and shows that don't have usable subtitles available. It's an open source, MIT-licensed Python service that runs in Docker or as a standalone application. Speech recognition runs locally using Whisper models through faster-whisper and stable-ts, with support for CPU processing and NVIDIA GPUs through CUDA.
sindresorhus.comOn-Device and In-Browser AI
macOS · iOS#Batch processing#Multilingual
2.3KUpdated 1 year agoMIT
Linux#Batch processing#GGUF#Hugging Face integration
4.6KUpdated 7 months agoMIT
Windows · Linux#Batch processing#Quantization#Speculative decoding
25.6KUpdated 2 hours agoMIT
#Batch processing#Hugging Face integration#Quantization
13.1KUpdated 2 years agoApache-2.0
macOS · Windows#Batch processing#Hugging Face integration#Multilingual
Aiko is a paid, native transcription app for macOS, iOS and visionOS that processes speech on your device with OpenAI's Whisper model. It's for people turning meetings, lectures or other recordings into text while keeping the audio local, including sensitive recordings.
AutoAWQ is a Python library for developers who want to compress and run LLMs on their own hardware using 4-bit Activation-aware Weight Quantization (AWQ). The project is archived and no longer maintained. It's open source under the MIT license. It installs as a Python package, with optional kernel or Intel CPU dependencies.
ExLlamaV2 is a local LLM inference library for developers and people hosting models on their own consumer GPUs. ExLlamaV2 is archived and no longer maintained; development continues in ExLlamaV3. The V2 library is free and open source under the MIT license, runs on Windows and Linux, and uses NVIDIA GPUs through CUDA. It supports multiple GPUs.
faster-whisper is a Python library for people building local speech transcription into their own software. It runs OpenAI's Whisper models through CTranslate2, with faster processing and lower memory use than the original Whisper implementation in the project's comparisons. It runs on a CPU. NVIDIA GPUs are supported too, and the code is open source under the MIT license.
insanely-fast-whisper is a command-line tool for people who want to transcribe audio on their own hardware, with a focus on processing long recordings quickly. It runs OpenAI's Whisper locally on NVIDIA GPUs or Apple Silicon Macs, including support for Windows with CUDA. The project is open source under the Apache 2.0 license.