PaddleSpeech

Open-source speech AI toolkit for local or self-hosted use on Linux, Windows and macOS, with streaming recognition, speech synthesis and Chinese text processing

Screenshot of PaddleSpeech website

PaddleSpeech is a Python toolkit built on PaddlePaddle for developers and researchers building speech applications on their own machines or servers. It covers speech recognition and synthesis, with streaming systems for both. The project uses the Apache 2.0 license and supports Linux, Windows and macOS, with Linux recommended. It supports CPU execution.

Its Chinese text processing is a distinct reason to consider it. The speech synthesis frontend normalizes text and converts written characters into pronunciations, with rules for characters that have multiple readings and for tone changes in context. The toolkit also restores punctuation in recognized text and supports English-to-Chinese speech translation.

The model selection includes Whisper large v3 and turbo, Conformer and DeepSpeech2 for recognition, plus WavLM, HuBERT and Wav2vec2 models. For speech synthesis, it includes Tacotron2, FastSpeech2 and ERNIE-SAT, with vocoders such as HiFiGAN and Parallel WaveGAN. Pretrained models sit alongside modules for training, inference and testing, so researchers can work on models within the same toolkit they use for deployment.

PaddleSpeech also handles speaker verification, keyword spotting and audio classification. Server and streaming server interfaces support applications that need ongoing audio processing. It includes subtitle generation examples that produce SRT files, code-switching recognition, and singing voice synthesis examples using DiffSinger.

Similar to PaddleSpeech