
PaddleSpeech is a Python toolkit built on PaddlePaddle for developers and researchers building speech applications on their own machines or servers. It covers speech recognition and synthesis, with streaming systems for both. The project uses the Apache 2.0 license and supports Linux, Windows and macOS, with Linux recommended. It supports CPU execution.
Its Chinese text processing is a distinct reason to consider it. The speech synthesis frontend normalizes text and converts written characters into pronunciations, with rules for characters that have multiple readings and for tone changes in context. The toolkit also restores punctuation in recognized text and supports English-to-Chinese speech translation.
The model selection includes Whisper large v3 and turbo, Conformer and DeepSpeech2 for recognition, plus WavLM, HuBERT and Wav2vec2 models. For speech synthesis, it includes Tacotron2, FastSpeech2 and ERNIE-SAT, with vocoders such as HiFiGAN and Parallel WaveGAN. Pretrained models sit alongside modules for training, inference and testing, so researchers can work on models within the same toolkit they use for deployment.
PaddleSpeech also handles speaker verification, keyword spotting and audio classification. Server and streaming server interfaces support applications that need ongoing audio processing. It includes subtitle generation examples that produce SRT files, code-switching recognition, and singing voice synthesis examples using DiffSinger.
Claim this page and we'll verify you by hand. PaddleSpeech gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PaddleSpeech?Promote it
Something wrong or outdated on this page?
1.4KUpdated 13 hours agoAGPL-3.0
macOS · Windows · Linux · Browser Extension#Multilingual#ONNX#OpenAI-compatible API
Sokuji is a free, open source speech translator for people joining meetings across languages. It sends your translated speech through a virtual microphone, so other participants hear ordinary call audio and don't need to install anything. Their replies appear as translated subtitles on your screen.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
5.4KUpdated 2 days agoMIT
macOS · Windows · Linux#Batch processing#MCP#Multilingual
14.4KUpdated 10 hours agoMIT
macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual
18.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
1.4KUpdated 1 month agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
SmartSub is a free, open-source desktop app for people who subtitle recordings or adapt videos into other languages. It combines local transcription, translation, subtitle editing and AI dubbing on Windows, macOS and Linux. Each stage also works independently.
Subtitle Edit is an MIT-licensed subtitle editor for Windows, macOS and Linux. It's for people creating captions, translating dialogue or fixing subtitles that don't match the video. Its core editing, conversion and video playback work offline on your device, with optional AI tools for transcription and translation.
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
SoniTranslate is a local AI video dubbing app for creators and translators who need speech in another language to follow the timing of the original video. Its Gradio browser interface brings transcription, translation and speech generation together, with speaker detection for recordings that contain multiple voices. Local installation is tested on Linux, and it can use an NVIDIA GPU or run in CPU mode.