ebook2audiobook turns non-DRM ebooks into narrated audio with chapters and metadata, for readers who want audio editions of their own books. It runs locally on Windows, macOS and Linux, with Docker support and a browser interface built with Gradio. It's open source under Apache 2.0.
Voice cloning lets you use an audio recording as the narrator, and the tool can clean background noise or music from the reference recording. Supported speech engines include XTTSv2, Bark, Fairseq, VITS, Tacotron2, Tortoise, GlowTTS and YourTTS. You can also use custom trained models or supplied fine-tuned presets. Language support depends on the engine, with options including English, Chinese, Arabic, Hindi and Swahili.
Input support covers EPUB, MOBI, PDF, Word documents and plain text, while OCR extracts text from scanned pages and images. EPUB and MOBI give the best results for automatic chapter detection. It also reads short text. Output formats include M4B, MP3, FLAC and WAV, in mono or stereo. For narration that needs more control, SML tags specify pauses and voice changes within the text.
Batch conversion and a command-line mode suit larger book collections, and interrupted conversions can resume. The tool can translate an ebook with Argos Translate before generating speech and connect to an Audiobookshelf server. It supports CPU processing and acceleration through CUDA, ROCm, Intel XPU, Jetson and Apple Silicon MPS. Minimum requirements are 2 GB RAM and, for GPU use, 1 GB VRAM; modern speech engines can be very slow on CPU.
Claim this page and we'll verify you by hand. ebook2audiobook gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find ebook2audiobook?Promote it
Something wrong or outdated on this page?
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
14.4KUpdated 24 hours agoMIT
macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual
1.5KUpdated 4 months agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
Subtitle Edit is an MIT-licensed subtitle editor for Windows, macOS and Linux. It's for people creating captions, translating dialogue or fixing subtitles that don't match the video. Its core editing, conversion and video playback work offline on your device, with optional AI tools for transcription and translation.
Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.