Alexandria Audiobook is a local AI audiobook generator for people who want separate narrator and character voices, with control over individual lines. It accepts EPUB, text and Markdown books, uses an LLM to identify speakers and delivery directions, and generates speech with a built-in Qwen3-TTS engine. It's open source under the MIT license.
The browser editor lets you correct speaker assignments, change text, preview audio and regenerate selected passages. An optional LLM review checks annotation errors. Character voices can come from presets, short reference recordings or written descriptions; automatic persona generation can assign voices based on the script. LoRA training creates reusable voice identities that follow delivery instructions, while an included dataset editor helps prepare voice samples. Cloned voices don't follow those instructions.
Script annotation requires a separate LLM through an OpenAI-compatible API, including LM Studio, Ollama or OpenAI. With local backends, the book text and speech processing stay on your hardware; cloud annotation sends the book text to the chosen provider. Speech can also run on a remote Qwen3-TTS server. Model downloads need an internet connection.
The app runs on Windows, Linux and macOS, with NVIDIA GPU support on Windows and Linux and AMD GPU support on Linux. Macs, including Apple Silicon, use the slower CPU mode, as does AMD hardware on Windows. GPU use needs at least 8 GB VRAM; 16 GB RAM is recommended. Docker supports NVIDIA server deployments. Exports include a combined MP3, chaptered M4B for audiobook players, and separate speaker tracks and labels for Audacity.
Claim this page and we'll verify you by hand. Alexandria Audiobook gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Alexandria Audiobook?Promote it
Something wrong or outdated on this page?
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
674Updated 4 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#GGUF#llama.cpp backend
8.6KUpdated 4 months ago
Windows · Linux · Docker · Web#Hugging Face integration#Voice cloning
VoiceCraft can generate speech or edit an existing recording using just a few seconds of a speaker's voice as reference. It runs locally and suits speech researchers, developers adding voice generation to their projects, and people who want to work with audio on their own hardware.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
3.8KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Batch processing#Voice conversion
Voice-Clone-Studio brings several speech models into one local browser interface for people making podcasts, audiobooks or custom voice recordings. It combines voice cloning, voice design and audio preparation, so you can compare engines without managing a separate app for each. It's open source under Apache 2.0.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.
Applio is a local AI voice conversion suite for musicians making AI covers, streamers changing their voice live, and creators working with speech. It converts recordings or microphone input into another voice using community models or models you train yourself. Its software uses the MIT license.