Podcastfy turns documents, websites and images into AI-generated audio conversations, with the option to write transcripts using a local LLM. It's an open-source Python alternative to NotebookLM's podcast feature for creators, educators and researchers who want control over the conversation format or need podcast generation inside their own software. It uses the Apache License 2.0.
Inputs can include PDFs, YouTube videos, supplied text and topics as well as visual material. Podcastfy can combine multiple sources into one conversation and produce short episodes or longer discussions. Topic-based generation can use web search to gather supporting material.
You can adjust the language, conversational style, structure and voices. That makes it useful for turning an article into a discussion, adapting lecture material for listening or creating audio explanations of research. Multilingual output includes French and Brazilian Portuguese.
The Python package and command-line interface support automated workflows, and the project includes Docker support. Transcript generation can use local HuggingFace models or cloud models from OpenAI, Anthropic and Google. The local option keeps transcript generation on your own hardware; choosing a cloud model sends that stage to an external service.
Audio generation is a separate stage that uses speech services from OpenAI, Google, ElevenLabs or Microsoft Edge. Podcastfy supports multi-speaker speech, so local transcript generation doesn't make the whole audio workflow offline.
Claim this page and we'll verify you by hand. Podcastfy gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Podcastfy?Promote it
Something wrong or outdated on this page?
1.5KUpdated 4 months agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API
Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
19.4KUpdated 10 months agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input#Voice cloning
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
214Updated 3 weeks agoMIT
Linux · Docker · Web#Home Assistant integration#Hugging Face integration#Multilingual
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
Wyoming Piper connects Piper's local text-to-speech engine to Home Assistant and other clients that use the Wyoming protocol. It's for people building a voice assistant on their own hardware who need speech generation as a self-hosted service. The project is open source under the MIT license.
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.