Favicon of Podcastfy

Podcastfy

Open-source AI podcast generator in Python with local HuggingFace models for transcripts and cloud speech services for multilingual audio.

Podcastfy turns documents, websites and images into AI-generated audio conversations, with the option to write transcripts using a local LLM. It's an open-source Python alternative to NotebookLM's podcast feature for creators, educators and researchers who want control over the conversation format or need podcast generation inside their own software. It uses the Apache License 2.0.

Inputs can include PDFs, YouTube videos, supplied text and topics as well as visual material. Podcastfy can combine multiple sources into one conversation and produce short episodes or longer discussions. Topic-based generation can use web search to gather supporting material.

You can adjust the language, conversational style, structure and voices. That makes it useful for turning an article into a discussion, adapting lecture material for listening or creating audio explanations of research. Multilingual output includes French and Brazilian Portuguese.

The Python package and command-line interface support automated workflows, and the project includes Docker support. Transcript generation can use local HuggingFace models or cloud models from OpenAI, Anthropic and Google. The local option keeps transcript generation on your own hardware; choosing a cloud model sends that stage to an external service.

Audio generation is a separate stage that uses speech services from OpenAI, Google, ElevenLabs or Microsoft Edge. Podcastfy supports multi-speaker speech, so local transcript generation doesn't make the whole audio workflow offline.

Similar to Podcastfy