Favicon of AudioNotes

AudioNotes

Local AI transcription and note-taking software runs through Docker or Python, using FunASR and Ollama to keep recordings and notes on your computer.

AudioNotes turns audio and video into transcripts and structured Markdown notes on your own computer. It's for people reviewing meetings, interviews, courses or voice memos who want to ask questions about a recording without sending it to a cloud AI service. The project is open source under the MIT license.

FunASR handles speech recognition locally, while Ollama runs the model that writes notes and answers follow-up questions. Recordings, transcripts, notes and conversation history stay on your machine. Internet access is needed to download the application and models; after those downloads, processing doesn't depend on cloud AI services. The browser interface accepts connections from your computer only by default.

You can upload existing media or record audio in the browser. AudioNotes keeps the full transcript alongside the notes, supports domain-specific vocabulary and saves conversation history so you can return to earlier material. Follow-up questions can focus on details such as responsibilities, dates or conclusions.

The default models are FunAudioLLM/Fun-ASR-Nano-2512 for transcription and qwen3.5:2b for notes and questions, with an emphasis on Chinese recognition and ordinary computer hardware. FunAudioLLM/Fun-ASR-MLT-Nano-2512 provides a multilingual recognition option.

AudioNotes runs through Docker or directly with Python. CPU processing is supported, with optional NVIDIA GPU acceleration on Linux. The Docker setup recommends at least 12GB of memory, or 16GB for longer recordings, plus about 10GB of disk space.

Similar to AudioNotes