Tools tagged with "Speaker diarization"

21 tools
Local-first AI video editor for macOS, Windows and Linux. Keep projects on your machine and edit through chat or a multitrack timeline. Free under AGPL-3.0.

2.1KUpdated 7 hours agoAGPL-3.0

macOS · Windows · Linux#Agent Skills#MCP#Multilingual

Favicon of Wan2GP

Wan2GP

5 videos
Local AI media generator with a browser interface, support for NVIDIA and AMD GPUs, and select models that run with 6 GB of VRAM.

9.7KUpdated 1 day ago

macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF

Local AI memory captures screen activity and audio on macOS, Windows, and Linux, with searchable history, on-device models, and MCP access for agents.

21.8KUpdated 21 hours ago

macOS · Windows · Linux#MCP#Multimodal input#Ollama integration

Favicon of VibeVoice

VibeVoice

2 videos
Open-source voice AI models for local transcription and speech generation, with MIT licensing, CPU inference, and streaming audio support.

54.5KUpdated 4 weeks agoMIT

#Hugging Face integration#Multilingual#Quantization

Offline speech-to-text app for iPhone, iPad and Apple Silicon Macs, with local speaker labels, transcript summaries on Mac and no account required.

whispernotes.appChat With Your Documents

macOS · iOS#MCP#Multilingual#Speaker diarization

Open-source dictation app for macOS, Windows, Linux and iOS. Transcribe offline with local models or choose cloud processing with your own API keys.

8.9KUpdated 10 hours agoMIT

macOS · Windows · Linux · iOS#Batch processing#MCP#Multilingual

A local audio transcription CLI that runs Whisper and Distil-Whisper on NVIDIA GPUs or Apple Silicon Macs. Open source under Apache 2.0.

13.1KUpdated 2 years agoApache-2.0

macOS · Windows#Batch processing#Hugging Face integration#Multilingual

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

5.1KUpdated 21 hours ago

macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API

Open-source speech recognition model for Mandarin, Cantonese, English, Japanese and Korean. Runs locally on CPU or GPU under the MIT license.

9.4KUpdated 3 weeks agoMIT

Docker#Batch processing#GGUF#Hugging Face integration

Local AI transcription app for macOS that keeps audio on your device, supports Whisper, Qwen and Parakeet, and offers optional cloud services.

goodsnooze.gumroad.comDictation and Voice Typing

macOS · iOS#Batch processing#Multilingual#Ollama integration

Favicon of Scriberr

Scriberr

1 video
Self-hosted audio and video transcription with Whisper, NVIDIA Parakeet and Canary. Runs offline on your hardware, with optional Ollama transcript chat.

3.1KUpdated 1 week agoMIT

macOS · Linux · Docker · Web#Ollama integration#OpenAI-compatible API#Speaker diarization

A self-hosted Python toolkit for curating AI training data across text, images, video and audio, with NVIDIA GPU support and an Apache 2.0 license.

1.8KUpdated 19 hours agoApache-2.0

Linux · Docker#Distributed execution#Hugging Face integration#Multilingual

A local speech-to-text browser app with Whisper backends, subtitle translation and speaker labeling. Open source under Apache 2.0, with Docker support.

2.9KUpdated 9 months agoApache-2.0

Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization

An open-source video translation and AI dubbing tool for Windows, macOS and Linux, with local offline models or cloud APIs. Licensed under GPL-3.0.

19.2KUpdated 2 days agoGPL-3.0

macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual

Favicon of WhisperX

WhisperX

1 video
Open source speech-to-text software that runs locally, aligns transcripts word by word, and can label speakers.

24.3KUpdated 4 days agoBSD-2-Clause

macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual

Offline transcription and translation software runs Whisper on Windows, Linux and macOS, with microphone capture and subtitle export.

21.8KUpdated 1 week agoMIT

macOS · Windows · Linux#Hugging Face integration#Multilingual#Speaker diarization

Open source data labeling and AI evaluation platform that runs locally or on your server, with custom annotation interfaces and model-assisted labeling.

28.4KUpdated 1 day agoApache-2.0

macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input

Favicon of Vibe

Vibe

1 video
Local transcription app for macOS, Windows and Linux. Uses Whisper, Nemotron and Parakeet, with Ollama analysis and optional Claude API summaries.

7.7KUpdated 4 weeks agoMIT

macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration

Open-source meeting bot API for Meet, Teams and Zoom, with live transcripts and AI agent access. Self-host with Docker or use the hosted service.

2.8KUpdated 2 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#Git integration#Human approval

Self-hosted speech-to-text API that runs in Docker on CPU or CUDA GPUs, with Whisper, Faster Whisper and WhisperX. Open source under MIT.

3.3KUpdated 2 months agoMIT

Docker · Web#Multilingual#Speaker diarization#Voice activity detection

Open-source speaker diarization toolkit with local PyTorch models, CUDA GPU support, and an optional hosted service that processes audio on pyannoteAI servers.

10.6KUpdated 3 months agoMIT

#Hugging Face integration#Speaker diarization#Voice activity detection