FluidVoice: offline macOS dictation architecture explained

Learn how FluidVoice handles dictation on Apple Silicon with two processing layers, vocabulary overrides, and macOS 15.0 or later requirements.

Player not loading? Watch on YouTube

FluidVoice is presented as an open source macOS dictation app that processes speech offline. The walkthrough explains its architecture rather than providing a step-by-step installation guide. The speaker says audio and text stay on the user's hardware and describes a zero-telemetry environment.

The stated requirements are Apple Silicon M1 through M4 and macOS 15.0 or later. Acoustic models remain loaded in memory to reduce startup delays. According to the speaker, a full local installation occupies nearly 5 GB of active RAM; a compressed 1 GB mini model for Macs with 8 GB of memory is described as under development.

The pipeline separates speech recognition from text cleanup. Core ML acoustic models produce raw text, then Fluid-1, a proprietary local LLM, handles punctuation, hesitation markers, and corrections. The speaker describes Fluid-1 as a 3.5 GB model trained on 100,000 dictation examples. A cited benchmark favors it over GPT 5.4, but the transcript does not explain the evaluation method.

User-defined vocabulary replacements run before AI processing. The walkthrough describes a JSONC mapping file that reloads after edits without a daemon restart. FluidVoice then uses macOS Accessibility access to type into the active window. It also discusses microphone permission failures and a terminal-based reset, though the transcript supplies no command to reproduce that fix.