Sherpa-ONNX audio transcription through Claude and Cowork

Learn two ways to transcribe audio with Claude: a temporary web sandbox and a reusable Cowork skill using downloaded speech models and timestamps.

Player not loading? Watch on YouTube

This tutorial shows how the presenter adds speech-to-text processing to Claude through code execution. He says Claude cannot transcribe attached audio out of the box and suggests Gemini for a quicker option with audio understanding.

The web and mobile method runs in Claude's server sandbox, rather than on the user's computer. The setup enables code execution, file creation and network egress, then uses a supplied prompt alongside an uploaded audio file. The presenter recommends Opus, though he suggests Sonnet may work. Further files can reuse the environment within that chat; a new chat requires another setup and model download. He does not establish how long the sandbox persists.

The desktop method uses Claude Cowork to create a transcription tool and install its SKILL.md file through the customization menu. Python and FFmpeg are dependencies. The exact prompts supplied with the video install Sherpa-ONNX, download models from its k2-fsa release repository, and use its OfflineRecognizer to run Parakeet for English or Whisper for other languages. Claude configures and calls this speech engine; the speech-to-text model is separate from Claude itself. This method uses a local AI speech model stored in a connected folder.

Folder access is a practical limitation: the new-session demonstration stops because the model folder is not connected. Installing the skill carries instructions, not the downloaded model. The final example produces timestamped text. The presenter notes that names and brands can still be wrong and suggests asking Claude to correct the transcript using context.