Player not loading? Watch on YouTube
Pete Warden demonstrates Moonshine Micro, a small version of the Moonshine voice library, on a Raspberry Pi Pico 2W with an RP2350 chip. He reports that the complete voice interface uses 520 kilobytes of RAM and 4 megabytes of flash. Speech processing runs offline while the device guides the user through connecting to Wi-Fi.
The demonstration starts with a spoken command, then asks the user to spell part of the network name and the password. The device repeats recognized characters, confirms the network and reads out its IP address after connecting. Warden says the recognition model handles about 50 commands, including letters, digits, password symbols and a delete command for corrections. This is a limited command interface, rather than a conversational chatbot.
The on-device system also uses voice activity detection and neural text-to-speech. Warden describes the synthetic voice as clearly nonhuman and cautions that the successful recording does not guarantee future recognition accuracy.
The open source project includes code, models, documentation and a training recipe for custom command sets. The hardware example uses a microphone, amplifier and speaker; Warden also describes using a Mac or PC's microphone and speakers to feed the RP2350. He says other chips are not yet running the system, though he hopes to see ports to more boards.