Player not loading? Watch on YouTube
Prince Canuma introduces MLX, an array framework for Apple Silicon, through accessibility projects and demonstrations of on-device AI. His motivation comes from helping his blind father access information where internet connections are unreliable.
The talk covers MLX VLM for image understanding and MLX Audio for speech recognition, speech synthesis and speech-to-speech applications. Canuma describes Python and Swift support, plus a modular voice pipeline where developers choose separate recognition, language and speech synthesis models to fit their hardware. He says Marvis can generate audio in less than 100 milliseconds.
Demonstrations include object detection, background blur and a Gemma 4 image chat interface using Gradio. The chat demo encounters startup trouble and attempts to download model files before producing an answer. Canuma reports that his Mac has 96 GB of VRAM, a useful qualification when assessing his performance claims. Community examples include a native voice app and chained video generation that he says can run on a MacBook with 16 GB of video RAM.
In the Q&A, Canuma explains that MLX uses the GPU and points to Core ML for Neural Engine access. He recommends Mactop for monitoring usage and cautions that model performance depends on the task. He also claims Turbo Quant reduces KV cache memory by 4x and allows a million-token context, depending on model size and hardware.