Cactus review: local transcription vs Gemini 2.5 Flash

Learn how Cactus handles edge inference and compare reported transcription latency: about 260 ms locally on an iPhone 12 Pro versus cloud batch processing.

Player not loading? Watch on YouTube

Cactus is an inference engine for mobile and edge devices. This review explains its approach to memory use and tests on-device speech transcription in a Swift app on an iPhone 12 Pro.

The speaker describes zero-copy memory mapping that reads model weights from storage and brings tensors into active computation as needed. Cactus uses a proprietary .cact format in place of GGUF for this approach. The review also describes an NPU-first design and a dashboard with models optimized for neural processing units. These are the speaker's explanations of the architecture, rather than independently verified memory or battery measurements.

Cactus's hybrid router uses confidence-based routing to choose between local inference and a cloud model. According to the speaker, requests that exceed the local model's capabilities or need a larger context window can move to the cloud without changing the application code.

The transcription demo uses the Swift Cactus package with Parakeet locally and Gemini 2.5 Flash in the cloud. The speaker reports roughly 260 ms latency for local streaming, compared with about 2,000 ms for a three-second cloud transcription batch. The comparison uses different models and processing modes: the speaker says Parakeet was unavailable on the cloud side. The landing page's roughly 100 ms demo claim is separate from the result reported on the older iPhone.