Apple Core ML 7: Stable Diffusion performance comparison

Learn how Apple's WWDC 2023 updates used 6-bit weights for Stable Diffusion, with CPU, GPU and Neural Engine output comparisons.

Player not loading? Watch on YouTube

Uploaded on June 15, 2023, this overview discusses Apple's Stable Diffusion Core ML implementation, linked as apple/ml-stable-diffusion in the source, and its WWDC 2023 presentation on Stable Diffusion for Apple silicon. The version context is Core ML 7, iOS 17 and macOS 14. It covers that period's announcements and benchmarks rather than a current setup guide.

The speaker explains how weight compression could reduce the memory needed for on-device image generation. He reports that Apple's 6-bit approach brought the demonstrated Stable Diffusion model below one gigabyte, alongside roughly 30% better Neural Engine performance. These are reported results for the implementation discussed, not guarantees for every model or device.

Image examples compare CPU plus Neural Engine execution with CPU plus GPU execution. The speaker judges their output to be similar, while acknowledging differences between 16-bit and 6-bit weights. He also discusses pruning, palettization and linear 8-bit quantization.

The technical explanation focuses on decompressing weights during runtime instead of expanding them all when loading the model. The stated behavior depends on the compute unit, layer type and hardware generation. Benchmarks discussed include the iPhone 14 Pro and M2 Ultra. SnapFusion provides a comparison with Snapchat's mobile implementation; the speaker suggests possible further speed gains, but does not demonstrate them here.