Player not loading? Watch on YouTube
Cormac Brick explains how LiteRT-LM supports on-device language models, contrasting larger models customized through prompts and skills with tiny models fine-tuned for specific app tasks. He describes Gemma 4 E2B and E4B as models whose effective resident parameter counts address RAM constraints. Their function calling and thinking capabilities underpin the skill examples.
AI Edge Gallery demonstrates chat, image questions and audio workflows on Android and iOS. Its skill architecture initially exposes short descriptions, then loads detailed instructions and JavaScript when needed. Brick says this reduces context overhead for smaller models. Constrained decoding limits tool-call output to supported tools. The demonstrations use the 4B model; he cautions that the 2B model may need simpler or fewer skills. Some skills run offline, while others call web services and require an API key.
The deployment discussion covers conversion and quantization through LiteRT Torch. CPU and GPU deployment can share an artifact, while NPU acceleration requires vendor-specific compilation. At the time of the talk, Brick describes Swift APIs as forthcoming and LiteRT-LM web support as work in progress.
For tiny models below one billion parameters, he recommends narrow tasks and fine-tuning. AI Edge Eloquent illustrates the approach with separate speech recognition and text polishing models. Brick describes the app as iOS-only; an attendee reports that it is unavailable in the European App Store.