Player not loading? Watch on YouTube
This tutorial pairs MiniCPM-V 4.6 with Gemma 4 E2B to turn a rough dashboard sketch into a single HTML file. Ollama runs the models, while a Python app uses Streamlit for image uploads, chat history, streamed code and downloads. The workflow is sequential: MiniCPM describes the image layout, then Gemma uses that description and the user's instructions to generate HTML. Text-only requests go directly to the coder.
The presenter chooses 4-bit versions on a machine with an Nvidia GPU and 4 GB of VRAM. His approximate memory estimates are 1 GB for MiniCPM and 1.2 GB for Gemma; these are tutorial estimates rather than measured requirements. An Ollama Modelfile configures the coder with a temperature near zero, an 8,000-token context window and a 4,000-token output limit. This custom configuration supplies instructions rather than demonstrating weight fine-tuning.
The demo uses a sketch generated with Gemini and requests a dark sales dashboard. The presenter shows the downloaded HTML with navigation, a daily/monthly graph toggle and hover effects. He describes the installed workflow as offline and discusses on-device support for iOS, Android and HarmonyOS, though the walkthrough uses a desktop. Benchmark advantages are claims presented in the video, not independent verification.