InternVL 2.5 1B local setup and image tests

Learn how the presenter loads InternVL 2.5 1B in Jupyter and tests image descriptions and OCR, with reported VRAM use of 4.6 GB.

Player not loading? Watch on YouTube

Uploaded on December 7, 2024, this tutorial covers InternVL 2.5's 1B variant. The presenter installs the multimodal model in a GPU VM and uses a Jupyter notebook to test local image inputs. The setup and capabilities shown belong to that dated demonstration.

The walkthrough starts with a virtual environment and dependencies including PyTorch, torchvision and Transformers. It then explains image preprocessing functions from the Hugging Face model card, including aspect ratio selection and image loading, before downloading the model and tokenizer. The VM has 48 GB of VRAM, but the presenter reports about 4.6 GB in use during testing. That observation is not a minimum hardware requirement. He suggests CPU use and quantization as options, but does not demonstrate them.

The local LLM tests cover a sunset illustration, multilingual OCR, a faucet and a colorful mural. The presenter judges the descriptions favorably, while noting a questionable color detail and an OCR spelling error. The model identifies the faucet as a kitchen fixture; he believes it is a laundry fixture. These examples give viewers concrete successes and errors to assess rather than a systematic benchmark comparison. Video processing appears in the setup explanation, but the inference tests use still images.