Player not loading? Watch on YouTube
This tutorial introduces FastEmbed for local AI embedding generation in Python. The presenter uses a ThinkPad without a dedicated GPU and describes the library as a lightweight option built around ONNX Runtime, without PyTorch or TensorFlow dependencies. Installation uses uv, with pip given as an alternative.
The text example creates a TextEmbedding instance, embeds three sentences, and converts the returned generator into a list and NumPy array. A distance matrix compares references to Apple, bananas, and Microsoft. In the reported results, the two technology company sentences sit closer together than the Apple and banana sentences. The tutorial then lists supported models and selects another model. FastEmbed accepts a defined model list rather than arbitrary Hugging Face repository identifiers.
For images, the presenter passes two JPEG paths to ImageEmbedding and prints the vectors, using Qdrant.ResNet50.ONNX as an example. An in-memory Qdrant client demonstrates FastEmbed integration through model selection and document retrieval. This example retrieves context; it does not generate an answer.
The final section explains FastEmbed-GPU and explicit CUDA provider configuration for an Nvidia GPU. The presenter cannot demonstrate successful GPU execution on this laptop, and the attempt fails because CUDA is unavailable. He also questions the library's state-of-the-art claim and recommends it for prototyping and lightweight workloads.