Sentence Transformers and FAISS: semantic search and RAG

Learn embeddings, cosine similarity and RAG by indexing 1,154 One Piece synopses with FAISS, then comparing Qwen 2.5 3B answers with retrieval.

Player not loading? Watch on YouTube

This tutorial builds semantic search over 1,154 One Piece episode synopses, then connects the results to Qwen 2.5 3B Instruct for question answering. The speaker explains sentence embeddings, mean pooling and cosine similarity before walking through Python scripts in VS Code.

The example uses Sentence Transformers with all-MiniLM-L6-v2 to encode the synopses. It normalizes the vectors and stores them in a FAISS index, where inner product gives cosine similarity for normalized vectors. Queries pass through the same embedding process. A search for the whale at the entrance of the Grand Line retrieves references to Laboon without naming him, while another example compares semantic retrieval with substring matching.

The RAG script adds retrieved synopses to the model's prompt. In the speaker's comparison, the model without retrieval confuses Captain Kuro with Blackbeard and gives an incorrect episode number. With context, it returns episode 12, which the speaker identifies as correct. Results remain uneven: a question about Luffy first meeting Zoro fails, while a Sanji cooking query finds a relevant episode. The demo shows how retrieval can help a local LLM answer questions about a dataset, but its usefulness depends on the documents it finds.