Unstructured tutorial: extract email and PowerPoint data

Learn how Unstructured partitions email and PowerPoint files in Python, with a Google Colab demo and an overview of S3 ingestion.

Player not loading? Watch on YouTube

Unstructured is the focus of this tutorial on preparing document text for LLM workflows. Uploaded on September 23, 2023, the video does not specify a package version or recording date. Its installation steps and connector examples reflect that period.

The presenter describes extracting files locally, though the worked examples use Google Colab with uploaded sample files. After installing Unstructured, he imports the partition function, passes it an email file and loops through the returned elements. He compares the output with an email containing headers, body text and footer material, saying the extraction keeps the content he needs.

A second example processes a two-slide PowerPoint presentation. The presenter installs an additional Unstructured module for PowerPoint, partitions the uploaded file and prints its elements. He also mentions PDF, text and image support, but does not demonstrate those formats.

The final section outlines S3 ingestion with source paths and an output location, then discusses Python, shell and REST API options. A brief documentation tour points to SharePoint examples. For a local AI pipeline, the relevant lesson is document extraction before model use: the tutorial does not run a local LLM, generate embeddings or demonstrate retrieval. The connector discussion is an overview rather than a completed cloud setup.