Player not loading? Watch on YouTube
Paul McWhorter's lesson 40 adapts an earlier Raspberry Pi face tracker to MediaPipe. The Python project combines camera detections with voice commands: "track" enables pan tilt servo tracking, while "release" turns the servos off. The supplied setup uses a Raspberry Pi 5 and SunFounder's AI Fusion Lab Kit with a prepared Raspberry Pi Bookworm image containing the class dependencies.
The tutorial uses the older mp.solutions.face_detection API in the supplied class environment. It replaces OpenCV Haar cascade detection with MediaPipe face detection, selecting model 0 for close faces and a minimum detection confidence of 0.5. McWhorter converts camera frames from BGR to RGB, turns relative bounding box coordinates into pixels, and calculates the face's offset from the frame center. He also fixes a leftover reference to the old face variable and disables tracking when detection disappears.
McWhorter reports about 33 frames per second in his demonstration, compared with four or five for the earlier unoptimized Haar cascade code and 12 after improvements. He attributes the demonstrated ceiling to camera speed; these figures describe his setup rather than a general benchmark.
The final exercise adds "blind" and "restore" commands. These toggle black circles over the eyes in the displayed image. He identifies eye keypoints 0 and 1 and draws filled circles with a 20-pixel radius. This is an on-device computer vision exercise, with spoken responses and servo control.