Player not loading? Watch on YouTube
Serj and Brett compare three approaches to prompt optimization: DSPy, Codex working directly on prompts, and Codex controlling DSPy. Their live example classifies customer service requests by intent using GPT-4.1 Nano. The comparison uses the same data splits and a locked test set.
The tutorial explains DSPy signatures, which define task inputs and outputs, then introduces inference methods such as chain of thought and majority voting. The speakers discuss few-shot example selection and instruction optimization, including MIPROv2 and GEPA. They use macro F1 to assess an imbalanced dataset and explain why separate training, validation and test data matter.
The combined AI agent approach produces the strongest result reported during the session. That outcome has limits: the speakers initially specify five attempts, but later inspect a sixth run, and the reported scores change as experiments continue. Coding assistant costs also fall outside their cost comparison, so the demonstration does not establish a universal winner or an equal-budget benchmark.
The discussion closes with a production workflow: save DSPy's optimized output, translate it into the serving framework, and evaluate it again on the same data. The speaker proposes this as a way to avoid a DSPy runtime dependency. This session covers application development rather than a local LLM installation or an offline deployment.