DSPy and GEPA: how prompt optimization works

Learn how DSPy selects examples and GEPA revises instructions, with reported math accuracy rising from 24% to 65% and limits around evaluation.

Player not loading? Watch on YouTube

DSPy replaces manual prompt revisions with a process driven by examples and evaluation. The speaker explains signatures through a support-ticket task: define the input and the urgency level you want back. Switching from prediction to chain of thought changes how DSPy builds the underlying instructions.

Compilation needs example inputs paired with desired answers and a metric that scores outputs. In the math example, DSPy keeps successful worked solutions and selects a few to include in the prompt. The goal is to find examples that help with unseen questions. The speaker reports that the first paper raised accuracy from 24% to 65% on a grade-school math task using the same model, without fine-tuning.

GEPA extends the approach by examining failed runs and their errors. A second model writes feedback, which guides revisions to the instructions. The speaker reports gains of up to 20% over reinforcement learning with as little as 1/35 of the attempts. These are reported experimental results, rather than a guarantee for other tasks.

The review spends its final section on practical limits. A poor metric can reward long, incorrect answers. Signature field names still influence output, and generated prompts can be difficult to debug. The speaker also describes mixed developer reactions to the framework. This is a conceptual explanation of prompt optimization, with no local deployment walkthrough or hardware requirements.