← back

How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads

2.5K views · Jul 24, 2026 · 19:29 min · Watch on YouTube ↗
Takeaway

Develop evaluation criteria from concrete failures and human agreement before scaling the rating process.

Summary

  • YouTube Ads recommends optimizing a focused set of LLM-friendly tools before evaluating a larger agent, with critique-and-remediation loops filling gaps.
  • Early manual inspection helps teams discover failure patterns and make major architectural changes before investing in scaled evaluation.
  • Begin with a small set of important tasks and negative checks, then expand the golden set as understanding improves.
  • Scaled raters need clear rubrics, examples, strong human agreement, and explanations that distinguish dimensions such as accuracy and brand safety.
agent-evalsrubricsyoutube-ads
Original description
Getting an AI agent to behave the way you want isn't just about writing better prompts. In real systems, behavior emerges from a loop: prompts, evals, iteration, and feedback. Small changes in any part of that loop can completely change outcomes.

The Google team shares lessons from building a seed-asset agent that turns messy advertising creatives — low-quality images, cluttered visuals, and heavy text overlays — into clean, reusable assets for downstream generative AI tools. They explain why prompting alone did not produce stable behavior, how evals became feedback signals rather than scorecards, how agent trace logs exposed why failures happened, and how they iterated without breaking problems they had already fixed.

Speakers:

Chris Souza — Google
Chris works on the Google team behind this seed-asset agent and its evaluation workflow.

Preetika Bhateja — Product Manager, Google/YouTube
Preetika works on ads, evaluations, agents, and LLM-as-judge systems.

Daniel Bump — Engineer, Google
Daniel focuses on image and video generation and computer vision.
X/Twitter: https://x.com/DanielJBump
LinkedIn:   / danielbump