← back

Prompt to Pipeline: Building with Google's Gen Media Stack — Paige & Guillaume, Google DeepMind

3.0K views · May 23, 2026 · 114:35 min · Watch on YouTube ↗
Takeaway

Build media workflows around customer needs while accounting for rapid changes in general model capabilities.

Summary

  • The workshop introduces Google’s generative-media stack through AI Studio and Antigravity examples.
  • The opening previews Gemini Flash Live for real-time conversation, Pro and Flash-Lite models, Nano Banana image generation/editing, and multimodal embeddings.
  • Paige traces an open-source path through scientific Python, TensorFlow, and Google’s model teams.
  • The discussion argues that durable product value comes from opinionated customer workflows as general models absorb capabilities previously requiring specialized tooling.
google-deepmindgenerative-mediagemini
Original description
A public domain book, a notebook, and three gen media models. Guom from Google DeepMind fed Wind in the Willows into Gemini, generated character portraits with Nano Banana, animated chapter scenes with VO, and scored each chapter with LIA, all live in the workshop.

The full three hour session covers more ground. Paige Bailey demos AI Studio's Build feature creating a bookshelf scanning app with Google login and Firestore from a single prompt, Gemini 3.1 Flash Light analyzing a dinosaur video frame by frame for under a dollar, and Genie 3 rendering a playable world with a pink sparkly squirrel on Regent's Canal. Ian Valentine closes with Gemma 4 running on device: 10 sub agents generating SVGs in parallel on a local 26B model, then open code building and debugging a game from a spec with no cloud API involved.

Speaker info:
https://x.com/DynamicWebPaige
  / dynamicwebpaige  
https://github.com/dynamicwebpaige
https://x.com/Giom_V
  / guillaumevernade  
https://github.com/Giom-V