Episodes from AI Engineer about Image Generation.

2026 State of AI Engineering — Barr Yaron, Amplify Partners
Jul 21, 2026 · 19:47
In the 2026 State of AI Engineering survey presented by Amplify Partners' Barr Yaron, 1,048 AI engineers reveal that cost is now a first-class engineering constraint—40% say it regularly shapes how ambitiously they use AI. Agents have exploded: 95% of teams now use agents, and 89% of those agents have write access, tripling from last year. Image generation adoption doubled to 36%, while audio shows the strongest intent-to-adopt at 56%. Open-weight models augment rather than replace closed models—45% use open-weight, but over 90% of them also use closed models. Evals remain the top infrastructure challenge, and inference is the most bought layer, while prompt management (61% built in-house) stays close to product logic. Teams report 97% net positive impact, but 59% fear long-term liabilities from AI code, and over a third say non-developers now ship features.

Any-to-Any: Building Native Multimodal Agents - Patrick Löber, Google DeepMind
May 20, 2026 · 16:21
Patrick Löber, a member of the technical staff at Google DeepMind, explains how to build native multimodal agents using the Gemini API ecosystem, covering multimodal understanding, native image and speech generation, and real-time interaction via the Live API. He demonstrates constructing a NotebookLM clone as an agentic system where a reasoning Gemini model decides whether to generate an infographic or podcast-style audio using function calls to specialized models like Nano Banana for images and a text-to-speech model for speech. The episode details practical implementation: uploading PDFs, video, and audio files; using context caching to reduce costs by 90%; and generating infographics or multi-speaker audio directly from prompts. Löber highlights that native generation models understand world context—like drawing arrows on a map to produce the Golden Gate Bridge—and that the Live API enables audio-to-audio interactions with a single architecture, supporting multiple languages and accents.

FLUX, Open Research, and the Future of Visual AI — Stephen Batifol, Black Forest Labs
May 8, 2026 · 22:32
Black Forest Labs (BFL), the team behind Stable Diffusion and Latent Diffusion, has released a series of open FLUX image models pushing toward visual intelligence. After FLUX.1 and the first open-source editing model FLUX Kontext (7–8 second edits), FLUX.2 achieved state-of-the-art text-to-image and multi-reference editing, and FLUX.2 Klein generates and edits in 300–500 milliseconds for near real-time use. BFL also published Self-Flow, a scalable self-supervised approach that trains multimodal models across images, video, audio, and actions without external encoders, outperforming baselines in all modalities. The episode explains how Self-Flow reduces artifacts (e.g., corrects text rendering and anatomy) and converges 70× faster with representation alignment. BFL’s roadmap includes world models that simulate geometry and interaction, aiming to train agents for robotics and automation.
Powered by PodHood