Episodes from AI Engineer about Vision Language Models.

Agents on the Canvas in tldraw — Steve Ruiz, tldraw
May 1, 2026 · 19:54
Steve Ruiz of tldraw details the evolution of AI on the infinite canvas, from early one-shot demos like MakeReal (draw to functional prototype) to multi-agent 'fairies' that collaborate, delegate, and even rewrite the canvas in real time. He explains how the tldraw SDK enables agents to act as virtual collaborators, using structured outputs and tool use to draw, animate, and modify designs. Ruiz demonstrates fairy agents that can work independently or in leader-follower mode, coordinating on tasks like creating wireframes or filling out forms. He also unveils a desktop prototype that lets Claude directly inject JavaScript into the canvas, enabling agents to edit code, modify UI, and even hack other apps like Spotify. The talk emphasizes the shift from sidebar agents to canvas-native collaborators, the challenges of vision model training for 2D spaces, and the safety trade-offs of giving LLMs runtime access in local-first apps.

Gemma 4 Deep Dive — Cassidy Hardin, Researcher, Google DeepMind
Apr 27, 2026 · 19:03
Cassidy Hardin, a researcher at Google DeepMind, details the Gemma 4 family of open-source models, claiming they set a new precedent for small-scale performance. The family includes two on-device effective models (E2B and E4B) with Per Layer Embeddings (PLE) stored in flash memory, and two larger models: a 26B mixture-of-experts (MoE) with 128 total experts (8 active) and a 31B dense model that ranked #3 on the LM arena leaderboard, outperforming models 20x its size. Architectural innovations include interleaved local/global attention (5:1 ratio, sliding windows of 512/1024 tokens), grouped query attention (8 queries per key-value in global layers), and native multimodal support with variable aspect ratios and resolutions for vision encoders (550M params for large, 150M for small) plus a 305M parameter conformer for audio. All models are released under Apache 2.0 license, available for self-hosting on Hugging Face and Ollama, or cloud deployment on Vertex AI.

Gemma, DeepMind's Family of Open Models — Omar Sanseviero, Google DeepMind
Apr 20, 2026 · 15:26
Omar Sanseviero presents Gemma 4, Google DeepMind's latest family of open models, which range from 2B to 32B parameters and introduce a novel per-layer embedding (E2B) architecture optimized for on-device inference. The models feature multimodal understanding (images, video, audio), multilingual support across 140+ languages, and are released under an Apache 2 license. Within a week, Gemma 4 reached 10 million downloads, contributing to over 500 million total downloads for the Gemma family and 100,000 community-derived models. Sanseviero highlights official variants like Shield Gemma for content safety and MedGemma for medical tasks, as well as community efforts such as AI Singapore's Southeast Asian language models and Sarvam's sovereign AI initiative for India. He emphasizes real-world applications including cancer therapy pathway discovery and fully offline agentic tasks on phones and Raspberry Pis, arguing that open models are rapidly enabling high-performance, private, and customizable AI across diverse use cases.

VoiceVision RAG - Integrating Visual Document Intelligence with Voice Response — Suman Debnath, AWS
Dec 6, 2025 · 1:23:52
Suman Debnath, a Principal ML Advocate at AWS, demonstrates how Colpali—a vision-based retrieval model that treats each document page as an image and generates multi-vector embeddings via patch-based late interaction—can be combined with voice synthesis for a more intuitive RAG system. He explains that Colpali bypasses traditional OCR and preprocessing by directly embedding document images, then scoring query-page relevance through dot-product similarity across patches. The workshop shows how to embed pages, store them in Qdrant with multivector support, and retrieve top pages. Debnath then wraps this retrieval pipeline using the Strands Agent framework, adding a speak tool to output answers in natural voice. A live demo answers a textbook question about trophic levels, first using Bedrock to generate text and then speaking the answer in a female voice, all without a predefined system prompt. The talk positions Colpali as a complementary technique for complex visual documents like IKEA instructions or scanned forms, not a full replacement for traditional RAG.

Vision AI in 2025 — Peter Robicheaux, Roboflow
Aug 3, 2025 · 17:24
Peter Robicheaux, ML lead at Roboflow, argues vision models are not smart because they fail at basic visual tasks like telling time on a watch or identifying a school bus direction. He attributes this to saturated evals like ImageNet and COCO, and to vision models not leveraging large pre-training like language models. Roboflow introduces RFDetter, a real-time detector using a Dynav2 backbone, achieving second SOTA on COCO and large gains on their new RF100VL dataset of 100 diverse domains. The benchmark shows specialist models fine-tuned on 10-shot examples outperform large VLMs like Qwen2.5 VL 72B, indicating VLMs are strong linguistically but weak visually. Roboflow's platform is free for researchers contributing data back; dataset at rf100vl.org.

What Is a Humanoid Foundation Model? An Introduction to GR00T N1 - Annika & Aastha
Jul 28, 2025 · 17:47
NVIDIA's Annika Brundyn and Aastha Jhunjhunwala introduce GR00T N1, an open-source Vision-Language-Action foundation model for humanoid robots. They argue that physical AI is key to addressing labor shortages in industries like healthcare and manufacturing, and that humanoid forms are needed because the world is built for humans. The model uses a data pyramid strategy combining limited real-world teleoperation data, synthetic simulation data, and internet video, with DreamGen for data multiplication. Its dual-system architecture, inspired by Kahneman's 'Thinking, Fast and Slow', has a System 2 planner for high-level reasoning and a System 1 for fast motor control at 120 Hz. The model is trained end-to-end via imitation learning and reinforcement learning, and features a generalist action decoder that enables cross-embodiment fine-tuning.

Robotics: why now? - Quan Vuong and Jost Tobias Springberg, Physical Intelligence
Jul 26, 2025 · 18:07
Physical Intelligence's Quan Vuong and Jost Tobias Springberg describe their mission to build a model that can control any robot to do any task, arguing that software intelligence is the main bottleneck in robotics. They explain Vision Language Action models (VLAs) as adaptations of vision language models that output robot actions instead of text. To train these models, they built a data engine from scratch, collecting 10,000 hours of successful episodes via teleoperation in six months. Their latest model, PAIO-5, achieves open-world generalization by training on data from multiple homes, matching or surpassing performance on held-out scenes. They demonstrate this with a policy that performs long-horizon tasks like cleaning an unseen bedroom for up to 10 minutes autonomously. They also highlight a remote coffee-making demonstration on a robot they never touched, showing model portability across hardware.

Building Multimodal AI Agents From Scratch — Apoorva Joshi, MongoDB
Jun 27, 2025 · 36:58
Apoorva Joshi, an AI developer advocate at MongoDB, leads a hands-on workshop on building multimodal AI agents from scratch. She defines agents as systems using LLMs to reason, plan, and execute tasks with tools, contrasting them with simple prompting and RAG. The workshop focuses on processing mixed-media documents (text and images) by converting each page to a screenshot, embedding it with Voyage AI's multimodal (VLM-based) model to avoid the modality gap of CLIP, and storing embeddings in MongoDB for vector search. The agent uses Google's Gemini 2.0 Flash as the multimodal LLM, implements REACT (reasoning and acting) for planning with feedback, and manages short-term memory via session IDs. Participants build a tool-calling agent that retrieves relevant page screenshots, passes them with conversation history to Gemini, and generates answers for queries like analyzing charts or extracting insights from documents.

Unveiling the latest Gemma model advancements: Kathleen Kenealy
Feb 9, 2025 · 16:25
Kathleen Kenealy, technical lead of the Gemma team at Google DeepMind, unveils the latest advances in the Gemma model family, including the launch of Gemma 2 in 9B and 27B parameter sizes, which outperform models two to three times larger, such as LLaMA 3 70B. She also introduces PALI Gemma, a multimodal model combining Siglip Vision Encoder with Gemma 1.0 for image-text tasks. The episode highlights Gemma's responsible-by-design approach, broad framework support (TensorFlow, Jax, PyTorch, etc.), and the release of the Gemma cookbook with 20 recipes. Kenealy emphasizes that Gemma 2 is optimized for easy integration and fine-tuning, available on Google AI Studio, and invites the community to build and share their projects.

Multi model multimodal and multi agent innovations in Azure AI: Cedric Vidal
Feb 6, 2025 · 28:56
Cedric Vidal, Principal AI Advocate at Microsoft, demonstrates Azure AI's multi-model, multimodal, and multi-agent capabilities in a session packed with live demos. He shows how GPT-4 Omni mixes text and vision to read handwritten French menus and translate them, and diagnoses infrastructure damage from photos for energy and insurance industries. A new video translation service lets him speak German, Spanish, Italian, and Japanese in his own voice while preserving tone, such as whispering or yelling. The Azure AI model catalog now offers 1,600 models, including serverless deployment options, and Phi-3 Vision, a small 3.8B parameter model running locally in a browser via WebGPU. He also demonstrates code interpreter analyzing a GPX file from a kitesurfing session, plotting a map with turn markers, and GitHub Workspaces preview generating a Java GUI from Python code.

EyeLevel Launch: Your RAG is Tripping, Here's the Real Reason Why
Feb 6, 2025 · 6:13
Benjamin Over, co-founder of EyeLevel.ai, argues that 95% of RAG hallucinations stem from content ingestion problems, not LLMs or prompts, and introduces his platform's approach to solve this. EyeLevel avoids vector databases entirely, instead creating semantic objects with auto-generated metadata and rewriting text for search and completion. Its pipeline uses a fine-tuned vision model to extract tables, figures, and text, then applies nine fine-tuned models for ingestion and search, including a re-ranking LLM. Air France uses EyeLevel to build a call-center co-pilot, achieving over 95% accuracy on hundreds of thousands of complex documents. In a study, EyeLevel reached 98% accuracy on real-world documents, outperforming popular solutions by up to 120%.

Moondream: how does a tiny vision model slap so hard? — Vikhyat Korrapati
Nov 14, 2024 · 19:26
Vikhyat Korrapati built Moondream, a tiny open-source vision language model under 2 billion parameters that matches LLaVA 1.5, a model four times its size, on VQA v2 and GQA benchmarks. He attributes this to focusing on image understanding over world knowledge and investing heavily in synthetic training data—a pipeline that generated 35 million images and used two orders of magnitude more compute on data than training. Key lessons: community engagement was critical, open-source builds trust, and safety guardrails should be application-layer, not baked in. He argues tiny models will dominate production due to cost and privacy advantages, and that prompting will replace custom model training for most vision tasks. Moondream raised a seed round from FullySysAscent and the GitHub Fund.

The Multimodal Future of Education: Stefania Druga
Oct 24, 2024 · 20:05
Stefania Druga, a research scientist on Google's Gemini team, argues that multimodal AI can transform education by making learning more interactive and critical. She presents Cognimates, an open-source platform she built that expands Scratch to let kids train custom AI models and program smart devices, showing a longitudinal study where kids became more skeptical of AI's intelligence after tinkering. Druga also introduces a benchmark for math misconceptions and demonstrates live demos using the Gemini API—a science tutor that responds to drawings, a math coach that avoids giving direct answers, and a curiosity-driven object identifier—all in under 100 lines of code. She emphasizes the need to cultivate AI literacy early, as 70% of generative AI users are Gen Z, and invites engineers to build tinkering tools that balance agency between learners and AI.

From Text to Vision to Voice Exploring Multimodality with Open AI: Romain Huet
Jul 10, 2024 · 23:39
OpenAI's Romain Huet demonstrates GPT-4o's real-time audio, vision, and screen-sharing capabilities at the AI Engineer World's Fair, showing how the omni-model achieves human-like conversational latency, understands emotion, and can be interrupted naturally. He walks through live demos of ChatGPT Desktop recognizing drawings, reading book pages, and pair-coding a responsive travel app using Tailwind CSS. Huet also previews Sora's text-to-video generation and OpenAI's unreleased Voice Engine, which can clone a voice from seconds of audio and translate it into multiple languages while preserving the original speaker's tone. He outlines OpenAI's four focus areas: advancing textual intelligence, faster/cheaper models (GPT-4o is twice as fast as GPT-4 Turbo at half the price), model customization via fine-tuning, and enabling multimodal agents. The talk emphasizes that today's models are the dumbest they'll ever be and urges developers to build for a future where AI works across text, voice, and video.

120k players in a week: Lessons from the first viral CLIP app: Joseph Nelson
Nov 22, 2023 · 15:59
Joseph Nelson (Roboflow CEO) details building paint.wtf, a viral AI Pictionary game using OpenAI CLIP that attracted 120,000 players in its first week, peaking at 7 submissions per second. The architecture uses GPT-3 to generate prompts, a browser canvas for drawing, and CLIP to judge similarity between text and image embeddings. Key lessons include CLIP's ability to read text, requiring a moderation hack that penalizes images more similar to handwriting than the prompt; CLIP's conservative similarity scores (ranging 8% to 48%); and the necessity of using CLIP itself to block NSFW content. Nelson also live-codes a minimal MVP in under 50 lines of Python using Roboflow's open-source inference server, demonstrating CLIP embedding and cosine similarity. The talk underscores foundation models' new paradigm of open-set understanding.

See, Hear, Speak, Draw: Logan Kilpatrick & Simón Fishman
Oct 24, 2023 · 18:43
OpenAI's Simón Fishman and Logan Kilpatrick argue that 2024 will be the year of multimodal models, showcasing how text can currently bridge vision, audio, and image generation. They demo a DALL-E vision loop where GPT-4 with vision describes an image, DALL-E 3 generates a synthetic version, then the models compare and iterate to improve accuracy—all in roughly 50 lines of code. Another demo extracts frames from a video, uses GPT-4 with vision to describe them, combines with Whisper's transcript, and generates a blog post with matching DALL-E images, making video content more accessible. They emphasize that while today's multimodal capabilities are separate, future unified models will directly reason across modalities, eliminating the need for manual integration.
Powered by PodHood