A product discussed on AI Engineer.

Teaching Gemini to Speak YouTube: Adapting LLMs for Video Recommendations to 2B+DAU - Devansh Tandon
Jul 16, 2025 · 22:51
Devansh Tandon, a Product Manager at Google leading YouTube's discovery system, details how YouTube adapted Gemini LLMs to power its recommendation engine for billions of daily active users. The team built SemanticID, a tokenization system that compresses video features into semantically meaningful tokens, creating a new language for YouTube content. They then continued pre-training Gemini on sequences of user watches to make the model bilingual in English and this video language. For generative retrieval, they prompt the adapted model with user demographics and watch history to output video recommendations as SemanticIDs, achieving 95%+ cost savings to serve at scale. Challenges include serving billions of users with low latency and handling video freshness—Taylor Swift's new music video must be recommendable within minutes. Tandon argues LLM-led recommendations are a bigger consumer application than search and hints at future interactive, steerable recommendations and even personalized content creation.

Keynote: Why people think "agent" is a buzzword but it isn't
Feb 22, 2025 · 28:07
Chip Huyen argues that agents are not a buzzword but a practical yet hard technology, facing three core challenges: the curve of complexity, tool use translation, and context management. Even simple queries require multiple steps, and models' success rates drop rapidly past 5 steps—newer reasoning models like DeepSeek R1 are pushing the boundary, but most still fail after 10. Tool use requires translating ambiguous natural language to precise API calls, worsened by poor documentation; she advises narrow functions and asking for clarifications. Context is another bottleneck: agents must juggle instructions, tool docs, and outputs, often exceeding a model's efficient context (many hallucinate beyond 30K tokens), necessitating external memory like RAG. Her benchmark shows planning-specialized models struggle with long context and vice versa, and she recommends breaking tasks into subtasks and using test-time compute scaling.

How to build the world's fastest voice bot: Kwindla Hultman Kramer
Feb 10, 2025 · 20:38
Kwindla Hultman Kramer, CEO of Daily, argues that colocating speech-to-text, LLM inference, and text-to-speech in a single compute container is the most effective way to achieve sub-500 millisecond voice-to-voice latency for conversational AI. He details how architectural flexibility and low-latency media transport are critical, citing measured bottlenecks like 30–40ms from macOS mic processing and typical voice-to-voice latencies of 600–700ms. To hit faster response times, his team uses Deepgram’s on-premises STT and TTS models via Docker and Llama 3 8B for LLM inference, achieving 500–700ms in an open-source demo. The talk introduces PipeCat, a vendor-neutral open-source framework for real-time multimodal AI that orchestrates components like transcription, endpointing, interruption handling, and text-to-speech. Kramer emphasizes that while frontier multimodal models are coming, orchestration layers remain essential for building production-grade voice bots, and shares that a recent latency demo gained over 175,000 views on Twitter.
Powered by PodHood