A product discussed on AI Engineer.

Small Bets, Big Impact Building GenBI at a Fortune 100 – Asaf Bord, Northwestern Mutual
Dec 23, 2025 · 22:50
Asaf Bord, AI Product Lead at Northwestern Mutual, shares how his team built GenBI, an LLM-powered analytics copilot, by flipping the logic from a single big bet to an incremental roadmap of small, fundable projects. Using real, messy data from the 160-year-old company, they deployed a modular architecture with metadata, RAG, SQL, and BI agents, each productizable independently. The RAG agent alone automated 80% of the 20% of BI team capacity spent on finding and sharing reports, saving roughly two full-time employees. Bord explains how a crawl-walk-run release strategy built trust with both users and leadership, starting with BI experts before expanding to business managers, and how each six-week sprint delivered tangible business value—like proving the ROI of enriched metadata through A/B tests against a semantic layer initiative. He also explores the future of SaaS pricing in the GenAI era, questioning whether per-seat models still make sense when individuals become 10x more effective.

From Mixture of Experts to Mixture of Agents with Super Fast Inference - Daniel Kim & Daria Soboleva
Jun 27, 2025 · 53:15
Daria Soboleva and Daniel Kim of Cerebras explain how Mixture of Experts (MoE) architectures enable scaling large language models efficiently by replacing monolithic feedforward networks with specialized experts, a technique used by GPT-4 and Claude. They then introduce Mixture of Agents (MoA), which combines multiple LLMs with custom prompts to outperform frontier models like GPT-4o on complex tasks, reducing a 293-second reasoning problem to 7.4 seconds using Cerebras' ultra-fast inference. The workshop guides participants to build their own MoA system, configure agents for bug fixing and performance optimization on a Python function, and achieve scores up to 120/120. Daniel details Cerebras' wafer-scale chip with 900,000 cores and distributed memory that eliminates memory bandwidth bottlenecks, enabling linear scaling and 15.5x faster inference on Llama 3.3-70B versus GPUs. Daria discusses ongoing research in diffusion models and sparsity, while Daniel notes plans for multimodal APIs and LoRA fine-tuning support.
Powered by PodHood