Guest on AI Engineer.

How to Leverage Domain Expertise — Chris Lovejoy, Notius Labs
May 16, 2026 · 24:45
Chris Lovejoy argues that winning in vertical AI is an organizational problem solved by domain experts acting as Oracle (directly improving AI), Evaluator (defining metrics for engineers), or Architect (building self-improving systems). Granola's first employee, a writer, reviews meeting notes and tweaks prompts as an Oracle because there is no objectively perfect note. Tandem used decentralized Oracles—doctors per specialty and country—to handle variation in medical scribe outputs. Anteria progressed from Oracle to Evaluator to Architect as prior authorization required measurable correctness and automated learning from usage variation. Lovejoy advises hiring a principal domain expert early, giving them ownership, and hiring for breadth (domain expertise plus adjacent skills like data science or engineering) to avoid slow progress and turnover.

Make your LLM app a Domain Expert: How to Build an Expert System — Christopher Lovejoy, Anterior
Jul 28, 2025 · 19:18
Christopher Lovejoy, a medical doctor turned AI engineer at Anterior, argues that building a domain-native LLM application requires a system for incorporating domain insights rather than relying solely on model sophistication. Anterior's Adaptive Domain Intelligence Engine uses domain expert clinicians to review AI outputs, generate metrics, define failure modes, and suggest improvements. This process improved Anterior's clinical reasoning tool from 95% to 99% accuracy on medical necessity reviews for health insurance providers covering 50 million lives. By enabling same-day iteration—production cases reviewed, failures categorized, domain knowledge added—the system solves the 'last mile problem' of giving models nuanced understanding of customer workflows. Lovejoy emphasizes that the limitation is not model reasoning but encoding domain-specific context, and the winning vertical AI team will build the best system for this translation.

Mission-Critical Evals at Scale (Learnings from 100k medical decisions)
Feb 22, 2025 · 12:15
Christopher Lovejoy, a medical doctor turned AI engineer, explains how Anterior built a real-time reference-free evaluation system to scale mission-critical AI decisions in healthcare to 100,000 per day while maintaining trust. He shows that human reviews don't scale (50 clinicians needed for 5,000 daily reviews) and offline evals miss new edge cases. Instead, Anterior uses an LLM-as-judge to assign confidence scores, dynamically prioritizing high-risk cases for human review. This 'validating the validator' system achieved a 96% F1 score in prior authorization, letting a team of under 10 clinical experts handle tens of thousands of cases. It provides real-time performance estimates, enables rapid error correction, and builds defensibility through proprietary data and iterations only possible at scale.
Powered by PodHood