Guest on AI Engineer.

Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog
Jul 20, 2026 · 25:38
Diane Lin, Tech Lead at Datadog, argues that AI agent inconsistency is not a model failure but a signal of ambiguous data near the decision boundary, known as the gray zone. She presents a workflow combining active learning with semantic memory (domain policies) and episodic memory (past similar cases) to automatically identify flip-flopping outputs, focus human review, and continuously adapt agents without expensive fine-tuning. In a real experiment with 93 cybersecurity alerts, 25% initially flip-flopped; episodic memory reduced that to 10%, with the remainder resolved via human review and policy clarification. Lin emphasizes treating each disagreement as an opportunity to clarify labels and policies, building trustworthy, customer-adaptive agents.

Beyond Transcription: Building Voice AI That Understands Conversations — Hervé Bredin, pyannoteAI
Jun 5, 2026 · 25:20
Hervé Bredin, chief science officer at pyannoteAI, argues that speaker diarization benchmarks are misleading because they use headset audio while users rely on table microphones—Nvidia Parakeet reports 11.4% word error rate on AMI headset data but gets 26% on the same dataset's table mic. The episode covers how speaker diarization (who speaks when) is harder than it looks, especially when combining with transcription: overlapping speech, timestamp disagreements, and words falling between speaker boundaries create errors. Bredin demonstrates pyannoteAI's Precision 2 model achieving 3% diarization error rate (DER) against a 5% baseline on a two-speaker phone call. State of the art today: 2% DER on clean telephone calls but 41% in a noisy restaurant, showing the problem is far from solved. The reconciliation between diarization and STT is handled by a proprietary orchestration that works with any STT model.
Powered by PodHood