open-rag-eval: RAG Evaluation without "golden" answers — Ofer Mendelevitch, Vectara
Jun 3, 2025 · 5:03
Ofer Mendelevitch from Vectara presents Open-RAG-Eval, an open-source framework that enables RAG evaluation without requiring golden answers or golden chunks, solving a major scalability problem. Backed by research with the University of Waterloo's Jimmy Lin Lab, it uses UMBRELA for retrieval scoring on a 0–3 scale that correlates well with human judgment, and AutoNuggetizer for generation via nugget creation, Vital/OK ratings, and an LLM judge analyzing the top 20 nuggets for support. Additional metrics include citation faithfulness and Vectara's HHEM hallucination detection model. Connectors are available for Vectara, LangChain, and LlamaIndex, with results viewable through an intuitive UI at OpenEvaluation.ai.