Watch and track your favorite playlist.
Curated by: CampusX (1206 videos)
In this hands-on session of our LLM Evaluation Masterclass, we begin constructing our complete RAG Eval Suite. To build a production-grade system, testing the final output isn't enough. You must independently audit the Retriever (to measure if it fetches the correct contexts) and the Generator (to measure if it creates accurate, grounded responses without hallucinations). Using our CampusX Course Doubt Solver project, we set up a modular codebase and implement component-level evaluation pipelines using DeepEval. This lecture was taken live for insiders, join here: https://youtu.be/j_G30FLmCcw Code: https://github.com/campusx-official/rag-eval-deepeval 📱 Grow with us: CampusX' LinkedIn: https://www.linkedin.com/company/campusx-official CampusX on Instagram for daily tips: https://www.instagram.com/campusx.official My LinkedIn: https://www.linkedin.com/in/nitish-singh-03412789 Discord: https://discord.gg/PsWu8R87Z8 E-mail us at support@campusx.in Chapters: 00:00 - Introduction & Recap: Building the RAG Eval Suite 01:43 - Agenda: Component-Level Evaluation (Retriever & Generator) 02:44 - Setting Up the Project Architecture & Folder Structure (/src, /evals, /goldens) 05:25 - Ingesting Transcripts & Configuring the UV Virtual Environment 08:45 - Building Component 1: The Retriever (retriever.py) 11:51 - Loading Data, Preprocessing Transcripts, and Chunking Parameters 17:35 - Executing Vector Database Storage (ChromaDB) & Basic Similarity Queries 21:01 - Deep Dive: The 2 Retriever Failure Modes (Misses vs. Noise) 24:50 - Recall vs. Precision in Information Retrieval 28:51 - Why Recall & Precision Are Reference-Based Evaluations 31:21 - The Flaw of Document-ID Golden Datasets (The Chunk Tweak Problem) 36:23 - The Solution: Structuring Golden Datasets with Questions & Ideal Answers 46:17 - How LLM-as-a-Judge Calculates Contextual Recall (Extracting Claims) 57:23 - How LLM-as-a-Judge Calculates Contextual Precision (Rank-Aware Scoring) 01:08:47 - Sourcing Golden Datasets: Hand-Authored vs. LLM-Assisted 01:12:30 - Demo: Generating Golden Test Cases with DeepEval's Synthesizer 01:21:40 - DeepEval Fundamentals: LLMTestCase, Metrics, and evaluate() 01:27:04 - Writing the Retriever Evaluation Script (eval_retriever.py) 01:33:40 - Running Baseline Evaluation Trial 1 (750 Chunk Size / 100 Overlap) 01:35:36 - Optimization Trial 2: Increasing Chunk Size (1000 Chunk / 150 Overlap) 01:38:15 - Optimization Trial 3: Adding a Sentence Transformer Cross-Encoder Re-ranker 01:42:25 - Optimization Trial 4: Upgrading the Embedding Model to Text-Embedding-3-Large 01:46:15 - Benchmarking Final Retriever Performance & Class Summary