Watch and track your favorite playlist.
Curated by: CampusX (12 videos)
In this session of our LLM Evaluation Masterclass, we kick off our deep dive into Application Evals by designing a complete, multi-tiered testing suite for a custom CampusX Course Doubt Solver Chatbot. This lecture was taken live for insiders, join here: https://youtu.be/j_G30FLmCcw Slides: https://1drv.ms/o/c/85452F67DAA1111C/IgCIRXGQVQm8R7-eF2h5lZ4NAayeX_9R3V4XHDDmsXYMnrk 📱 Grow with us: CampusX' LinkedIn: https://www.linkedin.com/company/campusx-official CampusX on Instagram for daily tips: https://www.instagram.com/campusx.official My LinkedIn: https://www.linkedin.com/in/nitish-singh-03412789 Discord: https://discord.gg/PsWu8R87Z8 E-mail us at support@campusx.in Chapters: 00:00 - Introduction & Milestone Recap: Shifting to Application Evals 03:55 - Why Focus on RAG & Agentic Application Evaluations? 05:34 - The #1 GenAI Interview Question: "How Do You Evaluate a RAG Chatbot?" 06:44 - Case Study Overview: Building the CampusX Course Doubt Solver 09:20 - Introducing the 3-Tier RAG Evaluation Suite Framework 10:10 - Level 1: Component-Level Evaluation (Retriever vs. Generator in Isolation) 13:46 - Evaluating the Generator: Faithfulness, Relevance, and Citation Accuracy 15:45 - Level 2: Pipeline-Level Evaluation & The RAG Triad 19:00 - Level 3: Application-Level Evaluation (Correctness, Completeness, & Tone) 20:14 - Safety & Operational Audits: PII Leakage, Jailbreak Defense, & Latency Limits 22:20 - Tooling Strategy: Why We Use DeepEval Over Ragas 24:43 - Automated Regression Testing: Establishing Baselines & CI/CD Release Gates 38:33 - MLOps Integration: Continuous Monitoring & The Self-Improving Feedback Loop 43:47 - How to Structure Your Answer in Technical Interviews