Watch and track your favorite playlist.
Curated by: CampusX (12 videos)
In this session of our LLM Evaluation Masterclass, we bridge the gap between staging and production. You will learn how to design a continuous, self-improving evaluation loop that monitors live production traffic using Online Evaluation frameworks. This lecture was taken live for insiders, join here: https://youtu.be/j_G30FLmCcw Slides: https://1drv.ms/o/c/85452F67DAA1111C/IgCIRXGQVQm8R7-eF2h5lZ4NAayeX_9R3V4XHDDmsXYMnrk 📱 Grow with us: CampusX' LinkedIn: https://www.linkedin.com/company/campusx-official CampusX on Instagram for daily tips: https://www.instagram.com/campusx.official My LinkedIn: https://www.linkedin.com/in/nitish-singh-03412789 Discord: https://discord.gg/PsWu8R87Z8 E-mail us at support@campusx.in Chapters: 00:00 - Course Recap: Eval Pipelines, Failure Points, and Evaluation Methods 02:42 - Introducing Today's Agenda: Offline Evals Vs Online Evals 03:27 - Understanding Offline Evaluation: Pre-Deployment Staging 05:53 - Three Massive Benefits of Offline Evals: Pre-Release Gating & Version Comparison 10:05 - Mitigating Code Regression: Making Sure Prompt Updates Don't Break Existing Logic 14:15 - Moving to Production: The 3 Operational Risks of Launching Live AI Systems 14:55 - Risk 1: Unanticipated User Inputs, Rants, & Adversarial Prompt Injections 16:41 - Risk 2: Emergent Systematic Failures, Concurrent Scaling, and Latency Spikes 18:23 - Risk 3: The Technical Reality of Data Drift & Obsolete Golden Datasets 21:54 - Defining Online Evaluation: Monitoring Live Production Traffic Without an Answer Key 23:39 - Side-by-Side Comparison Matrix: Offline Testing vs. Online Auditing 26:49 - Core Concepts: Distinguishing Correctness from System Normalcy 28:09 - The UPSC Grader Example: Analyzing Baseline Score Distribution Waves 34:53 - Handling Live Metrics: Using Thumbs Up/Down and Escalation Signals to Spot Failures 41:12 - Step 1 of the Online Pipeline: Non-Blocking Durable Logging 44:15 - Technical Engineering: Privacy Safeguards and PII Masking Processes 44:42 - Live Walkthrough: Navigating Tracing, Latency, and Cost Charts inside LangSmith 47:47 - Step 2 of the Online Pipeline: Stratified Sampling to Minimize API Evaluation Cost 01:01:40 - Running an LLM as a Judge on Production Traces 01:08:40 - LangSmith Platform Demo: Configuring Real-Time Safety & Toxicity Evaluators 01:14:35 - The Self-Improving Loop of Feeding Production Bugs Back into Offline Data 01:17:15 - Interactive Student Q&A & Lookahead to Model-Level Benchmarks