Uploads from CampusX

Watch and track your favorite playlist.

Curated by: CampusX (1206 videos)


Currently Playing: Online RAG Evaluation: Monitoring Production LLMs with LangSmith with Code Demo | CampusX

In this final session of our RAG evaluation series, we transition from pre-deployment offline benchmarks to continuous online evaluation and monitoring in production. You will discover which metrics can be tracked in production, why reference-based metrics cannot run online without ground-truth answers, and how to maintain evaluation consistency by running identical DeepEval quality tests in production via background polling. We demonstrate how to set up full-pipeline tracing with the LangSmith SDK, monitor real-time latency and cost telemetry, configure automated alerts, and close the data flywheel by converting failed production traces into versioned offline golden datasets. This lecture was taken live for insiders, join here: https://youtu.be/j_G30FLmCcw 📱 Grow with us: CampusX' LinkedIn: https://www.linkedin.com/company/campusx-official CampusX on Instagram for daily tips: https://www.instagram.com/campusx.official My LinkedIn: https://www.linkedin.com/in/nitish-singh-03412789 Discord: https://discord.gg/PsWu8R87Z8 E-mail us at support@campusx.in Chapters: 00:00 - Introduction & Recap: Transitioning from Offline to Online Evals 02:00 - What are Online Evals? Continuous Production Monitoring 03:30 - The Reference Constraint: Which Metrics Can and Cannot Run Online 07:35 - Selecting the Online Eval Suite: RAG Triad, Toxicity, Latency & Cost 11:45 - Connecting the Application to LangSmith (Environment Setup) 14:30 - Generating the First Trace & Diagnosing Missing Retrieval Steps 19:00 - Explicit Tracing with @traceable Decorators on Pipeline & Retriever 24:30 - Simulating Production Traffic to Collect Live Telemetry 27:30 - Building Custom Monitoring Dashboards (P95 Latency & Cost) 33:15 - Setting Up Production Alerts (Slack Integration & Thresholds) 40:05 - Online Safety Evals: Configuring LangSmith's Built-In Toxicity Evaluator 43:40 - Sampling Rates & Controlling Production Evaluation Costs 50:55 - Online Quality Evals: Why Online and Offline Evaluators Must Match 55:35 - Implementing the Background Polling Service (eval_online.py) 01:03:20 - Live Execution: Scoring Production Traces with DeepEval in Real Time 01:07:15 - Adding RAG Triad Quality Charts to the LangSmith Dashboard 01:08:45 - Closing the Flywheel: Turning Failed Traces into Golden Datasets 01:13:45 - Series Wrap-Up & Looking Ahead to Agent Evaluations


Tracks in this Playlist