LLM Evaluation

Watch and track your favorite playlist.

Curated by: CampusX (12 videos)


Currently Playing: Whats is LLM Benchmarking | Benchmark Saturation vs. Contamination | CampusX

In this session of our LLM Evaluation Masterclass, we decode the mechanics of model-level benchmarking. Using the classic GSM8K mathematical dataset as a practical lens, we walk through how to construct an automated evaluation harness to test raw model capabilities under strictly controlled conditions. This lecture was taken live for insiders, join here: https://youtu.be/j_G30FLmCcw Slides: https://1drv.ms/o/c/85452F67DAA1111C/IgCIRXGQVQm8R7-eF2h5lZ4NAayeX_9R3V4XHDDmsXYMnrk Code: https://colab.research.google.com/drive/1Jg-ujeyTfYn9yNqKSt7uvN0gJOtySiIg?usp=sharing 📱 Grow with us: CampusX' LinkedIn: https://www.linkedin.com/company/campusx-official CampusX on Instagram for daily tips: https://www.instagram.com/campusx.official My LinkedIn: https://www.linkedin.com/in/nitish-singh-03412789 Discord: https://discord.gg/PsWu8R87Z8 E-mail us at support@campusx.in Chapters 00:00 - Recap: Core LLM Capabilities & Shifting to Benchmarks 00:52 - Defining the LLM Benchmark: The Standardized Model Exam 01:45 - The 4 Core Architectural Components of Every Global AI Test 02:30 - Dissecting the GSM8K (Grade School Mathematics) Dataset 04:20 - Live Demo: Sourcing the GSM8K Repository on Hugging Face 05:30 - Component 2: Breaking Down Run Configurations & Guidelines 06:12 - Prompt Construction: Zero-Shot vs. Few-Shot (8-Shot Prompting) 08:24 - Enforcing Logic: Activating Chain of Thought (CoT) Frameworks 09:10 - Parameter Controls: Locking Temperatures & Managing Token Budgets 10:05 - Component 3: Evaluating Strict Pass@1 vs. Lenient Pass@K Scoring 11:18 - Statistical Certainty: Implementing Majority@K (Mode Selection) 12:15 - Tool Integration Constraints: Allowing Web Browsing & Code Interpreters 13:34 - Component 4: Scoring Mechanisms (Regex Parsing vs. Model Graded) 15:53 - Component 5: Data Aggregation Methods (Simple Means vs. Weighted Means) 17:46 - Exploring the Origins of AI Exams: Peer-Reviewed Research Papers 21:15 - Execution: How an Evaluation Harness Acts as the Exam Administrator 26:40 - The Plumbing Problem: Handling Rate Limits, Retries, and Token Batching 29:00 - Live Python Demo: Configuring the EleutherAI LM Evaluation Harness 30:35 - Executing the Code Command: Tracking Latency & Parsing JSON Trace Logs 32:40 - Reviewing the Output: Interpreting a 90% Accuracy Mean on 20 Samples 35:36 - The Flaws of AI Testing: Why You Cannot Blindly Trust Leaderboards 43:10 - Pitfall 1: Benchmark Contamination & Scraped Pre-Training Memorization 45:50 - Pitfall 2: Benchmark Saturation & Statistical Score Clustering 48:10 - Pitfall 3: Configuration Gaming & Concealing Domain-Specific Failures 50:45 - Final Summary: Implementing Your Own Validation Frameworks


Tracks in this Playlist