LLM Evaluation

Watch and track your favorite playlist.

Curated by: CampusX (12 videos)


Currently Playing: How to Use LLM Leaderboards | CampusX

When picking a backend model for your application, it's easy to look at the #1 ranked model on a leaderboard and assume it's the best choice. But public rankings can be heavily inflated due to data contamination, over-optimisation, and gaming of evaluation configurations. In this lesson of our LLM Evaluation Masterclass, we break down how LLM leaderboards actually work, who uses them, and why they should be treated as filtering tools rather than decision-making tools. This lecture was taken live for insiders, join here: https://youtu.be/j_G30FLmCcw 📱 Grow with us: CampusX' LinkedIn: https://www.linkedin.com/company/campusx-official CampusX on Instagram for daily tips: https://www.instagram.com/campusx.official My LinkedIn: https://www.linkedin.com/in/nitish-singh-03412789 Discord: https://discord.gg/PsWu8R87Z8 E-mail us at support@campusx.in Chapters: 00:00 - Introduction: What are LLM Leaderboards? 01:25 - 4 Reasons Why Leaderboards Exist: Trust, Cost, and Saturation Tracking 04:29 - Key Stakeholders: Who Uses Leaderboards and Why? 06:34 - How Frontier Labs Use Stealth Models (e.g., Nano Banana) 08:55 - Type 1: Benchmark-Specific Leaderboards (e.g., Humanity's Last Exam) 10:32 - Type 2: Multi-Benchmark Leaderboards (e.g., LiveBench & Artificial Analysis) 13:07 - Type 3: Human Preference Leaderboards (e.g., LMSYS Chatbot Arena) 16:04 - Type 4: Application-Specific Leaderboards (e.g., Berkeley Function Calling) 17:52 - 7 Critical Pitfalls: Why You Cannot Blindly Trust Leaderboard Scores 21:10 - Goodhart's Law in AI: When Metrics Become Targets 25:14 - Step-by-Step AI Engineer Framework: How to Correctly Read Leaderboards 28:54 - The Golden Rule: Leaderboards are Filtering Tools, Not Decision Tools 29:26 - Looking Ahead: Transitioning to Custom Model Evaluations


Tracks in this Playlist