AI subscriptions add up faster than you notice. I'm paying $236/month across seven AI tools — and every company I'm giving money to says they're losing money on me. The math only works because of VC subsidies that won't last. When the bill catches up, you'll wish you'd built your own stack. This playlist tracks what AI actually costs, why it's about to cost more, and what happens when you stop renting and start owning your compute. Topics covered: • What you're really paying every month (the subscription audit) • Why "$20/mo" turns into $200/mo (the enshittification cycle) • Self-hosting the AI tools you actually use • Replacing SaaS with a homelab — what works, what doesn't • When AI is genuinely cheaper than humans, and when it isn't If you've ever stared at a credit card statement wondering when "just one more subscription" became $100+/month, this playlist is for you.
Curated by: Codacus (15 videos)
A single Google paper wiped billions off memory chip stocks overnight. TurboQuant — a KV cache compression algorithm that gives you 6x memory reduction with zero accuracy loss. No training. No fine-tuning. Drop-in replacement. I tested it myself. On a 2021 M1 MacBook Pro with 16GB of RAM. 128K context? Easy. 256K? F16 crashes. TurboQuant runs with 5.7GB free. 768K? Squeezed through with 754MB to spare. 1 million tokens? 363MB free. The machine is on life support — but it's generating. This is the story of how a research paper turned my laptop into something it was never designed to be. Chapters: 0:00 The Headline 1:30 The Hidden Bottleneck (KV Cache) 3:15 How TurboQuant Works 5:00 The Pied Piper Comparison 6:15 Test 1: 128K Context (f16 vs turbo3) 7:30 Test 2: 256K Context (f16 crashes) 8:50 Test 3: 768K Context (pushing the limits) 9:45 Test 4: 1 Million Tokens 11:00 What This Means #turboquant #llm #localai #m1macbook #kvcache #googleresearch #QuantizedInference #llamacpp #contextwindow #aimemory