AI Subscriptions Are a Trap

AI subscriptions add up faster than you notice. I'm paying $236/month across seven AI tools — and every company I'm giving money to says they're losing money on me. The math only works because of VC subsidies that won't last. When the bill catches up, you'll wish you'd built your own stack. This playlist tracks what AI actually costs, why it's about to cost more, and what happens when you stop renting and start owning your compute. Topics covered: • What you're really paying every month (the subscription audit) • Why "$20/mo" turns into $200/mo (the enshittification cycle) • Self-hosting the AI tools you actually use • Replacing SaaS with a homelab — what works, what doesn't • When AI is genuinely cheaper than humans, and when it isn't If you've ever stared at a credit card statement wondering when "just one more subscription" became $100+/month, this playlist is for you.

Curated by: Codacus (15 videos)


Currently Playing: 1M Context in 500MB?! DeepSeek V4 + TurboQuant Explained

1 million tokens of context used to require ~4TB of KV cache. Now it fits in ~500MB. That’s not a small improvement — it completely breaks the economics of long-context AI. Two breakthroughs landed almost at the same time: → DeepSeek V4 (April 24, 2026) — open-weights MoE with a hybrid attention stack that reduces KV cache to ~7–10% at 1M context (MIT licensed) → TurboQuant (March 25, 2026) — a KV cache quantization method from Google that achieves ~5× compression with zero calibration (implemented by the open-source community) Stacked together, they shrink memory requirements by orders of magnitude. The result: Long-context AI is no longer a premium feature. Yes — V4 still trails models like Claude Opus 4.6 and Gemini 3.1 Pro in raw quality. Even DeepSeek admits that. But the real moat wasn’t quality. It was infrastructure cost. And that moat is gone. ⏱ Chapters 00:00 — 4TB to 500MB 00:50 — Why long context was expensive 01:55 — DeepSeek V4 (Hybrid Attention breakdown) 03:25 — TurboQuant (PolarQuant + QJL) 04:45 — Why they stack 06:10 — What this does to pricing 06:55 — Quality tradeoffs 07:50 — Hardware trends 08:40 — Wrap 🔗 Sources DeepSeek V4: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro TurboQuant: https://arxiv.org/abs/2504.19874 vLLM blog: https://vllm.ai/blog/deepseek-v4 Simon Willison: https://simonwillison.net/2026/Apr/24/deepseek-v4/ 📂 Research doc https://github.com/codacus/deepseek-v4-research If you’re into local AI, LLM optimization, and breaking hardware limits — subscribe. #DeepSeekV4 #TurboQuant #LLM #LocalAI #KVCache #LongContext #AI


Tracks in this Playlist