Videos about using Hermes Agent
Curated by: Tonbi's AI Garage (35 videos)
Hermes Agent has 8 hidden models running in the background — and tweaking one config can cut your auxiliary spend by 85%. Sign up for my FREE weekly newsletter, where I spill my unfiltered thoughts on the latest AI news, cool research, and projects I'm building: https://www.onchainaigarage.com/ 🐦 Follow Tonbi on X for real-time AI x blockchain updates! https://x.com/tonbistudio Most Hermes users never touch the auxiliary model config, but there are 8 background tasks running behind your main chat — compression, flush memories, web extract, vision, session search, skills hub, MCP, and approval — each of which can be independently pointed at a different model. By default they inherit your main model or fall back through an auto chain, which means you might be quietly burning frontier-model tokens on small tasks. In this video I walk through exactly where these models live in config.yaml, which tasks cost the most, and run a live head-to-head on the biggest cost driver (compression) between Claude Opus 4.6 and Kimi K2 through OpenRouter. ✅ The 8 auxiliary tasks explained — which ones cost the most (compression, flush memories, web extract, vision) and which models are best suited to each. ✅ Full config.yaml walkthrough — where to set provider/model per task, the top-level vs auxiliary compression gotcha, and how to wire up local models for zero cost. ✅ Live cost comparison: Opus 4.6 ($0.13) vs Kimi K2 ($0.02) on the same 50K context compression pass — 85% savings that compound across 10-20 compressions per day. 💻 Tonbi's GitHub: https://github.com/tonbistudio 🌐 Portfolio: https://www.tonbistudio.com Resources: 🔗 Hermes Agent (Nous Research): https://github.com/NousResearch/hermes-agent 🔗 Kimi K2 (Moonshot AI): https://openrouter.ai/moonshotai/kimi-k2 🔗 Ollama: https://ollama.com/ 🔗 LM Studio: https://lmstudio.ai/ Timestamps: 0:00 - Intro: The 8 hidden models you didn't know about 0:38 - How auxiliary tasks inherit your main model 1:47 - The 8 tasks and where each one fires 3:02 - Where the spend actually lives (top 4 cost drivers) 4:04 - Finding and editing config.yaml 5:27 - Choosing the right model for each task 7:42 - Pointing auxiliary tasks at a local model 10:34 - Live demo: Opus 4.6 compression cost baseline 13:04 - Switching compression to Kimi K2 via OpenRouter 15:41 - Compounding savings across a full workday Coming Next: More Hermes Agent deep dives, plus the full Master Class series starting in May with multiple episodes per week! 👀 Have you tweaked your Hermes auxiliary config? What models are you using for compression and vision? Drop your setup in the comments! If this saved you some money, please like, subscribe, and hit the bell for more Hermes Agent content! 🦐✨ #HermesAgent #NousResearch #AIAgent #CostOptimization #OpenRouter #KimiK2 #LocalLLM #AITools #LLM #ClaudeOpus #ConfigYaml #AIWorkflow