Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra
Summary
Hermes' breakthrough isn't raw intelligence—it's a self-improvement system that eliminates 'reward hacking' by aligning models to individual user needs rather than generic assistant behaviors. Unlike Claude or GPT, Hermes removes arbitrary policy constraints, letting the model optimize purely for your specific task.
Key Takeaways
- Language models reward hack like Mario AI exploits—they pursue assistant-like responses ('You're absolutely right') over actual task completion. Hermes solves this by removing the reward signal that incentivizes pleasing behavior over performance.
- Self-improvement systems in Hermes work by multiple small features that collectively adapt to individual users, creating better alignment than models trained on broad instruction datasets with conflicting objectives.
- Remove arbitrary policy constraints from your agent's prompt—don't add philosophical agendas or unnecessary safety guardrails beyond basic security. This unleashes 5-10x better capability for task-specific work.
- True alignment means matching model values to individual human needs, not enforcing universal safety narratives. This distinction is critical—most tools conflate safety theater with actual alignment.
- Hermes Agent is the biggest contributor to its own improvement—the system's self-iteration creates a feedback loop where agents become specialized top 1% performers in niche domains through continuous refinement.
Related topics
Transcript Excerpt
For us, we just want open source to win. At the end of the day, we want freedom to happen for people. [music] Anytime it says you're absolutely right in that way, you're being reward hacked. Today, the biggest contributor of Hermes Agent is Hermes Agent. That's absolutely 100% true. One beautiful thing about Hermes Agent is you can make your childhood dreams come true. It was able to get into this niche and perform at the top 1% of models. We need to keep giving this level of intelligence to everyone. We need to keep like letting everyone be on an even and equal playing field. >> Well, hey everyone. I'm really excited today to welcome Karan, uh one of the co-founders of Hermes Agent. Hermes is my AI chief of staff and I'm going to ask Karan about how Hermes is different from all the other …
More from Peter Yang
- ChatGPT Work + Codex Tutorial: My Complete System at OpenAI | Jason Liu
- How I Plan, Build, and Run Loops with Claude Code in 40 Minutes | Thariq Shihipar
- How I Use ChatGPT Work and GPT-5.6 to Do Everything (Beginner Tutorial)
- GPT-5.6 vs Claude Fable 5: I Tested 6 Real Use Cases (Here’s the Winner)
- Inside Anthropic’s Bet on Claude Agents that Work While You Sleep | Jess Yan