I hate Opus 5. It’s the best model, anyway.

Categories: AI, Product

Summary

We're hitting an "intelligence overhang" where incremental model improvements no longer translate to real-world leverage—expect the next year to shift focus from raw capability to speed, cost, and open source. Opus 5 excels technically but exhibits surprising neuroticism and human-dependency that reveals fundamental differences in how labs tune frontier models.

Key Takeaways

  1. Intelligence improvements are plateauing for practical use; builders should pivot focus to speed, cost optimization, and open-source models rather than chasing benchmark gains.
  2. Opus 5 demonstrates excessive conservatism and human-reliance in agent behavior—it constantly defers decisions and seeks permission rather than executing autonomously, signaling tuning toward caution over capability.
  3. Model personality analysis (beyond benchmarks) reveals core differences in lab philosophy; comparing how models handle autonomy, risk-taking, and decision-making shows what each lab actually optimizes for.
  4. Frontier model fatigue is real—new models releasing weekly with marginal gains creates decision paralysis; practical evaluation should focus on production use cases rather than benchmark scores.
  5. Agent systems need explicit autonomy tuning; Opus 5 required constant prompting to execute independently, suggesting builders should architect guardrails before delegating decision-making to models.

Related topics

Transcript Excerpt

You guys, I'm tired. What I'm tired of is models coming out every week. New models, new benchmarks, new frontier intelligence, new things to test. It's been a little bit of a run the past month. We've seen Fable come and go and come again. We've seen GPT 5 6. We've seen Sonnet 5. Lots of so many fives recently and just so many models. And I've been lucky. I've been able to test these models, been able to play with them for, you know, sometimes days, sometimes weeks. It just depends on who I'm working with. And it's been really interesting and exciting to have access to all this frontier intelligence. But I think we have an intelligence overhang. I really think that we're [music] running out of, and by we, I mean the average coder, average software engineer, average creator, average builder…

More from How I AI Podcast