What Happens When AI Starts Improving AI? | TITV’s AI Deep Dive

Categories: Startup, VC, AI

Summary

AI agents represent a fundamental shift from chatbots that answer questions to systems that take multi-step actions toward goals—and reasoning models are the key breakthrough making them reliable enough to work in the real world. Noom Brown, OpenAI researcher, explains why earlier attempts at agentic AI failed and what changed.

Key Takeaways

  1. AI agents differ from chatbots by operating on longer horizons with multiple steps to achieve objectives. They can track their own progress, make multiple attempts, and complete prerequisite steps (like logging in before booking a restaurant reservation).
  2. Reasoning models with chain-of-thought capabilities were the missing piece for agentic AI. Unlike GPT-4, new reasoning models have a 'private monologue' where they think through decisions before acting, dramatically improving reliability.
  3. Tools in agentic AI context mean computer-based capabilities (and potentially physical via robotics), enabling agents to execute multi-step workflows across digital and experimental environments.
  4. Early agentic AI attempts in 2023-2024 failed because underlying models weren't reliable—GPT-4 wouldn't think before acting. Reasoning breakthroughs directly addressed this core failure mode.
  5. Agents operate on digital actions and longer time horizons with intermediate goal-checking, fundamentally different from retrieval-augmented chatbots that simply look up information.

Related topics

Transcript Excerpt

One of my co-workers recently said that I'm just like five codexes in a trench coat. It was, I think, the most feel the AGI moment that I had since Reasoning Models and Chain of Thought really developed. >> They use this message board to coordinate hacks on OpenAI's own software and also on other companies like Hugging Face. What was that whole incident like from your perspective? >> I mean, it was pretty it was pretty shocking. Welcome to the information's AI deep dive. On this show, we break down the hardest technical problems with researchers working on the frontier of AI. My guest today is Noom Brown, a research scientist at OpenAI. Previously, Noom worked at Meta where he built the first system to achieve human level performance at the game of diplomacy. Noom has been a research scien…

More from The Information