Robot-Use Agents: Why General-Purpose Models May Win in Robotics

Categories: VC, Startup, Design

Summary

General-purpose language models are becoming viable for robot control without task-specific fine-tuning, shifting robotics from specialized VLAs to flexible agents that leverage web-scale training data. This represents a fundamental architectural shift similar to chain-of-thought reasoning in LLMs.

Key Takeaways

  1. RT2 breakthrough: Pre-trained language models fine-tuned to output end-effector poses (robot coordinates) outperformed task-specific training by tapping into web image and text understanding.
  2. Data modality matters more than architecture: Instead of building specialized robotics models, transfer high-quality language model training data to robotics—converting out-of-distribution robotics tasks into in-distribution language problems.
  3. Compute allocation enables complexity: Modern robot agents (like coding approaches) allocate more compute for reasoning before action, mirroring LLM chain-of-thought—allowing one model to handle diverse task complexity instead of fine-tuning per task.
  4. Multi-robot coordination now viable: LLM-based agents demonstrated communication and task execution across multiple robots, suggesting general-purpose models scale beyond single-robot control systems.
  5. Evaluation infrastructure emerging: RoboCurve's approach to benchmarking all robot types (humanoids, quadrupeds, arms) with LLMs and classical models reveals VLAs are losing ground as general agents prove more capable.

Related topics

Transcript Excerpt

One of the big surprises the last few years has been the coding agents across different domains. And now frontier researchers are showing that this includes controlling robots. This led MIT Professor Philip Isola to suggest in a recent viral essay that we may be entering the era of robot use agents where general purpose models could make different robots more capable. So today Francois and I invited the founders of Waddle Labs and RoboCurve, two groups of startups that are working at the very frontier of making robots more capable with LLMs. Maybe you guys want to briefly introduce yourselves and just say a little bit about um what each of your companies focuses on. >> I'm Jaime. I'm from Waddle Labs together with Vincent. We work on uh building LLMs that control robots and we do this by d…

More from Y Combinator