OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing

Categories: Startup, VC, AI

Summary

OpenAI's custom Jalapeno inference chip outperforms Nvidia's GB300 in both throughput and latency—a rare combination achieved through full-stack control from models to silicon. The chip reduces data movement through novel HBM4-based architecture, enabling cheaper inference and faster agentic workloads while OpenAI maintains Nvidia partnerships for training.

Key Takeaways

  1. Custom silicon enables full-stack optimization from models through software to hardware that general-purpose chips cannot achieve, allowing OpenAI to make specific trade-offs for inference workloads.
  2. Jalapeno combines high throughput AND ultra-low latency in one device—first in industry. High throughput reduces per-token costs at scale; low latency enables faster agent responses and coding tasks.
  3. Novel architecture reduces data movement between memory and cores through integrated affinity, fundamentally different from GPUs/TPUs. Programmability proven by shipping three models within months of hardware arrival.
  4. Focused on inference, not training, because inference compute is the growth bottleneck as weekly active users scale rapidly. Maintains Nvidia partnership for training while diversifying inference fleet.
  5. Strategic chip selection framework: map different models to optimal silicon based on cost-benefit and latency requirements. Goal is delivering more intelligence at lower inference cost per user.

Related topics

Transcript Excerpt

OpenAI is out with performance results for its first custom inference chip called Jalapeno. The chat GBT maker says in tests, its chip was faster and more efficient than NVIDIA's g b 300 system. Richard Ho, OpenAI's vice president of hardware, is with us here in San Francisco. I think let's start with the basics of the testing and the performance, benchmarks that you've published. What does it tell us about Jalapeno so far? First of all, thanks, Ed, for having me here. The benchmarks that we did were with an open source kind of benchmark called inference x, and we did that because it was a fair kind of neutral benchmarking method. What we're seeing with Jalapeno is that we're able to get both high throughput and low latency with the same device, which is a a kind of a first in the industry…

More from Bloomberg Technology