‘We made a chip and it is fast’: OpenAI’s Jalapeno shows big gains in speed
business-standard.com Aug 26, 2026

‘We made a chip and it is fast’: OpenAI’s Jalapeno shows big gains in speed

AI-summarised brief · reviewed before publication

OpenAI announced significant performance improvements from Jalapeno, its first custom artificial intelligence inference chip. The company stated that Jalapeno delivers higher throughput and lower latency compared to leading commercial systems. Testing against three large language models, including GPT-OSS 120B and Kimi K2.5 1T, revealed substantial gains. Jalapeno achieved 1.5 to 1.9 times more AI work per watt at peak throughput. End-to-end latency was reduced by 1.7 to 3.6 times across the tested models. For highly interactive workloads, performance improvements ranged from 2.1 to 4.1 times higher. The chip operates at a rated 700 watts, with sustained power remaining at or below 550 watts during tests. These results were measured using the InferenceX benchmark. OpenAI emphasized that the architecture provides more intelligence from every watt without sacrificing efficiency. This development marks a strategic shift for OpenAI, moving beyond building models to designing its own underlying infrastructure. The company aims to deploy its own silicon alongside accelerators from Nvidia and other providers. This in-house hardware initiative strengthens OpenAI’s control over its AI infrastructure and operational costs.

💡 Why It Matters

  • · OpenAI’s successful deployment of custom silicon reduces reliance on external hardware vendors like Nvidia, granting the company greater control over its supply chain and operational costs.
  • · By optimizing hardware specifically for its models, OpenAI can achieve superior efficiency and speed, creating a competitive moat that is difficult for rivals using off-the-shelf components to replicate.