OpenAI Jalapeno chip puts ChatGPT inference costs under fresh pressure
The OpenAI Jalapeno chip, officially called Jalapeño, puts a sharper number on a problem AI companies have avoided spelling out: every ChatGPT answer costs money. OpenAI and Broadcom unveiled the inference chip on June 24, saying it was designed for large language model inference, the process used when a model answers a prompt or runs a coding task.
According to OpenAI, Jalapeño is its first “Intelligence Processor” and was co-developed with Broadcom and Celestica for ChatGPT, Codex, API products and future agent systems. Reuters reported that OpenAI plans to deploy the chip by the end of 2026, with server systems built by Celestica and manufacturing handled through Taiwan’s TSMC.
The OpenAI Jalapeno chip is not aimed at training the next model from scratch. It is built for inference, where usage volume matters as much as raw model size. That helps explain why a year-old joke about polite prompts still matters. In April 2025, Vice reported that OpenAI CEO Sam Altman said users saying “please” and “thank you” to ChatGPT had cost the company “tens of millions of dollars.” The claim came from an X reply, not a detailed financial filing, but it pointed to the same pressure: more tokens mean more compute.
OpenAI said early lab samples were running machine-learning workloads at target production frequency and power, including GPT-5.3-Codex-Spark. It also said the chip was taped out in nine months with help from OpenAI models. A Broadcom release said in October 2025 that the companies planned a 10-gigawatt deployment of OpenAI-designed AI accelerators, starting in the second half of 2026 and finishing by the end of 2029.
The positive case for the OpenAI Jalapeno chip is control. If OpenAI can tune hardware around its own models, memory movement, networking and serving patterns, it may reduce latency and improve cost per query. That would matter for ChatGPT, Codex and enterprise tools where users expect faster answers without unpredictable usage bills.
The critical case is that custom silicon does not erase infrastructure risk. Reuters reported that Broadcom CEO Hock Tan said Jalapeño matches Nvidia Blackwell chips and Google’s tensor processing units, but OpenAI said final performance is still being measured. AI News reported that OpenAI’s infrastructure costs remain central to its financial outlook, while The AI Decode has covered how OpenAI losses have already raised questions about the cost of scaling ChatGPT.
The OpenAI Jalapeno chip therefore looks less like a side project and more like a cost test for the AI business model. The next question is whether the OpenAI Jalapeno chip can turn custom hardware into lower inference costs before user demand, power needs and capital spending rise again.
