Two people holding a plaque with a wafer labeled 'Jalapeño Intelligence Processor'.

OpenAI’s Fiery First Custom Chip Aims to Redefine LLM Inference

OpenAI Unveils Jalapeño, Its First Custom AI Chip Built for Faster LLM Inference

OpenAI has introduced Jalapeño, its first custom-designed AI chip, marking a major step in the company’s push to build more powerful and efficient infrastructure for large language models. The new processor was developed with Broadcom and is aimed specifically at AI inference workloads, the type of computing needed to run products such as ChatGPT, Codex, the OpenAI API, and future agent-based AI tools.

The announcement signals OpenAI’s growing ambition to control more of the technology stack behind its AI services. As demand for generative AI continues to rise, companies building advanced AI systems are looking for ways to reduce latency, improve performance, increase reliability, and scale computing capacity without relying entirely on third-party hardware suppliers.

Jalapeño is designed for the next generation of AI inference

Unlike general-purpose AI accelerators that are adapted to support many different workloads, Jalapeño was designed from the ground up for modern large language model inference. That means the chip is optimized for the real-time processing demands of conversational AI, coding assistants, API-based applications, and future AI agents that may need to reason, plan, and act across multiple steps.

OpenAI says Jalapeño is the first part of a multi-generation compute platform. The goal is not simply to release one chip, but to establish a long-term custom silicon roadmap that can support increasingly advanced AI models over time.

The company describes the chip as a blank-slate design built around the workloads OpenAI already runs every day. By focusing on LLM inference from the beginning, Jalapeño aims to deliver high throughput while also reducing response delays, an important factor for interactive AI products used by millions of people.

Built with Broadcom and supported by a wider hardware ecosystem

Jalapeño was brought to production with Broadcom, which is handling key production responsibilities for the chip. OpenAI says the processor moved from initial design to manufacturing tape-out in just nine months, a fast timeline for custom silicon development.

The platform will also be supported by Celestica, which is expected to help with board design, rack-level system integration, scalable manufacturing, and deployment infrastructure. Together, the partners are working to turn Jalapeño from a single chip into a complete AI computing platform that can be deployed at data center scale.

This approach reflects the broader direction of the AI industry. As AI systems become more complex, performance depends not only on the chip itself but also on memory, networking, rack design, power efficiency, and software integration. OpenAI appears to be building Jalapeño with that full-stack strategy in mind.

Designed for ChatGPT, Codex, the API, and future AI agents

OpenAI says Jalapeño is purpose-built for the workloads that power its core products, including ChatGPT, Codex, the API, and upcoming agentic AI services. Agentic AI refers to systems that can take more autonomous actions, complete multi-step tasks, and interact with tools or digital environments more dynamically than traditional chatbots.

These types of AI experiences require low-latency inference and reliable scaling. A delay of even a few seconds can affect the user experience, especially when AI tools are used for coding, research, productivity, customer support, or business automation.

By building a custom inference chip, OpenAI can optimize hardware around the exact needs of its models and services. This could help improve response speed, reduce operating costs, and make AI products more accessible as usage grows.

Early samples are already running AI workloads

OpenAI has not revealed the full technical specifications of Jalapeño, but the company confirmed that early engineering samples are already running machine learning workloads. One example mentioned is GPT-5.3-Codex-Spark, which is reportedly operating at production target frequency and power.

Images of the chip show multiple high-bandwidth memory areas and central compute silicon, suggesting that memory capacity and bandwidth are important parts of the design. This is expected, since LLM inference relies heavily on fast access to model parameters and efficient data movement.

While exact performance numbers have not been shared, OpenAI’s messaging suggests that Jalapeño is meant to combine the power of leading AI accelerators with latency closer to specialized inference systems. If successful, that could make it especially useful for large-scale interactive AI applications.

Deployment expected by the end of 2026

The first Jalapeño-based platforms are expected to be deployed by the end of 2026, with expansion planned in the following years. OpenAI is positioning the chip as the beginning of a long-term hardware strategy rather than a one-time experiment.

This timeline also reflects the increasing urgency around AI infrastructure. As more users and businesses adopt generative AI, the need for dedicated compute capacity continues to grow. Custom AI chips can help companies reduce bottlenecks, improve efficiency, and better align hardware with their software roadmap.

Why OpenAI is moving into custom AI silicon

OpenAI’s move into custom chip design comes at a time when demand for AI accelerators is extremely high. The AI industry has been heavily dependent on high-performance GPUs and accelerator systems, but supply constraints, cost pressures, and rapid model growth have pushed major AI companies to explore custom hardware.

A custom chip gives OpenAI more control over performance, cost, power efficiency, and long-term supply planning. It also helps diversify the company’s compute infrastructure, reducing reliance on any single hardware provider.

This does not necessarily mean OpenAI will stop using existing AI hardware from other suppliers. Instead, Jalapeño appears to be part of a broader strategy to create a more flexible and resilient compute portfolio. For a company operating some of the world’s most widely used AI services, that flexibility could become increasingly important.

A major step for AI infrastructure

The launch of Jalapeño highlights a major shift in the AI race. Leading AI companies are no longer focused only on building better models; they are also investing deeply in the chips, data centers, and platforms required to run those models at massive scale.

For OpenAI, Jalapeño represents a new phase in its infrastructure strategy. By designing a chip specifically for LLM inference and future agentic AI products, the company is preparing for a future where AI systems are faster, more responsive, and more deeply integrated into everyday digital tools.

If Jalapeño delivers on its goals, it could help OpenAI improve the performance of ChatGPT, Codex, API services, and future AI agents while supporting the growing demand for artificial intelligence worldwide.