ND.BUILDS // NOTES
· By Frank Milz
OpenAI and Broadcom unveil Jalapeño for LLM inference
OpenAI's first custom inference chip landed with Broadcom, aiming late-2026 deployment and better performance per watt.
On June 24, OpenAI and Broadcom unveiled Jalapeño, OpenAI's first Intelligence Processor built for large language model inference. OpenAI said the ASIC went from design to tape-out in about nine months, with early tests pointing to strong performance per watt. Broadcom handles silicon implementation and networking, with systems partners in the rack path.
## Why it matters Custom inference silicon is how labs fight GPU scarcity and unit economics. Microsoft, Meta, Amazon, and Google already walked versions of this path. OpenAI joining changes bargaining power across cloud and defense buyers who need reliable inference at scale.
## Caveat Public performance sheets were still thin at announce. Volume capacity is a late-2026 story. Until then, treat Jalapeño as strategy confirmed and silicon still scarce.
