Realtime AI News
OpenAI's Jalapeño inference ASIC stays in-house, with broader rollout left open
OpenAI's custom AI inference ASIC, Jalapeño, is being used internally, but the company has left the door open to a broader rollout, according to Tom's Hardware. OpenAI said it will have its 'hands full' with Jalapeño for 'a good long time', which helps explain why no external offering has been announced yet.
OpenAI has built a custom AI inference ASIC called Jalapeño, and for now it is reserved for the company's own use, according to a report from Tom's Hardware.
The same report notes that OpenAI has left the door open to a broader rollout. The company said it will have its 'hands full' with Jalapeño for 'a good long time', a remark that both explains the decision to keep the chip in-house for now and leaves room for a larger deployment later.
The details are thin so far. The report identifies Jalapeño as an inference part serving OpenAI's internal workloads, but it does not disclose a foundry partner, tape-out timing, production volumes, performance figures, or any schedule.
The reason this matters is that inference economics, not training, increasingly decide whether large-scale AI products make money. As agents, long-context requests and real-time multimodal calls multiply the volume of inference, the purchase, power and cooling costs of general-purpose GPUs become the heaviest fixed expense, and silicon tailored to a specific model's compute profile can deliver better efficiency on those workloads.
It is also a direction the industry has been moving in for years: companies with both frontier models and large-scale workloads increasingly shift part of their inference onto chips they design themselves, reducing reliance on any single general-purpose compute supplier.
What to watch next: whether Jalapeño stays an internal experiment or moves into production clusters, whether OpenAI publishes more architectural and performance detail, and whether the talk of being busy reflects a demand pipeline rather than a capacity constraint.
Why it matters
If the chip runs at scale inside OpenAI, the immediate effect is lower unit inference cost and less reliance on general-purpose GPUs. For chip and cloud suppliers, a rising share of in-house silicon among top customers would gradually reshape long-term purchasing, though today's public information is too thin to size that shift.
Nearby Updates
All09/28, 23:37
Google rolls out Gemini 3.8 Flash alongside a new cybersecurity-focused model
Google is rolling out Gemini 3.8 Flash and has introduced a separate model aimed at cybersecurity, according to ALM Corp. The paired releases continue Google's pattern of shipping a fast, cost-efficient Flash tier alongside specialised models for vertical use cases.
09/28, 23:30
Anthropic, Gamma, and Clay share what happens when enterprises actually deploy AI at TechCrunch Disrupt 2026
Anthropic, Gamma, and Clay share what happens when enterprises actually deploy AI at TechCrunch Disrupt 2026. Anthropic, Clay, and Gamma on what it takes for an AI product to go beyond the demo at the AI Stage at TechCrunchDisrupt 2026. Register to join and get 50% off a second pass.
09/28, 22:58
Insygna Joins Open Secure AI Alliance Alongside Nvidia, Microsoft and Okta to Push AI Agent Accountability
HRTech Series reports that Insygna has joined the Open Secure AI Alliance, whose participants include Nvidia, Microsoft and Okta, to advance accountability for AI agents. The move signals that compute, cloud and identity vendors are converging on the question of who answers for an autonomous agent's actions.
09/28, 22:05
Modulate raises $25M for its voice models and analysis suite
Modulate, a Boston-based voice intelligence startup, has raised $25 million in a round led by Future Ventures, with Hyperplane and Lakestar participating. The company runs more than 100 small models for transcription, emotional analysis, deepfake and AI music detection, and compliance review of voice agents in regulated industries.