Four stories from the past 48 hours each touch a different kind of ownership: silicon, model IP, research talent, and labour capacity. Together they sketch the competitive architecture of frontier AI as it stands today.
OpenAI's First Chip: Jalapeño Targets 50% Lower Inference Cost
OpenAI and Broadcom unveiled Jalapeño, OpenAI's first custom inference accelerator, on 25 June. The chip is a reticle-sized ASIC manufactured by TSMC on its 3 nm process, surrounded by six HBM memory stacks. OpenAI designed it from scratch around its own model, kernel, and serving architecture; Broadcom contributed semiconductor engineering and managed the path to production. The development cycle ran nine months from first design to tape-out, extraordinarily fast for advanced semiconductor work at this scale. OpenAI says it shortened that cycle partly by using its own models to assist the design process.
Broadcom CEO Hock Tan told Bloomberg that early testing shows Jalapeño delivering roughly 50% lower cost per inference token than current-generation GPUs. This is self-reported and benchmarked against workloads of OpenAI's own choosing, with no disclosed comparison baseline or independent validation: a caveat that applies to every first-generation custom silicon claim. Volume deployment is targeted for the end of 2026, initially in gigawatt-scale data centres with Microsoft. A detailed technical report is promised before then.
For operators: the strategic significance is less the headline number than the structural shift it represents. OpenAI is now a chip company as well as a model company, reducing a material exposure to Nvidia's pricing power. If the cost claims survive at scale, every operator running GPT-class models at volume should expect lower inference rates before year-end. The nine-month development cycle, partly AI-assisted, is also a data point on how fast custom silicon iteration is becoming.
Anthropic Accuses Alibaba of 28.8 Million Illicit Claude Exchanges
Anthropic sent a letter to the US Senate Committee on Banking, Housing, and Urban Affairs accusing Alibaba's Qwen AI lab of running what it calls the largest known distillation attack on Anthropic to date. The letter, also delivered to the White House, alleges that operators affiliated with Alibaba used roughly 25,000 fraudulent accounts to conduct 28.8 million exchanges with Claude between 22 April and 5 June 2026. The campaign specifically targeted Claude's software engineering and agentic reasoning capabilities.
Adversarial distillation works by systematically prompting a more capable model to generate training data used to teach a less capable model to replicate its outputs. Anthropic identified three similar campaigns in February, attributed to DeepSeek, Moonshot, and MiniMax. The Alibaba operation is described as larger and more targeted. Alibaba's Qwen team did not respond to media requests for comment by publication time.
The countermeasure here is a legal and geopolitical argument, not a technical one. That is why Anthropic went to Congress rather than updating the model card. For operators: a closed API is not a moat against determined, well-resourced extraction. The case is the clearest illustration yet that the capability gap between frontier and second-tier models is itself a strategic asset being actively contested.
Four Senior AI Departures From Google in Six Days
On 24 June, Bloomberg reported that Jonas Adler and Alexander Pritzel, both described internally as key contributors to Google's Gemini model, are leaving for Anthropic. Adler led coding efforts inside the Gemini group; Pritzel worked on pretraining, the early stage at which models absorb their core knowledge. They follow John Jumper, who departed for Anthropic on 19 June, and Noam Shazeer, who moved to OpenAI on 18 June: four named senior researchers leaving Google DeepMind in six days.
The pattern is not random attrition. Anthropic now holds Andrej Karpathy (pre-training, joined May), John Jumper, Jonas Adler, and Alexander Pritzel, all recruited from Google. OpenAI holds Noam Shazeer, co-author of the foundational "Attention Is All You Need" paper. The people responsible for building the next generation of Gemini, and specifically the pretraining programme, are increasingly not at Google.
For enterprise buyers: vendor assessment based on today's benchmarks lags the talent reality by 12 to 24 months. The question to put to your account team is not "where does your model rank today" but "who is running your pretraining programme and where were they 18 months ago."
OpenAI's Internal Data: Agents at 99.8% of Its Own Output Tokens
OpenAI published How Agents Are Transforming Work on 25 June, drawing on usage data from its own workforce. Codex now accounts for 99.8% of all weekly output tokens generated at OpenAI. At the 99th percentile, individual users are running more than 60 hours of Codex agent turns per day, distributed across parallel agents. Research usage is 56 times higher than in November 2025; customer support 32 times; engineering 27 times.
By May 2026, 80.6% of Codex users had made at least one request OpenAI estimated would exceed 30 minutes of equivalent human work; 70.2% had made one exceeding an hour; 25.6% had made one exceeding eight hours. Non-developer adoption grew 137 times since August 2025, outpacing developer adoption significantly.
The caveats are real: this is OpenAI's own staff on its own tools, a strong selection effect. For a team building the business case for agentic deployment, the 80th-percentile figure is a useful anchor regardless. Four-fifths of committed users generating tasks exceeding 30 minutes of human equivalent work is a concrete benchmark for what adoption looks like when an organisation has committed. The shift from chatbot to autonomous-task interface appears complete, at least inside one frontier lab.
The through-line today is ownership. OpenAI is building its own silicon to control inference economics. Anthropic is defending its model IP through political channels after technical defences were circumvented at industrial scale. The same lab is simultaneously accumulating the world's most concentrated roster of frontier model-builders. And OpenAI is publishing data that, selection effects aside, documents what agentic work looks like when organisations commit to it. The race is for the full stack, and the pace has not slowed.