The AI Brief

Vol. I · No. 94 · Thursday, August 27, 2026

Today's brief:

  • Nvidia reports $96.2 billion in Q2 revenue, up 106% year-over-year, with Data Center alone at $89 billion; the company guides Q3 to $108 billion, and its stock jumped 8% premarket, confirming AI infrastructure demand has not yet peaked.
  • Anthropic's $45 billion commitment to a one-year-old neocloud reveals the real pre-IPO bet: that locking in supply now, even from unproven builders, is less risky than showing up to an S-1 without the compute to back a $190 billion revenue projection.
  • OpenAI's Jalapeño custom inference chip posts 1.5×–1.9× higher throughput per kilowatt and 1.7×–3.6× lower latency than Nvidia's GB200 and GB300 racks in initial benchmarks, the first hard data on whether a frontier lab can out-run GPU infrastructure on inference.
  • Nvidia buying Hugging Face means the company that sells the picks now also owns the general store, giving it leverage to quietly tilt the open-source model ecosystem toward its own hardware.
  • Chinese lab Z.ai formally identifies Ox Alpha, a model that processed 62 trillion tokens anonymously on OpenRouter, as GLM-5.3-Flash, confirming it ran entirely on domestic Chinese chips and ranks 10th globally on the Artificial Analysis Intelligence Index.

Nvidia Posts $96 Billion Quarter, Guides to $108 Billion Next

Why it matters
Jensen Huang's guidance to $108 billion for Q3, 89% year-over-year growth, is the clearest evidence yet that AI infrastructure spend is still accelerating, not plateauing, with Vera Rubin racks now shipping to CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud.
What's at stake
For operators benchmarking AI infrastructure cost trajectories, Huang's explicit framing, "compute is revenue", signals that chip suppliers are positioning the GPU as a revenue-generating asset, not a cost input, which changes how hyperscaler pricing and capacity deals get structured for the next 12 months.
Detail

Nvidia reported revenue for the second quarter ended July 26, 2026, of $96.2 billion, up 18% from the previous quarter and up 106% from a year ago. Data Center second-quarter revenue was $89.0 billion, up 18% from the previous quarter and up 117% from a year ago. GAAP and non-GAAP gross margins were both 75.0%; GAAP earnings per diluted share were $2.46 and non-GAAP were $2.22.

Third-quarter revenue guidance is $108.0 billion, plus or minus 2%. Nvidia is not assuming any Data Center compute revenue from China in its outlook. Huang said AI "reached its inflection point," noting that the number of companies needing large GPU clusters has expanded dramatically: "This time last year, one lab alone was driving the buildout," he said. "Today, we have a golden age of new AI labs and startups, multiple frontier labs scaling in parallel, a thriving open-model ecosystem and physical AI coming online."

Nvidia announced the Vera Rubin platform is ramping into full production with racks running at partners including CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure. Nvidia stock jumped 8% during premarket trading on Thursday following the company's earnings release. Results were reported after the close on August 26. Nvidia's fiscal Q2 2027 ended July 26, 2026.


$45B
Anthropic's six-year compute commitment to British neocloud Nscale

Anthropic Locks In $45 Billion of Vera Rubin Compute in West Virginia

Why it matters
Anthropic is pre-IPO and must demonstrate supply-side capacity to match its $190–200 billion 2028 revenue projection; a $45 billion commitment to a two-year-old neocloud is a bet that Nscale's West Virginia build, funded separately, delivers on schedule, creating execution risk that will land squarely in the S-1.
What's at stake
For most operators, this is context on how frontier labs are pricing long-term capacity. For enterprises currently evaluating multi-year Claude API contracts, the late-2027 online date and Nscale's nascent track record are the relevant supply-chain variables, not the headline number.
Detail

Anthropic has agreed to spend $45 billion to rent AI cloud computing power from Nscale's flagship data center development in West Virginia, the latest move to secure capacity for its expanding business in advance of going public. Nscale will use Nvidia's Vera Rubin chips, which start coming online late next year, to meet Anthropic's needs; the commitment runs six years and represents about 460 megawatts of power.

Nscale was founded in May 2024 by Josh Payne, who previously co-founded Arkon Energy, an infrastructure company tied to crypto mining; the London-based firm raised just over $4.9 billion in a little over a year and reached a $14.6 billion valuation after a Series D round. Nscale is also preparing a US listing as soon as next month. Anthropic is projecting 2028 revenue of roughly $190 billion to $200 billion, per Reuters.

Bloomberg reports Anthropic has also signed other deals for computing power in recent months, including ones amounting to $50 billion with Fluidstack, $10 billion with Volta Infra Holdings, and $45 billion with SpaceX. The Nscale deal, first reported by Bloomberg, was confirmed by Reuters and CNBC citing separate sources familiar with the agreement.

Disclosure: Claude, which generates this brief, is built by Anthropic.


OpenAI's Jalapeño Chip Beats Nvidia Blackwell on Inference Efficiency in First Benchmarks

Why it matters
If Jalapeño's efficiency gains hold in production, where OpenAI has said safety monitoring alone consumes roughly 20% of supervised inference compute, the cost-per-token math shifts without a single new training run, narrowing the gap between OpenAI's cost structure and open-source alternatives at scale.
What's at stake
For most operators, this is a supply-side development that will take 12–18 months to show up in API pricing. For enterprises negotiating volume API contracts today, Jalapeño's pending end-of-2026 initial deployment is a credible basis to expect throughput-linked price declines in 2027, but the chip is OpenAI-captive silicon with no external purchase path.
Decode
Inference = the compute step that runs a trained AI model to generate an output (an answer, a code completion, an image). Distinct from training, which is the one-time process of building the model. For deployed AI products, inference is the ongoing cost. Throughput per kilowatt measures how much useful AI work, tokens generated, a chip produces per unit of electricity, the key efficiency metric for large-scale inference operations.
Detail

In initial benchmarks, Jalapeño delivered 1.5×–1.9× higher throughput per kilowatt and 1.7×–3.6× lower end-to-end latency than Nvidia's GB200 and GB300 rack systems, per Tom's Hardware reporting on the Hot Chips presentation. Richard Ho, OpenAI's head of hardware, said on a press call: "The bottom line is that the results show a very, very significant performance advance over state of the art."

Since announcing Jalapeño, OpenAI's first custom inference chip, the company has been testing the chip and the system built around it; Jalapeño delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two. Each Jalapeño package combines its compute die with six HBM4 stacks, delivering 216 GiB of memory and 15.4 TB/s of bandwidth.

Jalapeño is the first step in a multi-generation compute platform designed for initial deployment by the end of 2026 and expanding in the years ahead, combining OpenAI-designed accelerators with Broadcom silicon implementation, networking, and connectivity technologies; and Celestica's board, rack, and system expertise. Analysts told CNBC that it could put pressure on Nvidia in the fast-growing inference market and reduce OpenAI's reliance on Nvidia for some workloads. Jalapeño is OpenAI-captive silicon, not a product OpenAI sells, with no API, no rental market, no instance type; the chip serves OpenAI's own API traffic.

Caveat Benchmark figures are from OpenAI's own Hot Chips presentation tested on SemiAnalysis' InferenceX benchmark. Independent third-party replication has not been published.


Nvidia Agrees to Acquire Hugging Face for $12.9 Billion

Why it matters
Nvidia already supplies the chips that train and run most models hosted on Hugging Face's platform; owning the distribution layer for three million open-source models and one million datasets would extend Nvidia's position from the silicon layer into the developer ecosystem, the layer where model discovery, optimization for specific hardware, and deployment tooling converge.
What's at stake
For operators whose workflows depend on Hugging Face's neutral, community-governed model hub, the core question is preservation of access terms: Hugging Face's value rests on perceived neutrality across hardware vendors, and a Nvidia-owned hub creates incentives, if not outright requirements, to prefer Nvidia-optimized model formats, tilting the open-source playing field.
Detail

Nvidia has agreed to buy Hugging Face for $12.9 billion, a deal that would put the world's dominant AI chipmaker in control of one of the most important distribution hubs for open AI models; The Information first reported the agreement, citing a person familiar with the deal, and Reuters later confirmed it. The deal comes just two days after reports that Hugging Face was exploring a potential sale at a valuation of $13 billion or more.

The New York-based company hosts more than three million open-source AI models and one million datasets and had engaged a bank to gauge buyer interest. Hugging Face last raised in 2023 at a $4.5 billion post-money valuation in a round led by Salesforce Ventures with participation from Alphabet, GV, IBM Ventures, and others. The company earlier this year turned down a $500 million investment from Nvidia that would have valued it at $7 billion, saying it didn't want a single dominant investor to sway decisions.

The acquisition would give Nvidia something it cannot manufacture in a semiconductor fab: a massive community of AI developers and a repository that has become a default destination for models, datasets, and machine learning tools. The deal has not closed; Nvidia and Hugging Face did not immediately respond to requests for comment per Reuters.


China's Z.ai Reveals Viral "Ox Alpha" as GLM-5.3-Flash, Ran 62 Trillion Tokens on Domestic Chips

Why it matters
For a week, an anonymous model topped global usage charts on OpenRouter without any hardware disclosure; Z.ai's confirmation that it ran 100,000 domestic Chinese chips at the scale of 100 trillion tokens per day is the most concrete public demonstration yet that China's chip-independence push can sustain frontier-class inference traffic on non-Nvidia silicon.
What's at stake
For most operators, this is context on the durability of US export controls. For procurement teams evaluating open-weight models, GLM-5.3-Flash ships under MIT license at $0.15/$0.50 per million tokens with a 1-million-token multimodal context window, it ranks 10th globally on the Artificial Analysis Intelligence Index, ahead of DeepSeek V4 Pro Max, and warrants inclusion in model eval sweeps.
Decode
MoE (Mixture-of-Experts) = a model architecture where only a subset of parameters activates for each token, allowing large total parameter counts with lower per-query compute. GLM-5.3-Flash is a 320B-parameter MoE model with 18B active parameters per forward pass, enabling frontier-scale quality at sub-frontier inference cost.
Detail

China's Zhipu AI launched GLM-5.3-Flash, previously code-named Ox Alpha, saying the system ran entirely on a cluster of 100,000 domestically produced chips during a high-profile stealth trial. Before formal release, the model processed 62 trillion tokens on OpenRouter and agent platform OpenCode; on OpenRouter, the system processed more than 11 trillion tokens in its first three days, making it the platform's biggest launch to date.

GLM-5.3-Flash ranks 10th on the Artificial Analysis Intelligence Index, ahead of DeepSeek V4 Pro Max. Someone anonymously dropped Ox Alpha on OpenRouter and OpenCode on August 20, offering a 1-million-token multimodal context window, a 131,072-token output limit, and a capacity to process 100 trillion tokens per day entirely free of cost for a week. The model is 320B total parameters with 18B active, under MIT license, at $0.15/$0.50 API pricing.

The deployment marks a significant test of China's ability to handle large-scale global inference workloads on home-grown hardware, as Beijing seeks to reduce reliance on advanced processors from Nvidia amid tight export controls. Zhipu's shares closed more than 12% higher in Hong Kong on Thursday. Weights are published on Hugging Face under MIT License. Z.ai is the international brand name for Beijing-based lab Zhipu AI.

Caveat The claim that inference ran exclusively on domestic Chinese chips is from Z.ai's own launch blog. Independent hardware verification has not been published.