The AI Brief

Vol. I · No. 50 · Tuesday, July 14, 2026

Today's brief:

  • TSMC's N3 node and CoWoS packaging are fully booked through December, which means 2026 AI infrastructure capacity is already decided, and any operator still weighing whether to lock in GPU reservations has effectively run out of runway.
  • Global mutual funds park just 1.2% of assets in Chinese tech while China captures 16% of global AI revenue, and Goldman Sachs now says that gap is the trade.
  • Chinese models cost 10–25% of US frontier pricing not because of subsidies or loss-leader strategy, but because MoE architectures activate only 2–10% of parameters per query, making the price gap structural and permanent for cost-sensitive agentic workloads.
  • Google scrapped and rebuilt Gemini 3.5 Pro after it failed at recursive tool-calling, the exact capability the model was designed to lead on, making the unconfirmed July 17 launch a test of whether Google has actually solved agentic coding or just rescheduled the problem.
  • China's new companion AI rules drew a hard line between productivity agents (exempt) and persistent emotional agents (regulated), and ByteDance and Alibaba both chose full shutdown over retrofitting anti-addiction requirements into products whose core value depends on persistent memory.

TSMC Posts All-Time Revenue Record as AI Chip Supply Remains Locked Through Year-End

Why it matters
TSMC's N3 node, the process used by virtually every advanced AI GPU and ASIC shipping in 2026, is fully committed through December, meaning today's record revenue is not a growth ceiling to debate but a hard supply floor for all AI compute plans made after last year.
What's at stake
For most operators, this is context, not a decision. For those evaluating whether to lock in reserved GPU capacity or defer infrastructure buildout, the signal is clear: available capacity will not materially increase before 2027, and Thursday's full earnings call from TSMC CEO C.C. Wei is the next moment to watch for any revised outlook.
Detail

TSMC reported Q2 2026 consolidated revenue of approximately NT$1.27 trillion, roughly $39.6 billion, a 36% year-over-year increase and a record for the company's nearly four-decade history. Revenue in the three months to June 30 rose 36% year-on-year to T$1.270 trillion ($39.63 billion). June alone produced NT$442.68 billion, a jump of nearly 68% year-on-year, while first-half 2026 revenue rose nearly 36% to T$2.40 trillion.

The June month-over-month gain is the more telling figure. TSMC's Q2 revenue beat the top end of its projected range of $40.2 billion, and June revenue has typically fallen month-on-month over the past four years, making this increase especially notable. SemiAnalysis analyst Sravan Kundojjala was direct: "The demand supply situation in AI is still quite tight and TSMC is sold out on N3, which is targeted by all leading AI GPU and CPUs this year." SemiAnalysis estimates TSMC is on track to exceed $40 billion in AI-related chip revenue for the full year, roughly 25% of projected total revenue.

The bottleneck extends beyond N3 silicon. Before any AI accelerator can be deployed in a data center, the logic die must be integrated with High Bandwidth Memory stacks in a process called CoWoS (Chip-on-Wafer-on-Substrate), TSMC's proprietary 2.5D advanced packaging architecture. Total CoWoS demand is expected to reach approximately one million wafers in 2026, nearly three times the roughly 370,000-wafer level recorded across all of 2024. TSMC has been expanding CoWoS capacity at roughly 80% per year, but the supply-demand gap that ran as wide as 20% earlier this year is only expected to narrow to approximately 10% by year-end, according to TrendForce. TSMC commands a 73% share of the global pure-foundry market as of Q1 2026 and is set to report its full second-quarter earnings on Thursday, July 16.

The company had delayed the release of its June revenue figures to Monday from Friday due to disruptions caused by Typhoon Bavi. A single weather event in Hsinchu postponed a routine market-moving disclosure and illustrated, without ambiguity, a structural condition that no financial model can fully price: nearly every leading AI chip made anywhere in the world passes through one company, in one country.


1.2%
Global mutual fund allocation to Chinese AI tech, Goldman Sachs finds the figure dramatically misaligned with China's 16% share of global AI revenue, and formally recommends buying the value chain.

Goldman Sachs Calls Chinese AI Underweighted, Initiates Coverage on Zhipu With DeepSeek and ByteDance as Co-Favorites

Why it matters
Goldman's formal coverage initiation gives institutional allocators a named framework to reprice Chinese AI exposure, the same week that Chinese-origin models held above 46% of US enterprise token volume, turning a procurement trend into an investable thesis.
What's at stake
For most operators, this is context on the capital flows shaping which labs have the runway to compete at frontier. For investors with exposure to AI equities, the allocation gap Goldman names, China at 1.2% of tech portfolios versus 16% of AI revenue, is the number to stress-test.
Decode
MoE (Mixture of Experts) = an architecture where a model routes each input through a small subset of specialized sub-networks rather than activating all parameters. This produces near-frontier capability at a fraction of the compute cost, which is why Chinese labs running MoE can price aggressively without losing margin.
Detail

Goldman Sachs has three preferred Chinese AI models, only one of which is publicly traded: the investment bank initiated coverage on Hong Kong-listed Zhipu (also known as Knowledge Atlas) with a price target of HK$1,880. Its two other preferred Chinese AI model companies are DeepSeek and ByteDance, both privately held. Zhipu's GLM and DeepSeek's models generally ranked better than those from Alibaba, Tencent, and Minimax, especially in time to market and Arena score.

Goldman's core argument: since the end of 2022, the global market capitalization of AI-related public companies has increased by $34 trillion, while China's AI sector currently has a market cap of only about $4 trillion, contributing 10% of global AI market capitalization and 16% of global AI revenue, yet global mutual funds allocate just 1.2% of their holdings to Chinese technology.

Goldman dropped a comprehensive report dissecting the competitive positioning of China's AI large language model industry, led by analyst Ronald Keung and titled "Who Will Be the Long-Term Winner in China's AI Large Model Industry?," which introduces a three-dimensional framework evaluating companies across pricing power, cost advantages, and financial strength. Goldman Sachs forecasts that Chinese AI model API and subscription revenue will grow from an estimated 35 billion yuan in 2026 to 879 billion yuan in 2030, corresponding to daily token consumption increasing from 35 trillion to 460 trillion, an approximately 25-fold increase.

Sources: CNBC: Goldman Sachs Picks Its Favorite Chinese AI Models (primary) · CryptoBriefing: Goldman Sachs Releases Competitive Framework for Chinese AI Models
CaveatGoldman's revenue projections are proprietary research estimates, not independently verified market data.

Chinese Models Are Not Cheap Clones, They Price at $1/M Tokens Via Architecture, Not Subsidies

Why it matters
Enterprise procurement teams pricing AI workloads against US frontier models face a structural cost gap, not a temporary one: Goldman's framework shows Chinese labs achieve near-parity on coding and reasoning benchmarks by routing queries through MoE sub-networks at 2–10% of the active parameter count, producing genuinely lower inference cost rather than margin sacrifice or government subsidy.
What's at stake
For most operators, this reframes how to evaluate the 60–90% price gaps: they are not loss-leader pricing that will eventually normalize, but an architectural efficiency advantage. For those building cost-sensitive agentic pipelines at scale, the question shifts from "when will Chinese model prices rise?" to "which tasks justify US frontier premium?"
Detail

Chinese high-end models are running at roughly $1 per million tokens, compared to $4–8 for US equivalents, 10–25% of what American companies charge. At the low end, Chinese models go as cheap as $0.06–0.20 per million tokens. Chinese AI companies achieve this through smaller, more efficient architectures, often ranging from 2–10% of the parameter sizes of US models, and leveraging MoE architectures that route queries to specialized sub-networks rather than activating the entire model.

This is driven by architectural innovations and higher parameter efficiency, allowing Chinese models to achieve comparable capabilities at 2%–10% the parameter size and significantly lower cost. Goldman's framework found that Zhipu and DeepSeek currently dominate in foundational text models, while ByteDance has pulled ahead as the frontrunner in multimodal and video generation capabilities.

The enterprise adoption consequence is already visible: Goldman's report notes that the global enterprise AI usage paradigm is undergoing a fundamental shift from "token maximization" to "ROI priority." The former prevailed from late 2025 to early 2026, where enterprises equated high token consumption with organizational productivity; the latter focuses more on clear task boundaries, daily active agent count, backend process automation, and actual output. Chinese high-end model pricing sits structurally inside that ROI calculus in ways US frontier pricing does not. At the channel level, Alphabet's Gemini Enterprise Agent Platform and Amazon's AWS Bedrock already offer hosting services for Chinese AI models like DeepSeek, MiniMax, Moonshot, GLM, and Qwen , meaning the distribution barrier that once protected US model pricing has largely dissolved.

Sources: CNBC: Goldman Sachs Picks Its Favorite Chinese AI Models (primary) · CryptoBriefing: Goldman Sachs Competitive Framework for Chinese AI
CaveatPer-token pricing figures are Goldman research estimates; individual model pricing varies by provider and tier and changes frequently.

Update Gemini 3.5 Pro Is Three Days Away From Its Unconfirmed Target, Built From Scratch, Still Not Official

Why it matters
Google scrapped its original Gemini 3.5 Pro architecture entirely after internal tests revealed the model could not maintain structural consistency in recursive tool-calling, the defining requirement for agentic coding, making this not a routine delay but a signal about where the frontier's hardest unsolved problem actually sits.
What's at stake
For most operators, this is a watchlist item until Google publishes a model card, pricing, and API documentation. For those currently evaluating frontier model stacks, three major releases converge in ten days: a Gemini 3.5 Pro GA (July 17, unconfirmed), DeepSeek V4 stable release (July 24, confirmed), and ongoing GPT-5.6 rollout, making this the window to run comparative workload tests before committing to a provider.
Detail

Google DeepMind is targeting July 17 for the general availability of Gemini 3.5 Pro, but every specific claim circulating about the launch, including the date itself, the 2-million-token context window, and the benchmark numbers, comes from third-party reporting and unnamed internal sources, not from an official Google announcement. As of July 13, no model card, no pricing page, and no gemini-3.5-pro listing appear in the public Gemini API documentation. First covered in Vol. I, No. 49.

According to reporting from HackerNoon, the scrapped version of Gemini 3.5 Pro showed two specific failure modes: it could not maintain structural consistency when generating complex, multi-layered SVG scene layouts, and it broke down under complex, recursive tool-calling environments, the multi-step chains where an agent calls one tool, uses the result to call another, and so on across dozens of sequential decisions. Recursive tool-call stability is the defining requirement for an agentic coding model, which is the use case Google has staked the entire 3.5 generation on.

Three of the most-watched model events of 2026 converged on the same week, which means a Gemini 3.5 Pro slip from July 17 would land with greater competitive consequence than the June slip did: this time, the field around Google already has newer flagships in production. Leaked pricing places Gemini 3.5 Pro at $60 per million output tokens, putting it in the premium tier alongside Fable 5. DeepSeek V4-Pro at $0.87 per million output is roughly 69 times cheaper. Both figures are unconfirmed by either company. DeepSeek's legacy deepseek-chat and deepseek-reasoner aliases stop working on July 24, 2026, at 15:59 UTC with no announced extension; the required migration is to deepseek-v4-pro or deepseek-v4-flash.

Sources: TechTimes: Gemini 3.5 Pro Targets July 17 After Full Rebuild (primary) · HackerNoon: Google Delays Gemini 3.5 Pro to July 17 · Startup Fortune: Google Scraps Base Model
CaveatJuly 17 launch date, 2M context window, and all pricing figures are from third-party reporting; Google has not issued an official launch announcement as of publication.

Update China's Companion AI Law Takes Effect Today, Doubao and Qwen Shut Down Persistent Agents Rather Than Rebuild

Why it matters
Beijing's Interim Measures do not ban AI agents, they ban persistent emotional bonding, drawing a bright regulatory line between productivity agents (exempt) and companion agents (regulated), with compliance requirements so architecturally demanding that both ByteDance and Alibaba chose full shutdown over retrofit.
What's at stake
For most operators, this is context on how the first national regulatory framework targeting AI emotional interaction is being interpreted in practice: when compliance requires anti-addiction friction that is structurally incompatible with persistent-memory design, vendors pull products rather than rebuild them. Western operators building companion or persistent-memory products should note that China has now created a compliance template that others may reference, and that Character.AI and Replika face analogous pressure in US courts.
Decode
Persistent-memory agent = an AI system that retains information across sessions to maintain a consistent relationship with the user over time. The Interim Measures' mandatory anti-addiction interruptions and instant-exit requirements conflict structurally with this design: an agent cannot simultaneously maintain emotional continuity and interject friction.
Detail

ByteDance's Doubao and Alibaba's Qwen moved to disable customised agent features as China's new rules on humanlike AI interaction services took effect, with Doubao informing users that its agent feature would go offline on July 15 due to "product function adjustments." Qwen also issued a notice that its "humanlike interactive agents and user-created agent functions" would be disabled, with broader "Qwen agent functions and services" going offline by July 15. First covered in Vol. I, No. 49.

The regulation is the Interim Measures for the Administration of AI Anthropomorphic Interactive Services, co-issued on April 10, 2026, by the Cyberspace Administration of China and four partner agencies: the National Development and Reform Commission, the Ministry of Industry and Information Technology, the Ministry of Public Security, and the State Administration for Market Regulation. The measures require companion services to run anti-addiction systems, issue mandatory usage notifications and offer instant-exit mechanisms, alongside real-time detection of unhealthy dependence, demands that sit awkwardly with agents built to remember a user, stay consistent across sessions and keep an ongoing relationship going. Rather than retrofit the feature, ByteDance chose to shut it down.

Services that launch anthropomorphic functions or cross thresholds of one million registered users or 100,000 monthly actives must run security assessments covering eight areas, from training-data handling to minor protection, and file the reports with provincial regulators. The rules do not apply to workplace productivity agents. ByteDance is directing Doubao users to Maoxiang, a separate standalone app where they can create agents again; Alibaba has announced no equivalent migration path for Qwen. Doubao says data will be processed according to its privacy policy and will no longer be accessible or recoverable inside the app after October 15, 2026.