The AI Brief
Today's brief:
- TSMC's N3 node and CoWoS packaging are fully booked through December, which means 2026 AI infrastructure capacity is already decided, and any operator still weighing whether to lock in GPU reservations has effectively run out of runway.
- Global mutual funds park just 1.2% of assets in Chinese tech while China captures 16% of global AI revenue, and Goldman Sachs now says that gap is the trade.
- Chinese models cost 10–25% of US frontier pricing not because of subsidies or loss-leader strategy, but because MoE architectures activate only 2–10% of parameters per query, making the price gap structural and permanent for cost-sensitive agentic workloads.
- Google scrapped and rebuilt Gemini 3.5 Pro after it failed at recursive tool-calling, the exact capability the model was designed to lead on, making the unconfirmed July 17 launch a test of whether Google has actually solved agentic coding or just rescheduled the problem.
- China's new companion AI rules drew a hard line between productivity agents (exempt) and persistent emotional agents (regulated), and ByteDance and Alibaba both chose full shutdown over retrofitting anti-addiction requirements into products whose core value depends on persistent memory.
TSMC Posts All-Time Revenue Record as AI Chip Supply Remains Locked Through Year-End
TSMC reported Q2 2026 consolidated revenue of approximately NT$1.27 trillion, roughly $39.6 billion, a 36% year-over-year increase and a record for the company's nearly four-decade history. Revenue in the three months to June 30 rose 36% year-on-year to T$1.270 trillion ($39.63 billion). June alone produced NT$442.68 billion, a jump of nearly 68% year-on-year, while first-half 2026 revenue rose nearly 36% to T$2.40 trillion.
The June month-over-month gain is the more telling figure. TSMC's Q2 revenue beat the top end of its projected range of $40.2 billion, and June revenue has typically fallen month-on-month over the past four years, making this increase especially notable. SemiAnalysis analyst Sravan Kundojjala was direct: "The demand supply situation in AI is still quite tight and TSMC is sold out on N3, which is targeted by all leading AI GPU and CPUs this year." SemiAnalysis estimates TSMC is on track to exceed $40 billion in AI-related chip revenue for the full year, roughly 25% of projected total revenue.
The bottleneck extends beyond N3 silicon. Before any AI accelerator can be deployed in a data center, the logic die must be integrated with High Bandwidth Memory stacks in a process called CoWoS (Chip-on-Wafer-on-Substrate), TSMC's proprietary 2.5D advanced packaging architecture. Total CoWoS demand is expected to reach approximately one million wafers in 2026, nearly three times the roughly 370,000-wafer level recorded across all of 2024. TSMC has been expanding CoWoS capacity at roughly 80% per year, but the supply-demand gap that ran as wide as 20% earlier this year is only expected to narrow to approximately 10% by year-end, according to TrendForce. TSMC commands a 73% share of the global pure-foundry market as of Q1 2026 and is set to report its full second-quarter earnings on Thursday, July 16.
The company had delayed the release of its June revenue figures to Monday from Friday due to disruptions caused by Typhoon Bavi. A single weather event in Hsinchu postponed a routine market-moving disclosure and illustrated, without ambiguity, a structural condition that no financial model can fully price: nearly every leading AI chip made anywhere in the world passes through one company, in one country.
Goldman Sachs Calls Chinese AI Underweighted, Initiates Coverage on Zhipu With DeepSeek and ByteDance as Co-Favorites
Goldman Sachs has three preferred Chinese AI models, only one of which is publicly traded: the investment bank initiated coverage on Hong Kong-listed Zhipu (also known as Knowledge Atlas) with a price target of HK$1,880. Its two other preferred Chinese AI model companies are DeepSeek and ByteDance, both privately held. Zhipu's GLM and DeepSeek's models generally ranked better than those from Alibaba, Tencent, and Minimax, especially in time to market and Arena score.
Goldman's core argument: since the end of 2022, the global market capitalization of AI-related public companies has increased by $34 trillion, while China's AI sector currently has a market cap of only about $4 trillion, contributing 10% of global AI market capitalization and 16% of global AI revenue, yet global mutual funds allocate just 1.2% of their holdings to Chinese technology.
Goldman dropped a comprehensive report dissecting the competitive positioning of China's AI large language model industry, led by analyst Ronald Keung and titled "Who Will Be the Long-Term Winner in China's AI Large Model Industry?," which introduces a three-dimensional framework evaluating companies across pricing power, cost advantages, and financial strength. Goldman Sachs forecasts that Chinese AI model API and subscription revenue will grow from an estimated 35 billion yuan in 2026 to 879 billion yuan in 2030, corresponding to daily token consumption increasing from 35 trillion to 460 trillion, an approximately 25-fold increase.
CaveatGoldman's revenue projections are proprietary research estimates, not independently verified market data.
Chinese Models Are Not Cheap Clones, They Price at $1/M Tokens Via Architecture, Not Subsidies
Chinese high-end models are running at roughly $1 per million tokens, compared to $4–8 for US equivalents, 10–25% of what American companies charge. At the low end, Chinese models go as cheap as $0.06–0.20 per million tokens. Chinese AI companies achieve this through smaller, more efficient architectures, often ranging from 2–10% of the parameter sizes of US models, and leveraging MoE architectures that route queries to specialized sub-networks rather than activating the entire model.
This is driven by architectural innovations and higher parameter efficiency, allowing Chinese models to achieve comparable capabilities at 2%–10% the parameter size and significantly lower cost. Goldman's framework found that Zhipu and DeepSeek currently dominate in foundational text models, while ByteDance has pulled ahead as the frontrunner in multimodal and video generation capabilities.
The enterprise adoption consequence is already visible: Goldman's report notes that the global enterprise AI usage paradigm is undergoing a fundamental shift from "token maximization" to "ROI priority." The former prevailed from late 2025 to early 2026, where enterprises equated high token consumption with organizational productivity; the latter focuses more on clear task boundaries, daily active agent count, backend process automation, and actual output. Chinese high-end model pricing sits structurally inside that ROI calculus in ways US frontier pricing does not. At the channel level, Alphabet's Gemini Enterprise Agent Platform and Amazon's AWS Bedrock already offer hosting services for Chinese AI models like DeepSeek, MiniMax, Moonshot, GLM, and Qwen , meaning the distribution barrier that once protected US model pricing has largely dissolved.
CaveatPer-token pricing figures are Goldman research estimates; individual model pricing varies by provider and tier and changes frequently.
Update Gemini 3.5 Pro Is Three Days Away From Its Unconfirmed Target, Built From Scratch, Still Not Official
Google DeepMind is targeting July 17 for the general availability of Gemini 3.5 Pro, but every specific claim circulating about the launch, including the date itself, the 2-million-token context window, and the benchmark numbers, comes from third-party reporting and unnamed internal sources, not from an official Google announcement. As of July 13, no model card, no pricing page, and no gemini-3.5-pro listing appear in the public Gemini API documentation. First covered in Vol. I, No. 49.
According to reporting from HackerNoon, the scrapped version of Gemini 3.5 Pro showed two specific failure modes: it could not maintain structural consistency when generating complex, multi-layered SVG scene layouts, and it broke down under complex, recursive tool-calling environments, the multi-step chains where an agent calls one tool, uses the result to call another, and so on across dozens of sequential decisions. Recursive tool-call stability is the defining requirement for an agentic coding model, which is the use case Google has staked the entire 3.5 generation on.
Three of the most-watched model events of 2026 converged on the same week, which means a Gemini 3.5 Pro slip from July 17 would land with greater competitive consequence than the June slip did: this time, the field around Google already has newer flagships in production. Leaked pricing places Gemini 3.5 Pro at $60 per million output tokens, putting it in the premium tier alongside Fable 5. DeepSeek V4-Pro at $0.87 per million output is roughly 69 times cheaper. Both figures are unconfirmed by either company. DeepSeek's legacy deepseek-chat and deepseek-reasoner aliases stop working on July 24, 2026, at 15:59 UTC with no announced extension; the required migration is to deepseek-v4-pro or deepseek-v4-flash.
CaveatJuly 17 launch date, 2M context window, and all pricing figures are from third-party reporting; Google has not issued an official launch announcement as of publication.
Update China's Companion AI Law Takes Effect Today, Doubao and Qwen Shut Down Persistent Agents Rather Than Rebuild
ByteDance's Doubao and Alibaba's Qwen moved to disable customised agent features as China's new rules on humanlike AI interaction services took effect, with Doubao informing users that its agent feature would go offline on July 15 due to "product function adjustments." Qwen also issued a notice that its "humanlike interactive agents and user-created agent functions" would be disabled, with broader "Qwen agent functions and services" going offline by July 15. First covered in Vol. I, No. 49.
The regulation is the Interim Measures for the Administration of AI Anthropomorphic Interactive Services, co-issued on April 10, 2026, by the Cyberspace Administration of China and four partner agencies: the National Development and Reform Commission, the Ministry of Industry and Information Technology, the Ministry of Public Security, and the State Administration for Market Regulation. The measures require companion services to run anti-addiction systems, issue mandatory usage notifications and offer instant-exit mechanisms, alongside real-time detection of unhealthy dependence, demands that sit awkwardly with agents built to remember a user, stay consistent across sessions and keep an ongoing relationship going. Rather than retrofit the feature, ByteDance chose to shut it down.
Services that launch anthropomorphic functions or cross thresholds of one million registered users or 100,000 monthly actives must run security assessments covering eight areas, from training-data handling to minor protection, and file the reports with provincial regulators. The rules do not apply to workplace productivity agents. ByteDance is directing Doubao users to Maoxiang, a separate standalone app where they can create agents again; Alibaba has announced no equivalent migration path for Qwen. Doubao says data will be processed according to its privacy policy and will no longer be accessible or recoverable inside the app after October 15, 2026.