The AI Brief

Vol. I · No. 81 · Friday, August 14, 2026

Today's brief:

  • OpenAI previews Ultrafast, a Cerebras-powered API tier that runs GPT-5.6 Sol at 750 tokens per second, 14× standard speed, eliminating the frontier intelligence vs. latency tradeoff for real-time enterprise workloads.
  • Gemini 3.7 Flash jumped 16 points on the DeepSWE coding benchmark to 65.3% in three weeks while Google cut its price in half, making the workhorse tier, not the delayed Pro flagship, the model operators should actually be evaluating.
  • Anthropic makes Claude Code's auto mode the default today for Pro, Max, and Team plans, backed by data showing its classifier catches 89% of dangerous commands vs. 13.6% for human reviewers clicking through prompts.
  • Beijing's forced unwind of Meta's $2 billion Manus acquisition completes, with the Chinese agentic startup resuming independent operations and user data from the Meta era set for deletion, Tencent is in talks to take the largest stake.
  • OpenAI names Dali Rajic, former Wiz president and COO, as its second Chief Revenue Officer in under a year, adding another senior slot change in the months before its planned IPO.

OpenAI Runs GPT-5.6 Sol at 750 Tokens Per Second With Cerebras-Powered Ultrafast Tier

Why it matters
Speed at frontier intelligence has historically required a smaller model, Ultrafast breaks that constraint, putting GPT-5.6 Sol into real-time latency territory where incident response, live trading, and customer-support workflows previously could not use a flagship model.
What's at stake
For API customers in latency-sensitive verticals, Ultrafast changes what GPT-5.6 Sol can bid for; for rivals, including Anthropic and Google, Cerebras partnership gives OpenAI a speed axis none of their current flagship tiers matches at production scale.
Detail

OpenAI on August 13 shared an early look at Ultrafast, a new service tier running GPT-5.6 Sol up to 14× faster than standard processing; powered by Cerebras, Ultrafast generates up to 750 output tokens per second. It is launching first in the OpenAI API to a select group of customers, with expanded access to more businesses as capacity grows.

Until now, getting real-time speed typically meant choosing a smaller or more specialized model; Ultrafast points to progress in a new direction, more useful work per second. OpenAI cited incident response, customer service and support, financial market analysis, and e-commerce as primary target use cases. In demo scenarios, tasks like data collection, sorting, normalization, and contextualization that previously took one to two hours were compressed to 10 to 15 minutes, and at times operated near real-time.

GPT-5.6 Sol on Ultrafast mode is currently in a limited preview; OpenAI plans to expand access as capacity grows, inviting interested businesses to sign up for notifications. The partnership with Cerebras, a silicon vendor whose wafer-scale chips are purpose-built for inference throughput, is OpenAI's first public use of non-Nvidia hardware for a flagship model tier.


65.3%
Gemini 3.7 Flash's score on DeepSWE v1.1, the long-horizon software-engineering benchmark, up from 49.0% for Gemini 3.6 Flash, released just three weeks earlier.

Google Ships Gemini 3.7 Flash With a 16-Point Coding Leap and 50% Price Cut, Still No Word on 3.5 Pro

Why it matters
Google is accelerating its Flash cadence, three releases in three months, while its Pro flagship sits unshipped past 90 days of delay, a pattern that signals a deliberate shift toward the workhorse tier as the de facto competitive front.
What's at stake
For most operators evaluating frontier coding models, Gemini 3.7 Flash's DeepSWE and FrontierCode gains at $0.75 per million input tokens create a meaningful cost-performance argument against GPT-5.6 Terra and Claude Sonnet 5; for Google, each Flash upgrade without a Pro is a data point competitors will use to argue the flagship is structurally delayed, not merely late.
Decode
DeepSWE v1.1 = a benchmark measuring a model's ability to autonomously resolve real software engineering issues, finding bugs, writing fixes, and passing test suites, across long multi-step coding trajectories. Higher scores mean the model can handle more of a developer's pull-request workflow without human intervention.
Detail

Google released Gemini 3.7 Flash just three weeks after Gemini 3.6 Flash, introducing major improvements in coding, web development, and document reasoning, and slashed token pricing by 50% through the end of the year. The introductory price is $0.75 per million input tokens and $3.75 per million output tokens , reverting to $1.50/$7.50 in January 2027.

On DeepSWE v1.1, the model scores 65.3%, up from 49.0% for 3.6 Flash; FrontierCode 1.1 Main improves from 34.4% to 43.6%. On the Artificial Analysis Intelligence Index, Gemini 3.7 Flash landed at 56, above Claude Sonnet 5 at 55 and one point below GPT-5.6 Terra and Muse Spark 1.2. Google says Gemini 3.7 Flash outperformed comparable models from Anthropic and OpenAI across nine benchmarks. Caveat Benchmark comparisons are vendor-published; independent verification is pending.

Google released the new Flash model but did not say when it would release Gemini 3.5 Pro, which has been hampered by repeated delays. Reports cite persistent coding and reliability problems, senior researcher departures, and a possible complete retraining from its foundational pre-training phase due to a structural problem. Three Flash-tier upgrades between May and August, against a Pro flagship untouched since February, make clear which tier Google sees as the center of its lineup.


Anthropic Flips Claude Code to Autonomous-by-Default, Citing Classifier Safety Data Over Human Oversight

Why it matters
Anthropic is making the most consequential default change in Claude Code's history on empirical grounds, a classifier outperforming human reviewers at dangerous-command detection, and the choice reframes the human-in-the-loop assumption that underlies most enterprise AI governance policies.
What's at stake
For most operators, this is context for rethinking agentic governance defaults broadly. For teams running Claude Code on Pro, Max, or Team plans with broad interpreter allow-rules already written, auto mode's interaction with those rules changes behavior starting today, and the Enterprise rollout follows within weeks.
Detail

Starting August 14, new Claude Code sessions on Pro, Max, and Team plans run in auto mode by default, replacing per-step approval prompts with a background classifier that decides in real time which actions are safe to run; Anthropic's testing found the classifier caught 89% of dangerous commands, compared with just 14% under the old manual-approval workflow.

Auto mode remains opt-in for Claude Enterprise, the Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud's Agent Platform, and Microsoft Foundry, with Anthropic planning to make it the default across those services within the coming month. Anthropic has also stopped charging Pro, Max, and Team users for the extra tokens the classifier consumes.

Anthropic engaged Apollo Research for a two-week red-team pilot, then hardened the classifier by giving it more context about the repository environment, visibility, git state, and data-handling rules, after which Apollo re-tested including on a held-out attack set; the reported result was a miss rate falling from 12% to 7%. Teams with allow-rules broad enough to grant arbitrary interpreter execution should note that those rules are set aside while auto mode is active, specifically to prevent commands from skipping classification entirely by matching an overly broad allow-rule.

Disclosure: Claude, which generates this brief, is built by Anthropic.


Beijing Completes Forced Unwind of Meta's $2 Billion Manus Acquisition; Tencent in Talks for Largest Stake

Why it matters
Beijing's successful reversal of a completed, Singapore-domiciled acquisition, on grounds that the underlying technology and talent originated in China, establishes that offshore incorporation does not shield cross-border AI deals from Chinese regulatory jurisdiction, a precedent every investor in China-linked AI startups now has to price.
What's at stake
For operators with data in Manus, the August 23–24 deletion window is an immediate action item; for investors and acquirers considering China-linked AI assets globally, the NDRC's "Singapore-washing" ruling sets a structural constraint that changes how any such deal must be structured from day one.
Detail

Manus said on August 11 it will "soon resume operating as an independent company," after Chinese regulators in April demanded Meta unwind its $2 billion acquisition; Manus is a developer of general-purpose AI agents founded in China in 2022 before relocating to Singapore. China's National Development and Reform Commission (NDRC) issued its directive in April, instructing the parties to withdraw the transaction.

The NDRC order made clear that offshore incorporation does not shield a deal from Beijing's authority when the underlying technology and talent originated in China, a structure critics had called "Singapore-washing." Co-founders Xiao Hong and Ji Yichao were required to appear before Chinese officials in Beijing in March and have since been prohibited from traveling abroad.

User data generated by certain users on or after December 29, 2025, the acquisition close date, will be deleted later this month as part of the regulatory separation. Tencent has been in talks to become Manus's largest shareholder, Reuters reported in July. Manus had shared tools and data with Meta after striking the deal, but Meta has since erected an internal firewall as the deal unwinds. The unwinding is one of the first forced reversals of a completed AI acquisition under Chinese national-security grounds.


OpenAI Names Second CRO in Under a Year as Pre-IPO Revenue Execution Becomes the Central Test

Why it matters
OpenAI's revenue organization has now had two leaders in under twelve months, with COO Brad Lightcap and CEO of AGI deployment Fidji Simo also out this summer, creating a succession gap in the commercial layer at precisely the moment the company needs to convert 1 billion weekly users and 2 million business customers into the durable revenue a public market will underwrite.
What's at stake
For enterprise buyers mid-contract negotiation or renewal, repeated CRO turnover changes who owns the relationship and how reliably commitments made by the outgoing team survive; for potential IPO investors, the rate of senior departures is a diligence variable that will appear prominently in S-1 risk disclosures.
Detail

OpenAI has replaced Chief Revenue Officer Denise Dresser after just nine months on the job, tapping Wiz president and COO Dali Rajic to take on the frontier lab's top sales role. The move comes as part of a broader shake-up in which COO Brad Lightcap and CEO of AGI deployment Fidji Simo have both departed. OpenAI co-founder and president Greg Brockman has taken a larger role in management following Simo's departure and announced Rajic's arrival in a blog post.

Before joining OpenAI, Rajic was president and COO of Wiz, the cybersecurity firm Google acquired for $32 billion in its largest-ever acquisition. Brockman wrote that the next phase of OpenAI's business requires "relentless focus" on proving "measurable business impact" for every dollar customers spend, and that Rajic would turn what the company has learned "into repeatable execution."

OpenAI's products now reach more than one billion weekly active users and more than two million businesses, twice as many as a year ago. OpenAI's naming of its second CRO in less than a year is described by Bloomberg as a sign of efforts to bolster sales growth ahead of a highly anticipated Wall Street debut. Rajic's cybersecurity background at Wiz is notable as OpenAI continues to build out its Daybreak cyber-defense model line, the overlap is either strategic alignment or a coincidence of timing that will clarify quickly.