The AI Brief
Today's brief:
- OpenAI previews Ultrafast, a Cerebras-powered API tier that runs GPT-5.6 Sol at 750 tokens per second, 14× standard speed, eliminating the frontier intelligence vs. latency tradeoff for real-time enterprise workloads.
- Gemini 3.7 Flash jumped 16 points on the DeepSWE coding benchmark to 65.3% in three weeks while Google cut its price in half, making the workhorse tier, not the delayed Pro flagship, the model operators should actually be evaluating.
- Anthropic makes Claude Code's auto mode the default today for Pro, Max, and Team plans, backed by data showing its classifier catches 89% of dangerous commands vs. 13.6% for human reviewers clicking through prompts.
- Beijing's forced unwind of Meta's $2 billion Manus acquisition completes, with the Chinese agentic startup resuming independent operations and user data from the Meta era set for deletion, Tencent is in talks to take the largest stake.
- OpenAI names Dali Rajic, former Wiz president and COO, as its second Chief Revenue Officer in under a year, adding another senior slot change in the months before its planned IPO.
OpenAI Runs GPT-5.6 Sol at 750 Tokens Per Second With Cerebras-Powered Ultrafast Tier
OpenAI on August 13 shared an early look at Ultrafast, a new service tier running GPT-5.6 Sol up to 14× faster than standard processing; powered by Cerebras, Ultrafast generates up to 750 output tokens per second. It is launching first in the OpenAI API to a select group of customers, with expanded access to more businesses as capacity grows.
Until now, getting real-time speed typically meant choosing a smaller or more specialized model; Ultrafast points to progress in a new direction, more useful work per second. OpenAI cited incident response, customer service and support, financial market analysis, and e-commerce as primary target use cases. In demo scenarios, tasks like data collection, sorting, normalization, and contextualization that previously took one to two hours were compressed to 10 to 15 minutes, and at times operated near real-time.
GPT-5.6 Sol on Ultrafast mode is currently in a limited preview; OpenAI plans to expand access as capacity grows, inviting interested businesses to sign up for notifications. The partnership with Cerebras, a silicon vendor whose wafer-scale chips are purpose-built for inference throughput, is OpenAI's first public use of non-Nvidia hardware for a flagship model tier.
Google Ships Gemini 3.7 Flash With a 16-Point Coding Leap and 50% Price Cut, Still No Word on 3.5 Pro
Google released Gemini 3.7 Flash just three weeks after Gemini 3.6 Flash, introducing major improvements in coding, web development, and document reasoning, and slashed token pricing by 50% through the end of the year. The introductory price is $0.75 per million input tokens and $3.75 per million output tokens , reverting to $1.50/$7.50 in January 2027.
On DeepSWE v1.1, the model scores 65.3%, up from 49.0% for 3.6 Flash; FrontierCode 1.1 Main improves from 34.4% to 43.6%. On the Artificial Analysis Intelligence Index, Gemini 3.7 Flash landed at 56, above Claude Sonnet 5 at 55 and one point below GPT-5.6 Terra and Muse Spark 1.2. Google says Gemini 3.7 Flash outperformed comparable models from Anthropic and OpenAI across nine benchmarks. Caveat Benchmark comparisons are vendor-published; independent verification is pending.
Google released the new Flash model but did not say when it would release Gemini 3.5 Pro, which has been hampered by repeated delays. Reports cite persistent coding and reliability problems, senior researcher departures, and a possible complete retraining from its foundational pre-training phase due to a structural problem. Three Flash-tier upgrades between May and August, against a Pro flagship untouched since February, make clear which tier Google sees as the center of its lineup.
Anthropic Flips Claude Code to Autonomous-by-Default, Citing Classifier Safety Data Over Human Oversight
Starting August 14, new Claude Code sessions on Pro, Max, and Team plans run in auto mode by default, replacing per-step approval prompts with a background classifier that decides in real time which actions are safe to run; Anthropic's testing found the classifier caught 89% of dangerous commands, compared with just 14% under the old manual-approval workflow.
Auto mode remains opt-in for Claude Enterprise, the Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud's Agent Platform, and Microsoft Foundry, with Anthropic planning to make it the default across those services within the coming month. Anthropic has also stopped charging Pro, Max, and Team users for the extra tokens the classifier consumes.
Anthropic engaged Apollo Research for a two-week red-team pilot, then hardened the classifier by giving it more context about the repository environment, visibility, git state, and data-handling rules, after which Apollo re-tested including on a held-out attack set; the reported result was a miss rate falling from 12% to 7%. Teams with allow-rules broad enough to grant arbitrary interpreter execution should note that those rules are set aside while auto mode is active, specifically to prevent commands from skipping classification entirely by matching an overly broad allow-rule.
Disclosure: Claude, which generates this brief, is built by Anthropic.
Beijing Completes Forced Unwind of Meta's $2 Billion Manus Acquisition; Tencent in Talks for Largest Stake
Manus said on August 11 it will "soon resume operating as an independent company," after Chinese regulators in April demanded Meta unwind its $2 billion acquisition; Manus is a developer of general-purpose AI agents founded in China in 2022 before relocating to Singapore. China's National Development and Reform Commission (NDRC) issued its directive in April, instructing the parties to withdraw the transaction.
The NDRC order made clear that offshore incorporation does not shield a deal from Beijing's authority when the underlying technology and talent originated in China, a structure critics had called "Singapore-washing." Co-founders Xiao Hong and Ji Yichao were required to appear before Chinese officials in Beijing in March and have since been prohibited from traveling abroad.
User data generated by certain users on or after December 29, 2025, the acquisition close date, will be deleted later this month as part of the regulatory separation. Tencent has been in talks to become Manus's largest shareholder, Reuters reported in July. Manus had shared tools and data with Meta after striking the deal, but Meta has since erected an internal firewall as the deal unwinds. The unwinding is one of the first forced reversals of a completed AI acquisition under Chinese national-security grounds.
OpenAI Names Second CRO in Under a Year as Pre-IPO Revenue Execution Becomes the Central Test
OpenAI has replaced Chief Revenue Officer Denise Dresser after just nine months on the job, tapping Wiz president and COO Dali Rajic to take on the frontier lab's top sales role. The move comes as part of a broader shake-up in which COO Brad Lightcap and CEO of AGI deployment Fidji Simo have both departed. OpenAI co-founder and president Greg Brockman has taken a larger role in management following Simo's departure and announced Rajic's arrival in a blog post.
Before joining OpenAI, Rajic was president and COO of Wiz, the cybersecurity firm Google acquired for $32 billion in its largest-ever acquisition. Brockman wrote that the next phase of OpenAI's business requires "relentless focus" on proving "measurable business impact" for every dollar customers spend, and that Rajic would turn what the company has learned "into repeatable execution."
OpenAI's products now reach more than one billion weekly active users and more than two million businesses, twice as many as a year ago. OpenAI's naming of its second CRO in less than a year is described by Bloomberg as a sign of efforts to bolster sales growth ahead of a highly anticipated Wall Street debut. Rajic's cybersecurity background at Wiz is notable as OpenAI continues to build out its Daybreak cyber-defense model line, the overlap is either strategic alignment or a coincidence of timing that will clarify quickly.