The AI Brief
Today's brief:
- Anthropic and the White House remain deadlocked after Monday talks in Washington failed to restore Fable 5 and Mythos 5
- A new Berkeley benchmark finds frontier agents pass just 2.6% of the hardest real-world professional tasks
- On the hardest tier of real work, even the best agents barely register
- The DOJ files to dismiss the NAACP's xAI pollution lawsuit, citing Grok's role in active military operations
- OpenAI's Ona acquisition bets that persistent cloud execution is the missing layer for enterprise agents
Anthropic and the White House are deadlocked over Fable 5, with no resolution in sight
On Friday June 12, Commerce Secretary Howard Lutnick sent Anthropic CEO Dario Amodei a letter directing the company to suspend all access to Claude Fable 5 and Claude Mythos 5 by any foreign national, whether inside or outside the United States, including Anthropic's own foreign-born staff. The directive was issued by the Commerce Department's Bureau of Industry and Security under Lutnick's signature, ordering Anthropic to suspend access to both models for any foreign national, including Anthropic's own foreign-national employees. Unable to selectively comply in real time, Anthropic chose to shut down access entirely rather than risk blocking a wide swath of users.
The Washington Post reported that the Trump administration had weighed export controls on Anthropic weeks before forcing its models offline, after a dispute over the company sharing its technology with a suspected China-linked firm. David Sacks, a Trump adviser, said the government received a warning that Fable 5 could be jailbroken, and that when Anthropic was notified, Amodei said the jailbreak was not a serious risk and refused to fix it. The administration had previously pressed Anthropic to pause the release of the new models, but the company declined, leading to the formal export control letter.
Anthropic sent senior technical staff to Washington and met with administration officials on Monday. Anthropic and the Trump administration remain at an impasse over Fable 5, their most advanced AI model, after high-level discussions in Washington ended without a resolution. Anthropic disputes the government's rationale, arguing that finding a narrow potential jailbreak should not be cause for recalling a commercial model deployed to hundreds of millions of people, and that applying this standard across the industry would essentially halt all new model deployments for all frontier providers. Anthropic has stated it believes the government should have the ability to block unsafe deployments as part of a statutory process that is transparent, fair, clear, and grounded in technical facts, and that this action does not adhere to those principles.
The episode is the latest escalation in a deeper conflict: in February 2026, after failed contract renegotiations, Trump directed federal agencies to cease using Anthropic's AI, and Defense Secretary Hegseth designated Anthropic a "supply chain risk" — the first time that designation, historically reserved for foreign adversaries, was applied to an American company. The underlying dispute was the Pentagon's demand that Anthropic waive restrictions on using Claude for mass domestic surveillance and fully autonomous weapons without human oversight.
Disclosure: Claude, which generates this brief, is built by Anthropic.
A new benchmark grounded in real labor-market work finds frontier agents nowhere near job-ready
ALE is designed to evaluate AI agents on long-horizon, economically valuable, real-world tasks with verifiable outcomes, developed in collaboration with 250+ industry experts covering non-physical industries defined by the O*NET/SOC 2018 federal occupational taxonomy. The benchmark is organized around 55 subfields grouped into 13 industry clusters covering 1,000+ tasks; on the hardest tier, the average full pass rate across mainstream configurations is below 1% in the paper's primary results. The Hugging Face paper page places the hardest-tier aggregate at 2.6%.
On the hardest tier, both Claude Fable 5 and GPT-5.5 score 0%. Frontier agents, including Fable 5, GPT-5.5, and Composer 2.5, complete a meaningful fraction of professional tasks and cluster tightly on aggregate score across easier tiers, suggesting current systems handle routine professional work but fail on the complex, multi-step work that dominates high-value employment. Rather than scripting an agent step by step, ALE hands a frontier agent a real task on a real machine, lets it work to completion, and scores the artifacts against verifiable success criteria; each task is a genuine project that a domain expert has already shipped, converted into a code-graded, fully reproducible test.
The research is led by Dawn Song, professor at UC Berkeley, whose group grounded ALE on "economically valuable work" in the real labor market rather than abstract benchmark design. ALE uses rolling evaluation: every six months a new public subset is published with fresh instances while private tasks rotate in and retired tasks rotate out, to limit benchmark leakage.
ALE's evaluation design exposes why standard benchmarks overstate agent readiness
ALE spans 55 non-physical sub-industries grounded in the O*NET/SOC 2018 federal taxonomy. It targets what the researchers call the Generalist Computer-Use Agent (GCUA): a system given full access to both a graphical interface and a command line. The benchmark does not constrain how an agent solves a task; whatever a human could do on a computer the agent is free to do, and it is judged on the result rather than the method.
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. ALE's authors argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. The three-tier difficulty structure (near-term, full-spectrum, last-exam) is specifically designed so that the hardest tier remains unsaturated as models improve, functioning as a persistent signal of true frontier capability rather than a benchmark that gets solved and retired.
ALE is designed as a living benchmark, with its task pool growing continuously as new workflows and industries are onboarded. The benchmark is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP-relevant impact. The open-source nature of the evaluation harness, combined with code-graded scoring, means results are reproducible without model-judge variance, which is the primary reliability weakness in most existing agent evals.
DOJ files to dismiss NAACP's xAI pollution suit, citing Grok's role in active military operations
The Department of Justice, along with the state of Mississippi, is asking the court to dismiss the lawsuit the NAACP filed against xAI in April, which alleged the company was operating methane gas turbines to power its Colossus 2 data center in South Memphis without the proper permit. The NAACP asked the court for an injunction, citing increased "risks of asthma attacks and heart disease." As of June 15, 26 gas turbines were operating at the facility, emitting 16 tons of hazardous air pollutants and more than 1,000 tons of nitrogen oxides, making xAI's turbines collectively one of the largest, or potentially the largest, industrial source of nitrogen oxides in Shelby County, per the lawsuit.
In its filing, the Justice Department reportedly wrote that stopping xAI from running its turbines "threatens American national, economic, and energy security by seeking to shut off the power supply for artificial-intelligence innovation that supports the Department of War's military operations." It added that only four AI models support mission-critical operations across top-secret classified networks, with Grok being one of them. The Defense Department's chief digital and AI officer submitted a separate filing in support of xAI, detailing how Grok's Gov model supports "vital national security missions."
The NAACP filed an injunction request on May 6, alleging the turbines were operating without air pollution permits or necessary pollution controls. A DOJ deputy assistant attorney general wrote in a court notice that "it is the policy of the United States to sustain and enhance America's global AI dominance." The filing arrived the same day as the DOJ's June 15 deadline to intervene, per court records.
OpenAI acquires Ona to give Codex agents persistent, secure cloud workspaces
OpenAI announced on June 11, 2026 that it will acquire Ona, formerly known as Gitpod, a cloud execution and orchestration startup, bringing Ona's secure, persistent execution and orchestration technology into OpenAI's Codex ecosystem. Ona provides a platform that enables AI agents to run in cloud-based sandboxes that remain online when developers shut down their workstations, meaning agents' work is not interrupted.
OpenAI said Codex now serves more than 5 million weekly users, a roughly 400% increase from earlier this year, and that Ona has helped 2 million developers work in secure cloud environments. The addition of Ona's technology will allow Codex users to delegate work that may take hours or days to the coding agent without being tied to a single device or active session. The acquisition follows OpenAI's May 11 launch of the $4 billion OpenAI Deployment Company and a $150 million Partner Network announced June 14, signaling a consistent push toward full-stack enterprise deployment rather than model access alone.
Bloomberg and CNBC report the deal has not yet closed and financial terms were not disclosed. After closing, the Ona team will join OpenAI and work with the Codex team to advance secure, persistent enterprise execution capabilities.