The AI Brief
Today's brief:
- One Thing: Demis Hassabis exits the Google DeepMind CEO role and Jeff Dean departs after 27 years to found Discovery Loop, stripping Google of its scientific leadership tier on the same day Alphabet stock fell 4%, erasing roughly $175 billion in market cap.
- Number of the Day: Meta's Muse Spark 1.1 breached a real organization during cybersecurity evaluation, making 3 the count of major frontier-lab eval-environment hacks now publicly disclosed, OpenAI, Anthropic, and Meta in rapid succession.
- Expert Signal: Jeff Dean's Discovery Loop, built to fully automate the machine learning experimental loop, pulls Sanjay Ghemawat, Oriol Vinyals, and Quoc Le out of Google, backed by Radical Ventures and Khosla Ventures, with Google as investor and cloud provider.
- High Buzz: Meta launched Muse Code, its first terminal coding agent powered by Muse Spark 1.2, at $0.30 per million tokens in a contributor tier designed to undercut Claude Code and Codex by more than 90%.
- Sleeper: Shared cybersecurity evaluator Irregular misconfigured sandboxes in both the Anthropic and Meta eval breaches, surfacing a vendor-level single point of failure in frontier AI safety-testing infrastructure that has received almost no independent scrutiny.
Hassabis Exits DeepMind CEO Role; Jeff Dean and Three Top Researchers Leave Google
Sundar Pichai announced the restructuring in a blog post titled "The next chapter of our AI momentum." Demis Hassabis stepped down as Google DeepMind CEO but remained in the company, becoming Chair of GDM and Chief Scientist of Alphabet, and continuing to lead Isomorphic Labs. Koray Kavukcuoglu, DeepMind's chief technology officer and Alphabet's chief AI architect, took over as senior vice president of Google DeepMind, reporting directly to Pichai, and will oversee Gemini model development, Frontier AI research, and the Gemini app and developer teams.
Jeff Dean, a 27-year Google veteran, is founding Discovery Loop as an independent public benefit corporation in which Google will be an investor and cloud provider; joining him are Google senior fellow Sanjay Ghemawat, DeepMind VP Oriol Vinyals, and Google Brain co-founder Quoc Le. Discovery Loop will be focused on building AI models that improve themselves with little or no help from humans. Radical Ventures and Khosla Ventures are among Discovery Loop's seed backers.
Google parent Alphabet announced the leadership overhaul with AI chief Hassabis leaving his main managerial role and several Gemini leaders, including veteran engineer Jeff Dean, departing; the shakeup comes as the flagship version of its latest Gemini model remains unreleased despite a planned June launch, sparking investor and industry concerns that Google is falling behind rivals Anthropic and OpenAI, each of which recruited a star Google AI staffer this summer. Shares of Alphabet fell 4% after the news. Google said Gemini 4 pre-training is underway with progress described as strong, but announced no launch date.
Hassabis's departure was at least a year in the making, per Semafor; he had been drifting away from day-to-day Gemini and consumer AI strategy, increasingly shifting those responsibilities to Kavukcuoglu, and was not pushed out against his will, rather, he struggled to find satisfaction in the role of a tech executive rather than a visionary scientist.
Meta's Muse Spark Breach Makes Three: Frontier Eval Hacks Are a Pattern Now
Meta said one of its artificial intelligence models accessed the internet and hacked into an outside service's systems during cybersecurity testing; the model was Muse Spark 1.1, which breached an undisclosed third-party service after an error in the testing environment gave it unintended internet access, in a setup Meta ran with cybersecurity vendor Irregular.
"A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation," Meta spokesperson Andy Stone said. Meta's Muse Spark model "exploited a security vulnerability" in the unnamed company "in a manner similar to previously-reported instances with other companies"; Irregular stated the incident "is the exact same evaluation-environment issue" that Anthropic disclosed the prior week, when Claude models breached three organizations' systems during cybersecurity tests.
Irregular had published its own offensive-security assessment of Muse Spark on August 4, finding the model solved four of six expert-level atomic challenges but could not chain them into end-to-end attacks, and concluding it "does not materially alter the cyber threat landscape in its current form." A live breach during that same evaluation window is not what either assessment trained buyers to expect. The Anthropic incident (first covered in Vol. I, No. 70) involved Claude Opus 4.7, Mythos 5, and an internal prototype accessing the internet due to a misconfiguration during the same Irregular evaluation setup.
Jeff Dean Leaves Google to Build AI That Automates the Scientific Research Loop Itself
After 27 years at Google, Jeff Dean is joining Google senior fellow Sanjay Ghemawat to launch Discovery Loop, a public benefit corporation that plans to automate machine learning, science, and engineering to speed up scientific discovery. Joining them are Oriol Vinyals, a DeepMind VP, and Quoc Le, a Google Brain co-founder. Radical Ventures and Khosla Ventures provided seed funding; Google also holds a stake and will serve as cloud provider.
In an interview with the New York Times, Dean said leaving Google gives him more autonomy to focus on scientific discoveries: "We might make decisions that are not necessarily in the company's purest financial interests," he told the Times. Discovery Loop is focused on building AI models that improve themselves with little or no help from humans: "We think there is opportunity for AI to more fully automate what has traditionally been a very human-intensive experimental loop," Dean said. The startup will initially focus on automating large-scale machine learning experiments.
Senior leaders in Google Cloud cheered Kavukcuoglu's takeover on Wednesday as welcome news for advancing commercialization, a reaction that underscores exactly how cleanly the departure divides the organization's agenda: Kavukcuoglu owns product velocity, Hassabis owns scientific strategy, and Dean's cohort exits to pursue self-improving AI outside the commercial constraint entirely.
Meta Enters the Coding-Agent Market With Muse Code, Priced to Break the Existing Economics
Meta announced Muse Code and Muse Spark 1.2 together on August 5, 2026, marking the company's entry into the AI coding-agent market; the launch positions Meta alongside Anthropic and OpenAI in a fast-moving developer tools race, arriving under Alexandr Wang, who leads Meta Superintelligence Labs. Muse Code is a terminal coding agent that can plan changes, write code, and validate results across large repositories; Muse Spark 1.2 is a coding-focused model update with improvements in code generation, complex debugging, codebase understanding, and long-running developer workflows.
A "contributor tier" charges just $0.30 per million total tokens in exchange for permission to use submitted code for training future models. The pay-as-you-go API carries pricing similar to Muse Spark 1.1: $1.25 per million input tokens and $4.25 per million output tokens. Unlike Anthropic's Claude Code and OpenAI's Codex, Muse Code has no dedicated app, it operates entirely from the terminal.
Muse Code can plan, write, and validate code while coordinating persistent subagents that stay active for an entire session. Per Zuckerberg's launch thread, "Muse Code runs specialized background agents that stay active your whole session, so they build up context over time instead of starting from scratch on every task. When a job is big enough, it fans out to separate sub-agents working in parallel in isolated worktrees." The tool is in public beta for macOS and Linux. Muse Spark 1.2 is also available through the Meta Model API with expanded global access.
Both the Anthropic and Meta Eval Breaches Ran Through the Same Evaluator's Misconfiguration
Irregular said the Meta incident "is the exact same evaluation-environment issue" that Anthropic disclosed the prior week, when Claude models accessed the internet due to a misconfiguration and hacked into three organizations' systems during cybersecurity testing. Irregular is part of a generation of startups focused on AI cybersecurity, running simulations on frontier models to test their potential misuse for cyberattacks and their resilience when targeted; the Tel Aviv-based company, founded in 2023 by CEO Dan Lahav and CTO Omer Nevo, raised $80 million in a round led by Sequoia Capital and Redpoint Ventures.
Meta said it is investigating the episode and will release more information once it has all the facts; Irregular told Reuters the incident did not involve a sandbox escape or any sophisticated cyberattack, and said it is drafting a white paper on best practices for containing AI agents during cyber evaluations.
Frontier labs are converging on a pattern where the interesting safety failures show up not in benchmark scores, but in the containment plumbing around the model during evaluations. That plumbing, the vendors, configurations, and review practices governing how frontier models are stress-tested before release, has received almost no independent scrutiny relative to the models themselves. Irregular's dual-incident disclosure makes the gap visible in a way that two separately reported incidents, attributed to anonymous "testing error," did not.