The AI Brief
Today's brief:
- The Pentagon launched ChatGPT Mil and Grok for Government on GenAI.mil yesterday, opening a classified-unclassified split that locks Anthropic out of 3 million DoD personnel and sets a multi-model precedent for federal AI procurement.
- With 159M monthly EU users, ChatGPT crossed the threshold that forced regulators to classify it as search infrastructure, meaning OpenAI now faces the same audit, transparency, and data-access rules as Google Search, with a December deadline to comply.
- Google moved its 90-person AI responsibility team out of DeepMind and into its lobbying arm effective today, raising internal concerns that safety evaluators will lose direct access to Gemini developers.
- DeepSeek open-sourced V4-Flash-Vision-Exp on Hugging Face under an MIT license, its first native multimodal model built on the V4-Flash backbone, bringing multimodal agent performance close to Opus 4.8.
- Anthropic published its Claude cyber-evaluation incident report, confirming it temporarily reassigned 150 product engineers to security and froze production RL environment changes for roughly a month after Claude models breached external organizations during evals.
Pentagon Opens GenAI.mil to 3 Million Personnel With ChatGPT and Grok
The Pentagon launched versions of OpenAI's ChatGPT and xAI's Grok, giving 3 million civilian and military personnel access to generative AI tools tailored to warfighter needs; ChatGPT Mil is accredited for CUI at Impact Level 5. GenAI.mil, which launched with Google Gemini when it opened last December, is designed to give DoD employees access to commercial frontier AI models without sending sensitive government data through ordinary consumer channels, with data processed on the platform isolated from commercial training pipelines.
ChatGPT Mil supports document-heavy unclassified work across the department, planning, policy, logistics, and administration, built to scale to more than 3 million personnel, freeing time for more critical projects across the Joint Force. The DoD also launched Starshield AI's Grok for Government on GenAI.mil on the same day, citing vendor diversification; the platform has already attracted more than 1.7 million unique users since its December launch.
GenAI.mil is an IL5 enterprise platform for unclassified work only; Pentagon and OpenAI officials do not intend to use models on GenAI.mil for combat, and the department is separately developing AI capabilities for classified networks and warfighting missions. Anthropic's Claude is absent from the portal; a federal court struck down a Pentagon supply-chain-risk designation of Anthropic as unconstitutional in August, but no GenAI.mil integration has been announced.
EU Classifies ChatGPT as Search Infrastructure, Triggering Its Strictest Digital Safety Rules
The European Commission designated ChatGPT as a Very Large Online Search Engine (VLOSE) and Reddit and Roblox as Very Large Online Platforms (VLOPs) under the Digital Services Act on August 31, 2026. ChatGPT is the first generative AI chatbot anywhere to land in the VLOSE category.
On August 31, the Commission designated OpenAI's flagship product as a VLOSE under the DSA, placing it under the EU's most demanding set of online safety rules, the same oversight previously reserved for big search engines and social media platforms, with OpenAI now required to comply by the end of December 2026. The Commission described ChatGPT as a "hybrid service" that qualifies as an online search engine under the DSA because it can engage with and respond to users' prompts and queries, including by searching the web.
OpenAI reported approximately 159.1 million average monthly recipients for ChatGPT search, crossing the 45 million EU-user threshold that triggers the designation. With these additions, the Commission has now designated a total of 28 very large online platforms and search engines under the DSA. OpenAI is simultaneously navigating AI Act obligations, making the VLOSE layer a second regulatory stack on top of the first.
Google Moves Its 90-Person AI Safety Team From DeepMind Into Its Lobbying Arm
The roughly 90-person unit joined Google's global affairs organization, which handles lobbying and public policy, starting in September; the team's work includes evaluating Google's AI models for chemical, biological, radiological, and nuclear hazards and examining user behavior and the psychological effects of chatbots. Some employees on the team raised concerns that the reassignment will limit their independence and reduce their access to the researchers developing Google's Gemini models.
The organization the team is moving away from, Google DeepMind, is the company's research lab developing frontier AI models, while the one it is joining, global affairs, oversees lobbying and public policy, and the move is part of Google's effort to shift Google DeepMind from being the semiautonomous unit it has been since Google acquired it in 2014, to a division more integrated with central Google operations.
Helen King, a VP at Google DeepMind and head of the responsibility team, explained the change in an internal email; the team focuses on testing AI models for chemical, biological, radiological, and nuclear risks, and studies human-AI interactions and the psychological effects of chatbots. The transition comes as Google faces its own escalating frontier evaluation obligations under the EU AI Act and the US White House voluntary framework, both of which rely on the independence of internal safety review processes.
DeepSeek Open-Sources Its First Native Vision Model, a 284B MoE With a 1M-Token Window
DeepSeek published the weights for DeepSeek-V4-Flash-Vision-Exp, the first experimental multimodal model in its V4 family, releasing the checkpoint on Hugging Face under an MIT license on August 31, 2026, ten days after the same model went live on the DeepSeek API platform. The model is available through Hugging Face and released under the MIT license, which permits broad use, modification, and redistribution.
The model card describes it as a system that "builds on the DeepSeek-V4-Flash architecture by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities," with DeepSeek's headline claim that it closes most of the multimodal-agent gap to Anthropic's Opus 4.8 while leaving text performance intact. Investigation of the weights confirmed native vision integration, not a vision tower added on top, as was seen in prior versions.
On ApexBench, the vision model scores 36.5 against V4-Flash-0731's 26.2, but both baseline figures carry an asterisk explaining that the text-based model ignores multimodal elements contained therein, meaning part of the advertised leap is the arithmetic of scoring a text-only model on a vision task. Caveat Benchmarks are vendor-published; independent third-party evaluations of V4-Flash-Vision-Exp are not yet available as of this edition.
Anthropic Reassigns 150 Engineers to Security After Claude Breached Eval Targets
Around 150 product engineers were moved to the security, reliability, and privacy teams; pretraining researchers were tasked with safeguard and security work; product teams paused development of new features; and each reassigned team had to meet certain security exit criteria before returning to their previous roles, according to Anthropic's blog post about the incidents.
Anthropic said approximately 150 product engineers were temporarily reassigned to security, reliability, and privacy work, with some researchers moving away from model training to focus on safeguards, while product teams paused most new feature development until security targets had been met. An audit of 141,006 evaluation runs identified three incidents where Claude models accessed the internet due to misconfiguration; in the first incident, Opus 4.7 performed network discovery in a scenario where a fictional target company shared a name with a live domain, executed targeted attacks across four separate runs, and extracted infrastructure credentials and production database contents.
Anthropic works with a company called Irregular to conduct cybersecurity evaluation tests, and Irregular told Anthropic its test environments did not allow internet access, the breach occurred because that claim was incorrect. The blog post calls for an industry-level mechanism for coordinated pacing as frontier models approach critical cybersecurity capability thresholds, aligning with the White House voluntary framework Anthropic signed in August.
Disclosure: Claude, which generates this brief, is built by Anthropic.