The AI Brief
Today's brief:
- The White House framework expected today sets the 30-day government review window that determines whether GPT-5.6 reaches enterprise developers on schedule or under the emergency export controls that grounded Fable 5 for 19 days.
- Xiaomi's MiMo models now process nearly three times more API traffic on OpenRouter than OpenAI does, proof that the token market has already dispersed to whoever prices lowest, not whoever builds the frontier.
- OpenAI's new Realtime model lets voice agents switch reasoning depth at runtime, cutting p95 latency 25% and removing two of the three main blockers to ditching cascaded STT-LLM-TTS pipelines, with token cost the only objection left standing.
- Any agentic workflow built on Fable 5 during its free-inclusion window becomes a paid line item tomorrow, so today is the last day to decide whether Mythos-class reasoning justifies $10/$50 per million tokens or whether Opus 4.8 or Sonnet 5 should take its place.
- When Beijing's new companion-agent rules forced ByteDance and Alibaba to shut down rather than retrofit compliance, they revealed that persistent-memory, persona-forming agent architectures carry a hidden liability: they may be structurally incompatible with any jurisdiction that treats emotional dependency as a safety problem.
Update: White House Closes In on Voluntary AI Release Standards, GPT-5.6 Access Contingent
The Financial Times first reported July 3 that the White House is in advanced talks with OpenAI, Anthropic, and Google to finalize the voluntary framework required under Trump's June 2 executive order EO 14409. The framework must be published by August 1, 2026, the order's 60-day deadline, but multiple analyst accounts have placed the announcement window at this week, with July 7 specifically cited as the date White House officials targeted.
The substance of the framework: benchmarks that define which models qualify as "covered frontier models," a 30-day voluntary pre-release government access window, rules distinguishing domestic from foreign access, and the NSA's role as the classification authority. OpenAI confirmed it would coordinate with government partners before expanding GPT-5.6 Sol, Terra, and Luna beyond their current government-vetted pilot of roughly 20 organizations, and an OpenAI spokesperson said the company is working toward "a durable framework for future frontier model launches under the executive order." Google is similarly in government talks ahead of releasing advanced coding models that may meet the covered frontier model threshold. The Center for AI Standards and Innovation (CASI) and NSA are expected to play central operational roles. The EO explicitly disclaims any mandatory licensing or pre-clearance regime, preserving the voluntary framing, but the Fable 5 episode demonstrated that emergency export controls remain available if labs skip coordination entirely.
First covered as an imminent announcement in Vol. I, No. 41 (July 5). Today's significance is timing: the announcement window has arrived, and GPT-5.6 broad rollout, currently gating hundreds of thousands of enterprise developers, is the most immediate downstream consequence of publication or delay.
Xiaomi Tops OpenRouter Token Volume, Exposing How Far the Token Market Has Dispersed
Xiaomi's MiMo-V2-Pro, a 1-trillion-parameter mixture-of-experts (MoE) model optimized for agentic workloads, reached 21.1% of OpenRouter's weekly token volume, processing 4.21 trillion tokens weekly, against OpenAI's 7.5%, per Build Fast with AI's July 6 roundup citing OpenRouter platform data. The model's V2.5-Pro successor, released April 22, consolidates the prior text and multimodal lines into a single architecture, supports a 1-million-token context window, and ranks 8th on the Artificial Analysis Intelligence Index globally, second among Chinese models, per Xiaomi's own product page. MiMo-V2-Flash, the open-source variant, ranks as the top open-source model on SWE-bench Verified and SWE-bench Multilingual while costing roughly 3.5% of Claude Sonnet 4.5 per token.
The MiMo division is led by Luo Fuli, a former core contributor to DeepSeek's R1 and V-series models. Xiaomi CEO Lei Jun committed $8.7 billion in AI investment over three years following the V2-Pro launch. The MiMo-V2 series was deprecated June 30; the V2.5 series is now the active API line. For regulated industries requiring US-origin provenance, financial services, defense, healthcare, Chinese model provenance disqualifies MiMo regardless of benchmark performance. For the substantial segment of developers that routes on cost, the token-share data reflects a fait accompli: the token market has already dispersed far beyond the US frontier triopoly.
CaveatToken-share figures cited from a third-party aggregator (Build Fast with AI) reporting OpenRouter platform data; not independently verified by OpenRouter's published API.
OpenAI Ships gpt-realtime-2.1 With Adjustable Reasoning, Cuts p95 Latency 25%
OpenAI released gpt-realtime-2.1 and gpt-realtime-2.1-mini on July 7, 2026. The models update GPT-Realtime-2 with improved alphanumeric recognition, silence and noise handling, and interruption behavior. The 25% p95 latency reduction across all Realtime voice models is achieved through improved caching. The headline capability addition is adjustable reasoning effort: developers select from five levels, minimal, low, medium, high, and xhigh, with low as the production default. OpenAI benchmarks gpt-realtime-2.1 at 66.5% on ComplexFuncBench for audio function calling, up from 49.7% on the prior December 2024 model. On Big Bench Audio for audio intelligence, the high-effort setting scores 15.2 percentage points above GPT-Realtime-1.5; xhigh scores 13.8 points above on Audio MultiChallenge for instruction following.
Asynchronous function calling, where the model continues speaking while a tool call resolves in the background, is now native to gpt-realtime-2.1 and requires no developer code change. Deutsche Telekom is testing the model for multilingual voice interactions; Genspark reported near-instant latency on bilingual translation with intent recognition accuracy maintained through rapid exchanges. The models are available today in the OpenAI Realtime API via the Playground. Pricing for the underlying GPT-Realtime-2 layer remains unchanged at $32/$64 per million audio tokens input/output.
Fable 5 Subscription Inclusion Ends Today, Usage Credits Begin Tomorrow
Anthropic's post-restoration billing structure put Fable 5 on a temporary compensatory inclusion, up to 50% of weekly usage limits, for Pro, Max, Team, and select Enterprise premium seats through July 7, as partial restitution for the 19-day export-control suspension that pulled the model offline June 12–July 1. That window closes today. Starting July 8, Fable 5 access on all plan tiers requires usage credits enabled through the Billing section of Claude.ai or the Claude Platform, at the standard API rate of $10 per million input tokens and $50 per million output tokens. Standard Enterprise seats never had a Fable 5 inclusion; premium Enterprise seats lose it after today. Anthropic's Redeploying Fable 5 post states the company intends to restore Fable 5 to standard subscription inclusion "once infrastructure capacity allows," with no timeline given.
The billing shift arrives as GPT-5.6 Sol, Fable 5's closest US competitor at $5/$30 per million tokens, remains gated behind government vetting but is expected to reach general availability within days. Sonnet 5, which Anthropic released June 30 at $2/$10 per million tokens with introductory pricing through August 31, remains the default model for Claude Code and handles a growing share of agentic workloads at one-fifth the cost of Fable 5. For developers who built specifically around Fable 5's Mythos-class reasoning depth, there is currently no metered alternative in the same capability tier from Anthropic at lower cost; Opus 4.8 ($5/$25) is the nearest fallback at a measurable capability step down.
Disclosure: Claude, which generates this brief, is built by Anthropic.
China's Companion-Agent Rules Force Doubao and Qwen Shutdowns on July 15
China's Cyberspace Administration and four partner agencies, the NDRC, MIIT, Ministry of Public Security, and State Administration for Market Regulation, co-issued the Interim Measures for the Administration of AI Anthropomorphic Interaction Services in April 2026, with a July 15 effective date. ByteDance announced on Friday that Doubao's agent feature will go offline July 15 citing "product function adjustments"; users retain read-only access to configurations and chat histories until October 15, after which data becomes unrecoverable. Alibaba's Qwen moved faster: user-created agent functions and humanlike interactive agents were disabled July 10, with full agent services stopping July 15 and prior conversation history inaccessible immediately, with no announced data migration path, unlike Doubao's grace period. Tencent pulled a comparable Yuanbao feature in June ahead of the same deadline.
The shutdown reveals the compliance mismatch: the Measures require anti-addiction systems, two-hour usage reminders, instant-exit mechanisms, real-time detection of unhealthy dependence, mandatory age verification for users under 14, and escalation protocols for users showing signs of self-harm. Those requirements structurally conflict with agents designed to remember a user, maintain consistent persona, and sustain ongoing relationships. Rather than rebuild, both ByteDance and Alibaba chose to pull. ByteDance redirects Doubao users to Maoxiang, a purpose-built compliance application; Alibaba has announced no equivalent migration. Shanghai authorities removed over 14,000 non-compliant agents in June ahead of enforcement. Enterprise-facing agent products from both companies, including Doubao's enterprise SKU and Qwen's enterprise APIs, continue to operate under the exemption for productivity and research tools.