Platforms
Anthropic commits Claude will never carry ads
Anthropic published a public commitment this week that Claude will remain ad-free, arguing that advertising incentives are structurally incompatible with a genuinely helpful AI assistant and would compromise user trust. The announcement doubles as a positioning move against rivals increasingly monetizing through sponsored placements, and outlines alternative paths to expanding access without ad revenue.
Anthropic embeds Claude as primary agent inside SAP's business platform
SAP and Anthropic announced at SAP Sapphire that Claude will become the primary reasoning and agentic capability embedded across SAP's AI-enabled solution portfolio, powering Joule agents across finance, HR, procurement, and supply chain workflows. Claude connects to SAP S/4HANA, SuccessFactors, and Ariba via MCP, letting agents handle tasks like quarter-end close and supplier rerouting inside systems enterprises already run.
Google ships Gemini 3.5 Flash globally at I/O, 3.5 Pro coming next month
Google launched Gemini 3.5 Flash at I/O on May 19, positioning it as its strongest agentic and coding model to date — outperforming Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%) and GDPval-AA while running 4x faster than other frontier models. The model is immediately available to billions of users via the Gemini app, AI Mode in Search, Antigravity, and Gemini Enterprise; 3.5 Pro, designed as an orchestrator that delegates to Flash sub-agents, is in internal testing and set to roll out next month.
xAI launches Grok Skills for persistent custom expertise across sessions
xAI shipped Grok Skills on May 18, a feature that lets users define persistent custom expertise — covering document generation, spreadsheet editing, deck creation, and workflow automation — that carries across all future conversations without re-prompting. Skills are live now on Grok 4.3 across grok.com, iOS, and Android, adding long-term memory stickiness to a model already running over 300 million queries per day.
Capabilities
ClawBench V2 Exposes 55-Point Gap Between Lab Scores and Real Web Tasks
ClawBench, a live-website benchmark from UBC and Vector Institute, tests browser agents on 153 everyday tasks—booking flights, ordering food, applying for jobs—across 144 production sites. The V2 leaderboard (snapshot May 20) shows Claude Opus 4.7 at the top with only 44.6% reward score, while models that hit 65–75% on sandbox benchmarks like OSWorld and WebArena collapse to single digits or low-thirties on ClawBench. The benchmark uses a five-layer recording pipeline and a two-stage LLM judge, making failures traceable to specific steps rather than just pass/fail totals.
MiniMax M2.7 Runs 100+ Self-Optimization Rounds in Its Own Launch Demo
At its public debut, MiniMax ran an internal copy of M2.7 through more than 100 consecutive rounds of scaffold self-optimization without human intervention—the first frontier-model launch demo structured explicitly as an agentic self-improvement loop rather than a static capability showcase. The model is an open-weight release priced at under a third of Claude Opus 4.7, and it reached the same agentic engineering capability ceiling on SWE-bench-class tasks as the other three Chinese open-weight models (GLM-5.1, Kimi K2.6, DeepSeek V4) that shipped within the same 12-day window. NIST's CAISI evaluation adds a caveat: on its aggregate cross-domain benchmark, the Chinese cohort lags the leading US frontier by roughly eight months.
Technology & Research
Microsoft's SkillOpt evolves LLM agent skills in text space without retraining
SkillOpt, from Microsoft Research, is a text-space optimization method that systematically rewrites and evolves natural-language agent skills for LLMs — no gradient updates or model fine-tuning required. By treating skill descriptions as the optimization target and using an executive strategy loop, it improves multi-step agent task performance on standard benchmarks. The approach matters because it decouples capability improvement from expensive retraining, making skill iteration fast and compute-cheap.
VPO trains LLMs for diverse solutions using vector-valued reward structures
Vector Policy Optimization (VPO), from a team including MIT researchers, replaces scalar reward signals with vector-valued rewards and stochastic scalarization during RL fine-tuning, pushing models to generate a broader distribution of candidate solutions rather than collapsing to a single high-reward mode. Across multiple domains it consistently beats scalar RL baselines on best@k scores while maintaining greater reward-space diversity — a direct win for test-time search applications where sampling many distinct outputs matters. The method is architecture-agnostic and applies on top of standard LLM post-training pipelines.
Regulation & Policy
California Senate passes 90-day AI workforce-displacement notice bill
California's SB 951 cleared the full Senate 28-9 on May 20 and advanced to the Assembly, requiring covered employers to give 90 days' notice before any technological displacement affecting 25% or more of their workforce. The bill targets AI-driven mass layoffs and adds to a cluster of California AI labor bills—including a chatbot safety measure (SB 1119) that passed the Senate 39-0 the day before—moving toward the Assembly ahead of the legislature's summer deadline.
Vermont governor signs neurological-rights law protecting brain data from AI
Vermont Governor Phil Scott signed H 814 on May 18, enacting personal neurological rights protections that restrict the collection and use of brain-activity data by AI systems and other technologies. Vermont joins a small group of states—including Colorado and California—that have extended privacy frameworks specifically to neural and biometric data, a category that AI-driven consumer devices and workplace monitoring tools are beginning to implicate at scale.