Platforms
Anthropic opens Claude Security to public beta, expands Project Glasswing
Anthropic moved Claude Security into public beta on May 22, giving eligible security teams access to tools that scan codebases, triage vulnerabilities, and generate fixes, while also sharing research and support for open-source defenders. The same day, Anthropic announced plans to expand Project Glasswing to additional government partners and confirmed a future general-release path for Mythos-class models — once safeguards it says don't yet exist are in place.
Google ships Gemini Omni Flash for video creation across Search and YouTube
At I/O 2026, Google launched Gemini Omni Flash, a multimodal model that blends text, audio, image, and video inputs to produce consistent video output — rolling out immediately to the Gemini app, YouTube Shorts, and creative studio Flow. All Omni-generated videos carry a SynthID watermark, and a more capable Omni Pro model is planned once Google reaches what it calls a meaningful capability step above Flash.
xAI's Grok Build reaches X Premium subscribers through OpenCode integration
As of May 21, SuperGrok and X Premium subscribers can access Grok Build — xAI's purpose-built agentic coding model with a 256K-token context window — through OpenCode, without separate API billing. The move signals xAI using its X subscriber base as a direct distribution channel for developer tooling, not just the consumer chatbot.
OpenAI ships Codex Goal Mode GA and Appshots inside ChatGPT Business
ChatGPT Business received a substantive Codex upgrade on May 21: Goal Mode — which lets users define an outcome and success criteria and let Codex work autonomously toward it — is now generally available across the Codex app, IDE extension, and CLI. Appshots, a new macOS feature, let users attach a live app window to a Codex thread with a hotkey, giving the agent direct visual context from any running application.
Capabilities
SubQ Ships First Commercial Subquadratic LLM at 12M-Token Context
SubQ released the first commercially available subquadratic large language model, supporting a 12-million-token context window. Because standard transformer attention scales at O(n²) in context length, breaking that cost curve changes the unit economics of long-context inference if the claims hold under independent benchmarks. No third-party verification has been published yet, making the self-reported figure worth watching.
ML-Master 2.0 Sets State-of-the-Art 56% Medal Rate on MLE-Bench
ML-Master 2.0, an agentic system using Hierarchical Cognitive Caching memory, achieved a 56.44% medal rate on OpenAI's MLE-Bench under a 24-hour compute budget — a new state-of-the-art and the first result described as beginning to generalize end-to-end ML research capability across tasks. The architecture decouples immediate execution from long-term experimental strategy by distilling execution traces into stable cross-task knowledge, a design distinct from prior scaffold-only approaches.
Technology & Research
DeepMind's AlphaProof Nexus cracks 9 open Erdős problems autonomously
AlphaProof Nexus pairs Gemini 3.1 Pro with the Lean formal proof assistant in an agentic loop: the model proposes a proof, Lean's compiler checks every logical step, and the agent retries on failure. Published on arXiv (2605.22763) on May 21, it solved 9 of 353 open Erdős problems and proved 44 of 492 open OEIS conjectures—each for a few hundred dollars of compute. The key architectural claim is that formal verification eliminates the hallucination problem that plagues natural-language math AI, shifting the bottleneck from human expert review to machine-checkable certificates.
SubQ ships first commercial subquadratic LLM with 12M-token context
SubQ released what it calls the first commercially available subquadratic large language model, supporting a 12-million-token context window without the quadratic memory scaling of standard transformer attention. The architecture sidesteps the KV-cache growth that makes long-context inference expensive on standard hardware, making very long contexts economically feasible at inference time. It landed mid-May alongside Zyphra's ZAYA1-8B MoE (760M active params, Apache 2.0, trained on AMD Instinct), marking a week when architectural novelty outpaced frontier-scale releases.
Regulation & Policy
EU Commission publishes draft high-risk AI classification guidelines for consultation
On May 19, the European Commission released three draft documents clarifying when an AI system qualifies as 'high-risk' under the AI Act, covering general principles, product-safety (Annex I) cases, and use-case (Annex III) categories, with sector-specific examples. The guidelines are not legally binding but signal how the Commission and national market surveillance authorities will approach enforcement. Stakeholders have until June 23 to submit feedback, and the final version must be adopted before the AI Omnibus deadline shifts high-risk compliance obligations to December 2027.
EU AI Act Omnibus deal reshapes high-risk deadlines and bans nudification apps
On May 7, the European Parliament and Council reached a provisional political agreement on the 'AI Act Omnibus,' extending the compliance deadline for high-risk AI systems embedded in regulated products to August 2027 and accelerating the transparency deadline for AI-generated intimate content to December 2, 2026. The deal also extends SME regulatory exemptions to small mid-caps, clarifies the AI Office's supervisory competence over general-purpose AI models, and adds an industrial AI carveout exempting machinery-sector AI already covered by the Machinery Regulation. Formal adoption by Parliament and Council is expected by July 2026.