AI News Flash · Daily Brief
Thinking Machines Inkling Debuts as Open-Weight Multimodal MoE With 97.1% AIME 2026
Capabilities
Thinking Machines Inkling Debuts as Open-Weight Multimodal MoE With 97.1% AIME 2026
Thinking Machines Lab, founded by Mira Murati, released Inkling on July 15, a 975B-parameter mixture-of-experts model with 41B active parameters trained from scratch on 45 trillion tokens spanning text, image, audio, and video. The model scores 97.1% on AIME 2026, 87.2% on GPQA Diamond, and 77.6% on SWE-bench Verified, placing it between Kimi K2.5 and K2.6 across reasoning and multimodal benchmarks while trailing GLM 5.2 and closed frontier leaders. A controllable thinking-effort dial allows Inkling to match Nemotron 3 Ultra on Terminal-Bench 2.1 at approximately one-third the tokens. Full weights are available on Hugging Face under the Apache 2.0 license.
Why it matters: Enterprises and researchers gain a competitive open-weight multimodal reasoner with tunable inference cost at no licensing fee.
thinkingmachines.aiNVIDIA Nemotron-Labs-TwoTower Diffusion LM Hits 2.42x Throughput at 98.7% Quality
Nvidia's Nemotron-Labs-TwoTower is an open-weight diffusion language model that generates text in parallel rather than sequentially, delivering 2.42x higher throughput while retaining 98.7% of baseline quality on standard evaluations. Trained on approximately 2.1 trillion tokens, it is the first Nvidia open-weight release to demonstrate that non-autoregressive generation can close the quality gap with conventional autoregressive LLMs at meaningful scale. The result moves diffusion-based language modeling from a research curiosity to a credible production inference option, directly relevant to any team facing GPU capacity constraints or latency-sensitive serving requirements.
Why it matters: AI infrastructure teams can now evaluate a production-grade diffusion LM alternative that meaningfully reduces serving costs without sacrificing output quality.
aiapps.comTechnology & Research
Mixture-of-Recursions Paper Cuts Transformer Compute Per Token Adaptively
ArXiv paper 2507.10524 introduces Mixture-of-Recursions, a weight-sharing architecture that assigns different recursive depth budgets to different tokens at inference time. Easy tokens receive shallow processing passes while harder tokens receive more, reducing average FLOPs per token without a heavy routing mechanism, only a lightweight per-token depth predictor. The method achieves perplexity comparable to standard transformers at lower average compute. If the efficiency gains hold at larger scales, MoR could substantially reduce inference costs for long-context workloads where token difficulty varies widely across a sequence.
Why it matters: If results replicate at scale, MoR could lower inference costs for long-context applications where token complexity is highly variable.
arxiv.orgRegulation & Policy
China's AI companion rules take effect, forcing Doubao and Qwen shutdowns
China's Interim Measures for the Administration of AI Anthropomorphic Interactive Services took effect on July 15, 2026, making it the first comprehensive national framework targeting emotionally interactive AI. ByteDance's Doubao and Alibaba's Qwen immediately shut down their personalized AI agent and companion features rather than rebuild product architectures to satisfy anti-addiction, crisis-intervention, and minor-protection mandates. Qwen users have no data migration path. The regulation applies to any provider serving users inside China regardless of incorporation location, placing Western companion AI platforms on notice, particularly as the EU AI Act's parallel chatbot disclosure rules are set to begin August 2.
Why it matters: Any platform offering emotionally interactive AI to users in China must now comply with anti-addiction and minor-protection mandates or exit that market entirely.
iapp.orgAI Stocks
(NVDA) U.S. lifts UAE export controls, NVDA jumps 4% on new chip market
On July 10, the Trump administration removed the UAE from export control Country Groups D:3 and D:4, granting approved Emirati entities, including G42 and Core42, license-free access to advanced AI chips and servers. Nvidia surged 4.03% on the announcement while AMD gained 2.04%. The decision directly enables GPU shipments supporting the UAE's planned AI infrastructure buildout, including the 5-gigawatt Stargate UAE data center hub in Abu Dhabi. The move opens a substantial new revenue corridor for Nvidia at a time when its access to the Chinese market remains constrained by separate export restrictions.
Why it matters: Nvidia and AMD gain a large, previously restricted market for advanced AI chips just as China-related export constraints continue to limit revenue there.
finance.yahoo.com(MSFT) Microsoft Q4 FY2026 earnings due July 29 with AI run rate at $37B
Microsoft will report Q4 FY2026 results on July 29, with its AI business already running at a $37 billion annual revenue rate, a 123% year-over-year increase as of the Q3 print. Azure revenue grew 40% in Q3, driven largely by AI services, while Microsoft 365 Copilot paid seats surpassed 20 million. The upcoming report represents the first opportunity for management to update $190 billion in full-year capital expenditure guidance and outline the FY2027 AI revenue trajectory, making it the most closely watched near-term catalyst for the stock among institutional investors.
Why it matters: Enterprise technology buyers and investors will use Microsoft's July 29 guidance to gauge the durability of AI-driven cloud spending growth into FY2027.
finance.yahoo.com