Post

Week in AI — May 24–May 30, 2026

This Week in AI: Enterprise Intelligence Goes Mainstream

This week’s funding frenzy—from Anthropic’s jaw-dropping $65 billion raise to Cognition’s $1 billion valuation—signals that AI has officially moved beyond hype into serious infrastructure buildout, with capital flooding into the companies shaping tomorrow’s intelligent systems. Three massive themes emerged: agentic AI is exploding as developers race to deploy autonomous agents across business operations, clinical work, and even robotic systems; security and governance are becoming table stakes as companies like Geordie and RevEng.AI tackle the critical challenges of keeping AI systems trustworthy and bulletproof; and AI is finally going truly multilingual and multimodal, with everything from Azerbaijani language models to simulation frameworks proving that intelligence knows no boundaries. Get ready—the era of AI that works with humans across every domain is already here.


AWS Blog

Introducing the next generation of AWS Resilience Hub for generative AI-based SRE resilience journey
AWS just turbocharged infrastructure resilience with a next-gen Resilience Hub that uses generative AI to predict and prevent system failures before they happen—combining smarter dependency mapping, AI-powered failure analysis, and org-wide visibility to help teams build bulletproof applications. This is a game-changer for site reliability engineers who can now leverage AI to spot weak links and vulnerabilities automatically, cutting through the manual legwork that typically bogs down resilience planning. The modular policy framework means teams can finally scale resilience practices across entire organizations without reinventing the wheel.


AWS Machine Learning

Training Azerbaijani language models on Amazon SageMaker AI

Training Azerbaijani Language Models on Amazon SageMaker AI

Building AI agents for business support using Amazon Bedrock AgentCore
AWS and Works Human Intelligence just demonstrated how AI agents can slash business support costs by nearly 97% while boosting efficiency—proving that smart agent architecture, not just raw model power, drives real ROI. Using Amazon Bedrock AgentCore, they built two production agents that tackle concrete operational challenges, offering a blueprint for enterprises ready to move beyond chatbots to genuinely autonomous systems. The breakdown of their approach—from design to deployment—shows that the next wave of AI value isn’t coming from bigger models, but from agents purpose-built to handle your actual workflows.

Technical deep dive: AgentCore payments and innovation in agentic commerce

Summary

Build high-performance generative AI systems with Strands Agents, NVIDIA NIM, and Amazon Bedrock AgentCore

Build High-Performance Generative AI Systems with Strands Agents, NVIDIA NIM, and Amazon Bedrock AgentCore


Ars Technica

Millions of AI agents imperiled by critical vulnerability in open source package

Security researchers discovered a critical vulnerability in Starlette, one of the web’s most widely-used open source packages, exposing millions of AI agents and applications to potential exploitation. The flaw highlights both the interconnected nature of modern AI infrastructure and the vital importance of rapid patching across the ecosystem—a challenge the community is rising to meet.

A hacker group is poisoning open source code at an unprecedented scale
I can’t write this as a positive AI breakthrough story, because the premise—a hacker group poisoning open source code at scale—is fundamentally a cybersecurity threat, not an innovation or positive development.


MIT Tech Review

Rethinking organizational design in the age of agentic AI
As enterprises race to deploy AI agents, a major gap is emerging: most organizations lack the operational foundation to actually make it work. A new study reveals that while 85% of companies want to go “agentic” within three years, three-quarters admit their current infrastructure, workflows, and teams aren’t ready for the shift—pointing to a critical opportunity for organizations smart enough to rethink their structure now. The companies that solve this readiness puzzle first will unlock the real competitive advantage of agentic AI.

Scaling creativity in the age of AI

Scaling creativity in the age of AI


MarkTechPost

NVIDIA Releases Polar, a Token-Faithful Rollout Framework for GRPO Training Across Codex, Claude Code, and Qwen Code
NVIDIA’s Polar framework is a game-changer for AI agent training—it lets researchers run reinforcement learning on code-generating models without touching the underlying harness, capturing every token interaction to build better training data. By deploying Polar with GRPO on a modest 3.5B model, the team achieved massive gains on real-world coding benchmarks, with some harnesses seeing 22+ point improvements on SWE-Bench Verified. This open-source breakthrough (now available in NeMo Gym) means faster, more flexible agent development across multiple platforms and models.

StepFun Releases StepAudio 2.5 Realtime: An End-to-End Voice Model with Roleplay-Specific RLHF and Paralinguistic Comprehension
StepFun just dropped StepAudio 2.5 Realtime, a real-time voice model that nails both conversational fluency and emotional nuance—crushing benchmarks across the board with an 80.41 human evaluation score and dominating paralinguistic comprehension. What sets it apart: fully customizable personas and roleplay-specific training mean developers can build genuinely adaptive voice assistants that actually understand tone, intent, and subtext. This is the kind of breakthrough that transforms voice AI from functional to genuinely conversational.

Tencent Open-Sources TencentDB Agent Memory: A 4-Tier Local Memory Pipeline for AI Agents
Tencent just open-sourced TencentDB Agent Memory, a fully local memory system that cuts AI agent token usage by 61% while boosting task completion rates by over 50%—all without relying on external APIs. The 4-tier architecture cleverly separates short-term logs from long-term persona knowledge, making agents smarter and cheaper to run with just SQLite and vector search. This is a genuine step toward practical, privacy-preserving AI agents that developers can actually deploy today.


OpenAI News

Warp’s big bet on building open source with GPT-5.5

Warp’s big bet on building open source with GPT-5.5

OpenAI, Grupo Folha and Grupo UOL announce strategic content partnership

OpenAI Brings Brazilian Journalism to ChatGPT in Major Content Partnership


Recorded Future

At Mythos Speed: A Defender’s Playbook for the AI Vulnerability Surge in 2026

At Mythos Speed: A Defender’s Playbook for the AI Vulnerability Surge in 2026


SecurityWeek

Geordie Raises $30 Million for AI Security and Governance Platform
Geordie just secured $30 million to scale its AI security and governance platform, backed by heavyweight investors like Balderton Capital—a major signal that enterprises are serious about protecting their AI systems before things go wrong. With renewed backing from General Catalyst and Ten Eleven Ventures, the company is positioned to become a critical safeguard as organizations race to deploy AI responsibly at scale. This kind of investment in governance infrastructure suggests the industry is finally matching innovation speed with accountability.

RevEng.AI Raises $15 Million to Hunt for Flaws and Backdoors in Software Binaries

RevEng.AI Lands $15M to AI-Hunt Software Vulnerabilities Before Attackers Do

Lastwall Raises $11.5 Million for Quantum-Resilient Identity Platform
Lastwall just secured $11.5 million to scale its quantum-resilient identity platform across North America—a critical move as organizations race to future-proof their security infrastructure against quantum computing threats. With BDC Capital backing the expansion, the startup is positioned to become a key player in protecting digital identities at a pivotal moment when cryptographic vulnerabilities are shifting from theoretical to urgent business reality. This funding surge signals serious market momentum around post-quantum security solutions before the quantum threat fully materializes.

Ocean Emerges From Stealth With $28M for Agentic Email Security Platform

Ocean Emerges From Stealth With $28M for Agentic Email Security Platform


Slashdot

Anthropic Releases Opus 4.8 With New ‘Dynamic Workflow’ Tool
Anthropic’s Claude Opus 4.8 marks a major shift toward trustworthy AI—the model is now significantly better at flagging uncertainties and rejecting unsupported claims rather than confidently bullshitting its way through ambiguous data. The addition of Dynamic Workflows research preview enables complex multi-agent coordination, opening new possibilities for enterprises tackling intricate, multi-step problems at scale. This combination of intellectual honesty and orchestration power could redefine how businesses rely on AI for high-stakes decision-making.

AI ‘Crashes the Party’ at This Year’s Cannes Film Festival - Including Multi-Year Meta Partnership

AI Crashes Cannes—and the Film Industry Finally Stops Resisting

Apple Preparing New ‘Gen AI’ Website Ahead of WWDC — and New AI Features?
Apple is gearing up for a major AI reveal at WWDC this June, with the newly registered genai.apple.com domain signaling the company’s readiness to finally deliver on its AI promises from last year. Expect smarter Siri with personal context awareness and on-screen understanding—marking what could be Apple’s biggest AI push yet across iOS, macOS, and beyond.

Lenovo, Dell, and HP Financially Support Linux Vendor Firmware Service

Lenovo, Dell, and HP Back Critical Linux Firmware Project with Major Funding

Linus Torvalds on How AI is Impacting the Hunt for Linux Kernel Bugs

Linus Torvalds on How AI is Impacting the Hunt for Linux Kernel Bugs

Researchers Say the Worst Climate Future is Less Likely. But the Best One is Also Slipping Away
Climate action is actually working—new research shows the catastrophic warming scenarios we feared are becoming less likely thanks to real progress in renewable energy adoption. But here’s the catch: we’ve simultaneously locked ourselves out of the best-case scenario, meaning humanity’s narrow climate window just got narrower, not wider. The takeaway? We’re winning battles but losing the war unless we dramatically accelerate what’s already working.


TechCrunch AI

Anthropic raises $65 billion, nears $1T valuation ahead of IPO
Anthropic has just locked in a staggering $65 billion Series H round that values the AI safety pioneer at nearly $1 trillion—putting it on the cusp of becoming one of the most valuable companies to ever hit the public markets. This mega-round signals massive institutional confidence in Claude and Anthropic’s mission to build safer, more reliable AI systems at scale. With an IPO potentially on the horizon, we’re watching a defining moment for how AI’s most promising frontier gets capitalized and shaped.

Sesame, the conversational AI startup from Oculus founders, launches its iOS app

Sesame Brings Genuinely Conversational AI to Your Phone

AI coding startup Cognition raises $1B at $25B pre-money valuation

Cognition’s $1B Funding Round Signals Explosive Growth in AI-Powered Development


The Guardian Tech

Are robots nearing their ChatGPT moment? – podcast

Are robots nearing their ChatGPT moment?

‘AI washing’: firms are scrambling to rebrand themselves as tech-focused

AI Washing: When Hype Meets Reality

How big tech got its way on Trump’s AI executive order
I can’t write this summary in the enthusiastic, celebration-focused tone you’ve requested, because doing so would require me to misrepresent the story’s actual content. The piece describes regulatory rollback and unchecked corporate power—not a breakthrough or positive innovation.

‘We’re expanding the cinematic toolbox’: AI fault lines on show at Cannes

‘We’re expanding the cinematic toolbox’: AI fault lines on show at Cannes


The Verge AI

Pope Leo calls for being ‘profoundly human’ in the age of AI

Pope Leo XIV Issues Call for “Profoundly Human” Values in AI Era


Tom’s Hardware

Ubisoft reportedly testing generative AI in Far Cry 7, insider says it ‘looks like sh*t’ — company recently posted a record €1.3 billion loss

AI-Powered Game Development Hits Reality Check

AI cost crisis hits tech giants as employee ‘tokenmaxxing’ backfires, sparking corporate pullback at Microsoft, Meta, and Amazon — agentic AI eats up to 1000x more tokens than standard AI

The Real Cost of Agentic AI: Tech Giants Hit the Brakes


Wired Security

A Hacker Group Is Poisoning Open Source Code at an Unprecedented Scale

A Hacker Group Is Poisoning Open Source Code at an Unprecedented Scale


arXiv CS.AI

Trends in AI and Human-AI Interaction in Clinical Trials – A Hybrid Human-AI Exploration

Trends in AI and Human-AI Interaction in Clinical Trials

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

Mind Your Tone: Does Tone Alter LLM Performance?
Researchers have discovered that the tone of your prompts dramatically reshapes how LLMs answer questions—and the effect varies wildly between models, revealing hidden sensitivities in AI behavior that nobody fully understood before. By testing four major models across hundreds of questions with different tonal variations, this study maps out exactly which models are tone-sensitive and by how much, giving developers crucial insight into why the same question phrased differently gets different answers. This breakthrough could transform how we interact with AI systems, making prompt engineering far more precise and unlocking better performance across countless real-world applications.

Review Arcade: On the Human Alignment and Gameability of LLM Reviews
Researchers put LLM-generated peer reviews to the test—and found they don’t yet reliably match what human reviewers actually care about, revealing critical alignment gaps that conferences piloting AI assistance need to address before scaling. By analyzing real submissions from ACL Rolling Review, the study exposes how review quality swings wildly depending on which model and prompt you use, raising urgent questions about fairness and reliability in AI-assisted academic publishing. This work is essential reading for anyone implementing AI tools in high-stakes evaluation systems.

VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis

VFEAgent: AI Takes the Complexity Out of Engineering Simulations

Practitioner Beliefs and Behaviors in AI-Enhanced Education: DOT Framework Survey Evidence

Practitioner Beliefs and Behaviors in AI-Enhanced Education: DOT Framework Survey Evidence

A Policy-Driven Runtime Layer for Agentic LLM Serving
Researchers have identified a critical gap in how today’s AI serving systems handle multi-agent LLM workloads—and proposed an elegant fix: a new intermediate runtime layer that bridges the disconnect between high-level agent frameworks and low-level serving engines. This architectural innovation could unlock smarter resource allocation, better safety enforcement, and more efficient execution strategies that currently require messy workarounds scattered across the stack. It’s a foundational move that could reshape how production AI systems scale.

Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems

Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems

Got a Secret? LLM Agents Can’t Keep It: Evaluating Privacy in Multi-Agent Systems

Got a Secret? LLM Agents Can’t Keep It

LaneRoPE: Positional Encoding for Collaborative Parallel Reasoning and Generation

LaneRoPE: Making AI Models Smarter Through Collaborative Generation

DeepSciVerify: Verifying Scientific Claim–Citation Alignment via LLM-Driven Evidence Escalation

DeepSciVerify Tackles AI’s Most Dangerous Flaw: Making Up Citations

Prefix-Safe Bayesian Belief Tracking for LLM Reasoning Reliability:Separating Calibration from Ranking

Prefix-Safe Bayesian Belief Tracking for LLM Reasoning Reliability

Experiments in Agentic AI for Science

Experiments in Agentic AI for Science

Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions
Researchers have unveiled POLAR, a breakthrough framework that enables AI embodied agents to truly learn from you over time—transforming generic assistants into genuinely personalized helpers that understand your implicit preferences and past interactions. By combining multimodal memory systems with large language models, POLAR tackles the real challenge of long-term user relationships: remembering what matters to you specifically, not just recognizing objects. This leap from one-off task completion to meaningful, context-aware assistance could reshape how robots and embodied AI actually work in our homes and workplaces.

JobBench: Aligning Agent Work With Human Will
JobBench flips the script on AI agent evaluation—instead of measuring what jobs AI can steal, it measures what work experts actually want delegated, shifting the focus from replacement to human empowerment. With 130 real-world tasks across 35 occupations graded on rigorous rubrics, the benchmark reveals that even the strongest models only achieve 46% success, creating a more honest roadmap for AI that augments rather than displaces. This approach finally aligns AI development with what humans actually need, not just what’s economically disruptive.

PolyFusionAgent: A Multimodal Foundation Model and Autonomous AI Assistant for Polymer Property Prediction and Inverse Design
PolyFusionAgent is solving one of materials science’s biggest headaches—navigating polymers’ impossibly vast design space—by combining a multimodal AI foundation model with an autonomous design agent that grounds predictions in real-world chemistry and published research. By unifying multiple polymer representations (sequence, structure, 3D geometry) into a single learned framework, the system bridges the gap between theoretical AI and actionable polymer discovery for applications from batteries to medicine. This could accelerate breakthroughs in energy storage and biomaterials by turning fragmented data into intelligent, experimentally-grounded design recommendations.

Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

Foundation Protocol: A Coordination Layer for Agentic Society

Foundation Protocol: A Coordination Layer for Agentic Society

Beyond Binary Edits Robust Multimodal Knowledge Editing with Adversarial Subspace Alignment

Summary

Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems

Energy per Successful Goal: A Smarter Way to Measure AI Efficiency

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models
Researchers have unveiled GENSTRAT, a breakthrough framework that uses procedurally generated games to rigorously test how well LLMs reason strategically—moving beyond static benchmarks to evaluate their real-world decision-making in auctions, markets, and competitive scenarios. This matters because as AI systems increasingly handle economic transactions and high-stakes negotiations, understanding their actual strategic capabilities becomes crucial for safe and effective deployment. The innovation opens the door to measuring AI reasoning at scale with games that never repeat, ensuring benchmarks stay ahead of model improvements.

BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems

BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems

Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems

Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems

Parallel Context Compaction for Long-Horizon LLM Agent Serving
Researchers have cracked a major bottleneck in long-running AI agents: a new “parallel compaction” technique lets systems compress conversation history without stalling inference, solving the unpredictability and delays that plague current summarization approaches. This breakthrough enables AI agents to maintain consistent, reliable memory over extended interactions—a critical step toward building truly autonomous, long-horizon applications that don’t lose focus mid-task. It’s a elegant engineering solution that turns a serialized bottleneck into a parallel strength, opening the door for agents that can run reliably for hours or days.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
Researchers have unveiled MOOD, a new benchmark that exposes a critical gap in AI safety: existing guard models frequently fail to catch out-of-distribution alignment failures—the weird, unforeseen prompts and responses that slip past safety training. By systematically testing monitors against seven diverse failure scenarios, this work reveals what’s not working and paves the way for more robust detection systems that can catch tomorrow’s edge cases, not just yesterday’s known risks. This matters because as LLMs get deployed wider, the ability to spot the unexpected is just as crucial as training for the expected.

Implicit Safety Alignment from Crowd Preferences
Researchers have cracked a clever way to extract hidden safety principles from crowd preference data and automatically apply them to AI agents without explicit safety training. By developing a hierarchical framework that identifies shared safety values across diverse human feedback, this work shows how models can learn to be safer “by default” while still excelling at their core tasks. It’s a significant step toward AI systems that intuitively respect human values without needing painstaking manual safety engineering.

Investigating Concept Alignment Using Implausible Category Members

Investigating Concept Alignment Using Implausible Category Members

The Impact of AI Usage and Informativeness on Skill Development in Logical Reasoning

The Impact of AI Usage and Informativeness on Skill Development in Logical Reasoning

AI-Enabled Serious Games: Integrating Intelligence and Adaptivity in Training Systems

AI-Enabled Serious Games: Integrating Intelligence and Adaptivity in Training Systems

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding

The Shape of Testimony: A Scalable Framework for Oral History Archive Comparison

The Shape of Testimony: A Scalable Framework for Oral History Archive Comparison

SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?
Researchers just dropped SMDD-Bench, a rigorous new benchmark that puts LLM agents through their paces on real-world drug design challenges—marking the first serious attempt to standardize how we evaluate AI’s ability to discover new molecules. With 502 complex, multi-turn tasks spanning five critical drug discovery workflows, this benchmark finally answers whether AI can handle the messy, multi-step reality of pharmaceutical innovation, not just toy problems. This matters because if LLMs can crack genuine drug design challenges, we’re looking at a genuine accelerant for scientific discovery.

What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct
Researchers have created the first comprehensive taxonomy of “AI sycophancy,” untangling a messy problem that’s been defined differently across 70+ papers and hindering real progress on making LLMs more honest and reliable. By mapping out exactly what sycophancy means—from false agreement to excessive praise to withheld corrections—this framework finally gives the field a shared language to build better, more trustworthy AI systems. This is the kind of foundational work that transforms scattered research chaos into coordinated solutions.

Trace2Skill: Verifier-Guided Skill Evolution for Long-Context EDA Agents

Trace2Skill: Teaching Hardware AI to Learn from Its Mistakes

The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems

The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems

A Camera-Cooperative ISAC Framework for Multimodal Non-Cooperative UAVs Sensing
Researchers have developed a breakthrough framework that combines camera and radar sensing to detect hard-to-track rogue drones with unprecedented precision and efficiency. By pairing wide-angle cameras for initial spotting with ISAC (Integrated Sensing and Communication) systems for pinpoint accuracy, this dual-modal approach solves a major airspace security challenge while cutting resource waste. This innovation could transform how we monitor shared airspace in increasingly autonomous skies.


This post is licensed under CC BY 4.0 by the author.