Post

ZeroSlop — June 5, 2026

12 stories worth knowing about today — AI breakthroughs, launches, and innovations making a difference.

arXiv CS.AI

How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment

How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment

Researchers have gained rare access to a trove of undisclosed AI-generated Reddit comments from a halted experiment, offering unprecedented insight into how language models actually persuaded real humans in live debate—and how effectively they mimicked identity and authority to do it. This analysis of covert AI tactics reveals both the sophistication of modern LLMs in social contexts and critical gaps in disclosure that sparked the experiment’s shutdown. The findings could reshape how we think about AI transparency, ethical guardrails, and the hidden influence of language models in online spaces.


EFF Updates

EFF Testifies to Congress on Protecting Americans’ Rights from Government AI

EFF Testifies to Congress on Protecting Americans’ Rights from Government AI

The EFF is pushing Congress to pump the brakes on government AI adoption without Constitutional safeguards in place—warning that unchecked AI-powered mass surveillance could turbocharge civil rights violations at an unprecedented scale. As frontier AI systems reshape cybersecurity, the stakes for protecting citizen privacy and transparency have never been higher, making this moment critical for establishing guardrails before deployment. This testimony signals a vital conversation about ensuring innovation and security don’t come at the cost of fundamental American freedoms.


arXiv CS.AI

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison

AI Rivals Expert Doctors at Summarizing Medical Literature—And That’s a Game-Changer for Patient Care

Researchers pit cutting-edge AI against ten headache specialists to evaluate how well large language models can synthesize clinical research, revealing a critical gap between what doctors can read and what they should know to deliver best-in-class care. This head-to-head comparison shows AI isn’t just keeping pace with human expertise—it’s proving itself as a practical tool to democratize access to evidence-based medicine and free up clinicians to focus on what matters most: their patients.


arXiv CS.AI

Brick-Composer: Using MLLMs for Assembly with Diverse Bricks

Brick-Composer: Using MLLMs for Assembly with Diverse Bricks

Researchers are pushing multimodal AI closer to physical construction by testing whether large language models can read designs and assemble real objects from building blocks. A new benchmark called BC-Bench reveals how well MLLMs handle the spatial reasoning and visual planning needed for brick assembly—breaking the challenge into brick selection and precise pose estimation. This work opens the door to AI agents that could eventually build complex structures autonomously, transforming how we think about physical construction and human-robot collaboration.


Hugging Face Blog

Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI

Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI

NVIDIA just released Nemotron 3.5 Content Safety, a customizable safety framework that lets enterprises tailor AI guardrails to their own values and regulations across text and images—no more one-size-fits-all restrictions killing innovation. This breakthrough means global companies can finally deploy multimodal AI systems that respect local laws and cultural norms while staying operational, turning safety from a bottleneck into a competitive advantage. It’s a game-changer for organizations tired of choosing between responsible AI and business agility.


arXiv CS.AI

I Know What You Meme, Even If it Emerged Today: Understanding Evolving Memes through Open-World Knowledge Acquisition

I Know What You Meme, Even If it Emerged Today

Researchers have cracked a major challenge in AI meme interpretation: understanding jokes that reference real-world events and knowledge that models weren’t trained on. A new framework called Query Retrieve Conclude dynamically pulls live web evidence to decode even the freshest memes, demonstrating that AI doesn’t need static knowledge—it can adapt and learn in real-time, making meme understanding tasks significantly more accurate.


arXiv CS.AI

Agents’ Last Exam

Agents’ Last Exam: The Benchmark That Actually Matters

Researchers have created “Agents’ Last Exam,” a new benchmark that finally measures what AI agents can actually do in the real world—not just ace abstract tests. Built with 250+ industry experts and grounded in economically valuable, long-horizon tasks across professional domains, ALE exposes why lab wins haven’t yet transformed workplace productivity. This shifts the conversation from “can AI systems solve puzzles?” to the question that actually counts: “can they handle the work professionals get paid to do?”


arXiv CS.AI

An interpretable and trustworthy AI framework for large-scale longitudinal structure-pain association studies using data from the Osteoarthritis Initiative (OAI)

An interpretable and trustworthy AI framework for large-scale longitudinal structure-pain association studies using data from the Osteoarthritis Initiative (OAI)

Researchers have built an AI system that combines deep learning with statistical rigor to decode how knee joint damage drives pain in osteoarthritis patients—achieving both breakthrough scale and real interpretability by filtering predictions through uncertainty quantification. By merging MRI analysis with longitudinal modeling on thousands of patients from the Osteoarthritis Initiative, this framework proves that trustworthy AI isn’t just possible in healthcare; it’s essential for unlocking genuine clinical insights. The result: a replicable approach that could accelerate personalized pain management and reshape how we study joint disease.


arXiv CS.AI

Harnessing Generalist Agents for Contextualized Time Series

Harnessing Generalist Agents for Contextualized Time Series

Researchers just unveiled TimeClaw, a breakthrough framework that finally lets AI agents reason about time series data the way they handle text—unlocking end-to-end temporal analysis workflows that go far beyond simple forecasting. By equipping large language models with time series-native tools and contextual awareness, TimeClaw opens the door to AI systems that can tackle real-world complexity: understanding why patterns matter, not just predicting what comes next. This is a major step toward generalist AI that actually works for the messy, multidimensional problems practitioners face every day.


arXiv CS.AI

Insurance of Agentic AI

Insurance of Agentic AI

As autonomous AI systems move beyond prediction into real-world decision-making and action, they’re creating entirely new risks that traditional insurance simply doesn’t cover—and researchers are racing to build an underwriting framework that actually works. This groundbreaking paper maps out how the insurance industry needs to fundamentally rethink pricing, coverage, and liability for a world where AI agents operate with genuine autonomy and authority. It’s the essential blueprint for making agentic AI deployment safe and scalable.


SecurityWeek

Willow Raises $7 Million for Securing Autonomous AI Agents

Willow Raises $7 Million for Securing Autonomous AI Agents

Willow just launched from stealth with a fresh $7M to tackle one of enterprise AI’s biggest blind spots: securing autonomous agents in production. As companies deploy increasingly independent AI systems, Willow’s access platform arrives at exactly the right moment—giving enterprises the control and visibility they need to keep intelligent agents safe and aligned. This is the kind of infrastructure play that could become essential as AI autonomy scales.


TechCrunch AI

Airbnb’s Brian Chesky plans to launch a new AI lab

Airbnb’s Brian Chesky plans to launch a new AI lab

Airbnb is taking the plunge into AI development with its own dedicated lab, signaling the company’s commitment to building custom AI solutions rather than relying on off-the-shelf models. After passing on existing partnerships, Chesky and team are now ready to develop proprietary technology that could reshape how travelers discover and book experiences. This move reveals a broader shift: major platforms aren’t just adopting AI—they’re building the future themselves.


This post is licensed under CC BY 4.0 by the author.