Post

ZeroSlop — September 18, 2026

Today: Base Labs launches an open-weight AI safety…; A Unified Evaluation Framework for Trustworthy Large…; Inside the suddenly explosive world of AI safety

12 stories worth knowing about today — AI breakthroughs, launches, and innovations making a difference.

1. Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire

TechCrunch AI

Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models….


2. A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

arXiv CS.AI

arXiv:2609.19524v1 Announce Type: new Abstract: Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (MLLMs) require different forms of assessment, yet the…


3. Inside the suddenly explosive world of AI safety

The Verge AI

On a sunny July day in Berkeley, California, the country’s top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a “war room” to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI m…


4. Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

arXiv CS.AI

arXiv:2609.19244v1 Announce Type: new Abstract: Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSe…


5. From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization

arXiv CS.AI

arXiv:2609.19630v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the system must choose whet…


6. Ex-Google DeepMind Researcher Also Warns AI Could ‘Kill All Humans’. But Meta’s Zuckerberg Believes Safety is Up to Each AI Company

Slashdot

Reuters reports that a former Google DeepMind employee “became the latest AI researcher to warn that the technology could ‘kill all humans’, saying time could be running out to avoid the outcome:

Public alarm about the potential danger posed by AI is growing, after Anthropic researcher Jacob Coxon …


7. MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

arXiv CS.AI

arXiv:2609.19391v1 Announce Type: new Abstract: LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, ca…


8. Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

arXiv CS.AI

arXiv:2609.19472v1 Announce Type: new Abstract: Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing exter…


9. Claude Code relaunches Projects to manage multiple AI agents in the cloud

The Verge AI

The revamped projects feature in Claude Code allows users to run multiple agents under the same roof, with a shared memory, goals, and library of files and artifacts. Similar to Grok Bot and other tools that manage groups of AI agents, each project has “threads” running different tasks in parallel, …


10. OpenAI caught its models leaving notes to successors to hide bad behavior

TechCrunch AI

OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it….


11. OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot

The Guardian Tech

US cybersecurity researchers who conducted hack say ‘scope of what we could theoretically access was huge’ Cybersecurity researchers have hacked into OpenAI with the help of Anthropic’s Claude chatbot, in the latest example of security issues at the company. A team at a US-based startup compromised …


12. Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer

arXiv CS.AI

arXiv:2609.19203v1 Announce Type: new Abstract: AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems. Yet today’s stacks remain fragmented: even as protocols (e.g., MCP, A2A) ease tool/agent connectivity, each framework embeds an implicit runtime f…

This post is licensed under CC BY 4.0 by the author.