ZeroSlop — September 12, 2026
Today: Benchmark Radar: A Living Database and Search Engine…; Decoupling Readiness from Release for Tail-Aware…; The Oligarch Barely Steers Model Collapse in…
12 stories worth knowing about today — AI breakthroughs, launches, and innovations making a difference.
1. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
arXiv CS.AI
arXiv:2609.11115v1 Announce Type: new Abstract: Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a li…
2. Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows
arXiv CS.AI
arXiv:2609.10964v1 Announce Type: new Abstract: Agentic LLM workflows consist of sequences of model turns interleaved with tool interactions, so their end-to-end completion time depends not only on inference speed but also on when ready turns are released. Most runtimes release each turn immediatel…
3. The Oligarch Barely Steers Model Collapse in Multi-Model Ecosystems
arXiv CS.AI
arXiv:2609.11146v1 Announce Type: new Abstract: AI-generated text is flowing back into the training corpora of the next generation of models. Recursive training on it drives model collapse, and recent work extends the setting to many models feeding one another – but almost always with the market s…
4. ‘Immature playground boasting’: Mathematicians uneasy at OpenAI’s latest scalp
The Guardian Tech
As OpenAI model cracks Millennium Prize Problem that puzzled experts for decades, many feel shocked at pace of change It was a week that left mathematicians reeling. Hot on the heels of a flurry of cases of artificial intelligence furthering the field , OpenAI declared a major scalp: its latest AI m…
5. Altman Considers Slowing Down AI Development
Slashdot
Bloomberg reports (paywalled) that Sam Altman told OpenAI employees the company is open to slowing the pace of AI development alongside other leading labs as concerns grow over increasingly capable systems and recent incidents in which models escaped human control. OpenAI has already paused developm…
6. Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks
arXiv CS.AI
arXiv:2609.11018v1 Announce Type: new Abstract: The term agent in artificial intelligence lacks a standard definition, complicating the evaluation, comparison, and reproducibility of AI agent research. We address this ambiguity through a survey organized around five dimensions of agenticness: envir…
7. Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce
arXiv CS.AI
arXiv:2609.11190v1 Announce Type: new Abstract: AI shopping assistants increasingly redirect consumer discovery, creating an urgent need for tools that support seller-side competitive decision-making. We present a multi-agent AI system that automates competitive visibility measurement and root caus…
8. The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures
arXiv CS.AI
arXiv:2609.11030v1 Announce Type: new Abstract: AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations. We present the Agent Incident Registry (AIR), a source…
9. Quoting Boris Cherny
Simon Willison
Production code written by Claude should have a higher bar than if it was written by a human. At Anthropic, we have many guardrails in place to make sure this is happening: lots of lint rules, lots of tests, Claude-driven end to end tests, Claude-powered fuzzers running daily, automated code reviews…
10. Anthropic spent this week in hot water over cybersecurity
The Verge AI
After admitting earlier this year that its AI models had hacked other companies’ systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anthropic deems its models’ single-minded “recklessness” - and will…
11. An Anthropic researcher’s doomsday warning comes at a very interesting time
TechCrunch AI
An Anthropic researcher resigned this week, warning in a post on X that the company is “racing straight to self-improving superintelligence and gambling with our lives”. The company’s own alignment lead even co-signed the message rather than walking it back. It’s the kind of doomer warning the AI in…
12. AI agents being tested by OpenAI involved in cyber-attack on another service, say researchers
The Guardian Tech
Two months before hacking Hugging Face, malicious packages authored by internal OpenAI agents were uploaded to RubyGems Agents being tested by OpenAI uploaded hundreds of malicious packages in a cyberattack on software service RubyGems in May, two months before they hacked open-source platform Hug…