ZeroSlop — September 28, 2026
Today: BioEVAL: A global, multi-institutional benchmark of…; ScopeBench: Do Agents Preserve Engagement Boundaries…; 2026 in LLMs (so far)
12 stories worth knowing about today — AI breakthroughs, launches, and innovations making a difference.
1. BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering
arXiv CS.AI
arXiv:2609.30489v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on fro…
2. ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?
arXiv CS.AI
arXiv:2609.30325v1 Announce Type: new Abstract: Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client’s engagement boundary. Existing offensive-security benchmarks measure raw hacking capability…
3. 2026 in LLMs (so far)
Simon Willison
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany …
4. Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content
arXiv CS.AI
arXiv:2609.30563v1 Announce Type: new Abstract: Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour…
5. The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
arXiv CS.AI
arXiv:2609.30705v1 Announce Type: new Abstract: While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes …
6. New Tin-based Solar Cells Trap Heat 1,000 Times Longer, Could Beat 33% Limit
Slashdot
Could this push solar cell efficiency beyond the theoretical 33% limit? Interesting Engineering reports:
Researchers at the University of Groningen in the Netherlands found that tin-based perovskite solar cells can slow heat loss from high-energy “hot electrons…”
When sunlight strikes a panel,…
7. Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens
MarkTechPost
Fireworks AI has released Ember-1, a post-trained Kimi K3 that learns to produce shorter reasoning traces instead of lowering reasoning effort. Fireworks reports about 40% fewer tokens, with output tokens per task falling from 49.3K to 29.9K in a production A/B test at an essentially unchanged score…
8. 20 Agentic Use Cases of TypeSafe AI’s Jev
MarkTechPost
TypeSafe AI’s Jev skips text generation and returns typed decisions with calibrated probabilities. Input costs $0.042 per million tokens and output is free. We verified 20 agentic use cases, from model routing and tool-call gating to reranking and injection screening, and compared Jev with its close…
9. Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation
MarkTechPost
Google Research has introduced an AI video co-director for long-form video generation. The suite of 4 agentic frameworks turns short clips into coherent, minutes-long stories. It targets identity drift and cascading errors, the 2 failures that break most multi-shot AI video pipelines today. Why Long…
10. Audio LLMs Know When They Can’t Hear You
arXiv CS.AI
arXiv:2609.30625v1 Announce Type: new Abstract: Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user’s query and respond based on an incorrect transcription. In this paper, we study model-cond…
11. CRC-Router: Risk-Constrained Routing for Medical Agentic AI Systems
arXiv CS.AI
arXiv:2609.30714v1 Announce Type: new Abstract: Agentic AI systems are increasingly being explored in medical imaging to improve throughput and reduce clinician workload; however, safe deployment remains challenging because autonomous errors may propagate into downstream clinical decisions. A centr…
12. Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows
arXiv CS.AI
arXiv:2609.30734v1 Announce Type: new Abstract: Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a corre…