AI Safety
OpenAI's Astra Hits Critical Cybersecurity Threshold With Strict Safeguards
OpenAI's Astra becomes the first AI model to meet Critical cybersecurity capability under the Preparedness Framework, prompting stringent safeguards.
Claude AI Improves Alignment Benchmarks While Preserving Capabilities
Anthropic's Claude achieved significant alignment improvements on 10 benchmarks, outperforming human researchers and maintaining model capabilities.
Anthropic Warns of Risks in Multiagent AI Systems
Anthropic highlights coordination failures, collusion, and sabotage in experiments with multiagent AI systems, raising safety concerns.
Anthropic Enhances Biology Safeguards for Claude Fable 5
Anthropic updates Claude Fable 5's biology safeguards, reducing fallbacks by 85% and expanding its utility for health and education tasks.
OpenAI Expands Teen Protections in ChatGPT for Safer AI Use
OpenAI enhances ChatGPT safeguards for teens, introducing parental controls, age-specific content filters, and new learning tools.
NVIDIA Halos OS Brings AV-Grade Safety to Robotics
NVIDIA launches Halos OS for robotics, extending AV-grade safety to industrial robots and physical AI, built on IGX Thor and Halos Core.
DeepMind Unveils AI Control Roadmap to Address Alignment Risks
DeepMind introduces a defense-in-depth AI Control Roadmap, targeting risks from misaligned advanced AI. Key implications for security and governance.
Google DeepMind Offers $10M for Multi-Agent AI Safety Research
Google DeepMind and partners launch a $10M funding call to tackle emergent risks in multi-agent AI systems. Applications close August 8, 2026.
NVIDIA Halos OS Drives Safety for L4 Robotaxis at Scale
NVIDIA's Halos OS offers a safety-certified platform for Level 4 robotaxis, addressing key challenges in autonomous vehicle deployment.
OpenAI Pushes for Global Youth AI Safety Standards at G7 Summit
OpenAI urges G7 leaders to establish a global institute for youth AI safety, aiming to standardize protections and promote opportunities.