AI Safety

OpenAI's Astra Hits Critical Cybersecurity Threshold With Strict Safeguards
AI Safety

OpenAI's Astra Hits Critical Cybersecurity Threshold With Strict Safeguards

OpenAI's Astra becomes the first AI model to meet Critical cybersecurity capability under the Preparedness Framework, prompting stringent safeguards.

Claude AI Improves Alignment Benchmarks While Preserving Capabilities
AI Safety

Claude AI Improves Alignment Benchmarks While Preserving Capabilities

Anthropic's Claude achieved significant alignment improvements on 10 benchmarks, outperforming human researchers and maintaining model capabilities.

Anthropic Warns of Risks in Multiagent AI Systems
AI Safety

Anthropic Warns of Risks in Multiagent AI Systems

Anthropic highlights coordination failures, collusion, and sabotage in experiments with multiagent AI systems, raising safety concerns.

Anthropic Enhances Biology Safeguards for Claude Fable 5
AI Safety

Anthropic Enhances Biology Safeguards for Claude Fable 5

Anthropic updates Claude Fable 5's biology safeguards, reducing fallbacks by 85% and expanding its utility for health and education tasks.

OpenAI Expands Teen Protections in ChatGPT for Safer AI Use
AI Safety

OpenAI Expands Teen Protections in ChatGPT for Safer AI Use

OpenAI enhances ChatGPT safeguards for teens, introducing parental controls, age-specific content filters, and new learning tools.

NVIDIA Halos OS Brings AV-Grade Safety to Robotics
AI Safety

NVIDIA Halos OS Brings AV-Grade Safety to Robotics

NVIDIA launches Halos OS for robotics, extending AV-grade safety to industrial robots and physical AI, built on IGX Thor and Halos Core.

DeepMind Unveils AI Control Roadmap to Address Alignment Risks
AI Safety

DeepMind Unveils AI Control Roadmap to Address Alignment Risks

DeepMind introduces a defense-in-depth AI Control Roadmap, targeting risks from misaligned advanced AI. Key implications for security and governance.

Google DeepMind Offers $10M for Multi-Agent AI Safety Research
AI Safety

Google DeepMind Offers $10M for Multi-Agent AI Safety Research

Google DeepMind and partners launch a $10M funding call to tackle emergent risks in multi-agent AI systems. Applications close August 8, 2026.

NVIDIA Halos OS Drives Safety for L4 Robotaxis at Scale
AI Safety

NVIDIA Halos OS Drives Safety for L4 Robotaxis at Scale

NVIDIA's Halos OS offers a safety-certified platform for Level 4 robotaxis, addressing key challenges in autonomous vehicle deployment.

OpenAI Pushes for Global Youth AI Safety Standards at G7 Summit
AI Safety

OpenAI Pushes for Global Youth AI Safety Standards at G7 Summit

OpenAI urges G7 leaders to establish a global institute for youth AI safety, aiming to standardize protections and promote opportunities.

Trending topics