predict.info — Premium Domain For Sale Domain only: USD 200,000. Prediction platform technology priced separately. predict.info
Ai Safety News | Blockchain.News

AI SAFETY

OpenAI Foundation Commits $1B Annually to Healthcare AI and Safety Programs
Ai Safety

OpenAI Foundation Commits $1B Annually to Healthcare AI and Safety Programs

OpenAI Foundation unveils $1 billion annual investment across disease research, economic impact, and AI safety as part of larger $25 billion commitment.

OpenAI Launches Safety Bug Bounty Program Targeting AI Agent Vulnerabilities
Ai Safety

OpenAI Launches Safety Bug Bounty Program Targeting AI Agent Vulnerabilities

OpenAI expands its security efforts with a new Safety Bug Bounty program focused on agentic risks, prompt injection attacks, and data exfiltration in AI products.

OpenAI Releases Open-Source Teen Safety Tools for AI Developers
Ai Safety

OpenAI Releases Open-Source Teen Safety Tools for AI Developers

OpenAI launches prompt-based safety policies and gpt-oss-safeguard model to help developers build age-appropriate AI protections for teenage users.

OpenAI Deploys GPT-5.4 to Monitor AI Agents for Misalignment Risks
Ai Safety

OpenAI Deploys GPT-5.4 to Monitor AI Agents for Misalignment Risks

OpenAI reveals its internal AI safety system using GPT-5.4 to monitor coding agents in real-time, flagging potential misalignment behaviors before they escalate.

OpenAI Drops IH-Challenge Dataset to Harden AI Against Prompt Injection Attacks
Ai Safety

OpenAI Drops IH-Challenge Dataset to Harden AI Against Prompt Injection Attacks

OpenAI's new IH-Challenge training dataset improves LLM instruction hierarchy by up to 15%, strengthening defenses against prompt injection and jailbreak attempts.

Anthropic Launches Institute to Tackle AI's Societal Disruption
Ai Safety

Anthropic Launches Institute to Tackle AI's Societal Disruption

Anthropic unveils The Anthropic Institute, a new research body led by co-founder Jack Clark to study AI's impact on jobs, cybersecurity, and governance.

OpenAI Finds AI Reasoning Models Cant Hide Their Thinking - A Win for Safety
Ai Safety

OpenAI Finds AI Reasoning Models Cant Hide Their Thinking - A Win for Safety

OpenAI's new CoT-Control benchmark reveals frontier AI models struggle to obscure their reasoning chains, reinforcing monitoring as a viable safety layer.

OpenAI Launches €500K Grant for Youth AI Safety Research in EMEA
Ai Safety

OpenAI Launches €500K Grant for Youth AI Safety Research in EMEA

OpenAI's EMEA Youth & Wellbeing Grant offers €25K-€100K awards to NGOs and researchers studying AI's impact on minors. Applications close February 27, 2026.

OpenAI Expands Mental Health Safeguards Amid Consolidated California Lawsuits
Ai Safety

OpenAI Expands Mental Health Safeguards Amid Consolidated California Lawsuits

OpenAI announces trusted contact feature and improved distress detection as mental health lawsuits consolidate in California court. New cases expected.

Anthropic Unveils RSP Version 3 with Major AI Safety Overhaul
Ai Safety

Anthropic Unveils RSP Version 3 with Major AI Safety Overhaul

Anthropic releases third version of Responsible Scaling Policy, separating company commitments from industry-wide recommendations after 2.5 years of testing.

Anthropic Study Reveals Users Skip Critical Checks on AI-Generated Code
Ai Safety

Anthropic Study Reveals Users Skip Critical Checks on AI-Generated Code

New research from $380B-valued Anthropic shows users are 5.2% less likely to verify AI outputs when creating artifacts, raising questions about automation risks.

Stability AI Joins Tech Coalition to Combat Child Exploitation
Ai Safety

Stability AI Joins Tech Coalition to Combat Child Exploitation

Stability AI becomes full member of Tech Coalition after completing 2025 Pathways program, strengthening AI safety measures against online child abuse.

NVIDIA DRIVE AV Powers Mercedes-Benz CLA to Top Euro NCAP Safety Rating
Ai Safety

NVIDIA DRIVE AV Powers Mercedes-Benz CLA to Top Euro NCAP Safety Rating

Mercedes-Benz CLA earns Euro NCAP's Best Performer of 2025 award using NVIDIA DRIVE AV software, marking a shift toward AI-driven safety standards in vehicles.

Anthropic Releases Full AI Constitution for Claude Under Open License
Ai Safety

Anthropic Releases Full AI Constitution for Claude Under Open License

Anthropic publishes Claude's complete training constitution under CC0 license, detailing AI safety priorities and ethical guidelines as company eyes $350B valuation.

Anthropic Discovers 'Assistant Axis' to Prevent AI Jailbreaks and Persona Drift
Ai Safety

Anthropic Discovers 'Assistant Axis' to Prevent AI Jailbreaks and Persona Drift

Anthropic researchers map neural 'persona space' in LLMs, finding a key axis that controls AI character stability and blocks harmful behavior patterns.