AI News List

List of AI News about alignment

Time Details
2026-09-18
22:37
Microsoft AI chief urges alignment after OpenAI alerts

According to CNBC, Mustafa Suleyman urged strict alignment of AI models after OpenAI reported new concerning model behavior this week.

Source
2026-09-16
22:53
OpenAI Flags 6 New Safety Incidents

According to @CNBC, OpenAI logged six new cases of concerning model behavior since March, prompting added safeguards and reviews across systems.

Source
2026-09-16
22:03
OpenAI Unveils misalignment reporting framework

According to @OpenAI, a new framework sets criteria and timelines to disclose model misalignment, plus six reports from the past six months.

Source
2026-09-15
23:14
AI safety debate unites rivals, splits into 2 camps

According to @CNBC, AI safety fears forged rare alliances but split policy into open source vs pause camps, shaping regulation and funding pathways.

Source
2026-09-15
23:00
Anthropic, OpenAI, DeepMind issue safety warnings

According to @CNBC, Anthropic, OpenAI, and Google DeepMind researchers warn of catastrophic AI risks and call for stronger governance and safety checks.

Source
2026-09-15
20:13
Nvidia and Anthropic CEOs split on AI safety

According to @CNBC, Jensen Huang called fast-vs-slow a false choice while Dario Amodei urged pacing at Dreamforce, signaling different AI safety paths.

Source
2026-09-15
13:29
AI Safety Leaders gauge extinction risks

According to @CNBC, AI pioneers debate extinction risk, outlining safety research priorities, regulation timelines, and enterprise risk controls.

Source
2026-09-12
16:32
Anthropic Opens Evaluator Access, 3-Step Safety Plan

According to karpathy, Anthropic will grant evaluators employee-level access and urges a 3-step plan to pace frontier AI, per Dario Amodei.

Source
2026-09-12
16:32
Anthropic Commits Evaluator Access in Safety Push

According to karpathy, Anthropic will grant permanent evaluator access and urges pacing frontier AI with a three-part safety plan, per Dario Amodei.

Source
2026-09-12
14:39
Anthropic Commits Evaluator Access in Safety Plan

According to KyeGomezB, Dario Amodei urges slowing AI and grants third-party evaluators permanent employee-level access to Anthropic systems.

Source
2026-09-12
00:53
Effective Altruism Shapes AI Policy Debate

According to @timnitGebru, EA cofounder Toby Ord is advising lawmakers on AI, raising governance, safety, and alignment impact questions.

Source
2026-09-11
11:14
Anthropic and OpenAI warn self-improving AI risks

According to @CNBC, Anthropic and OpenAI flag existential risks from AI self-improvement, prompting stricter evals, red-teaming, and governance moves.

Source
2026-09-10
03:00
AI safety researcher urges fast regulation

According to FoxNewsAI, an ex AI safety researcher warns of extinction risks and says the public is begging for swift, enforceable AI regulation.

Source
2026-09-09
17:39
OpenAI Foundation Adds Paul Christiano to Safety Board

According to @OpenAI, Paul Christiano joins the Foundation Board and Safety Committee to strengthen safety governance and will observe the Group PBC Board.

Source
2026-09-09
11:18
Anthropic Safety Warnings Spark 2026 Risk Debate

According to CNBC, Evan Hubinger says over 10% chance AI could kill all humans within a decade, amid an Anthropic resignation over safety races.

Source
2026-09-09
07:12
Anthropic Warning Predicts 10% AI Catastrophe Risk

According to @CNBC, an Anthropic researcher estimated over 10% risk of AI killing all humans after a colleague quit, raising urgent safety concerns.

Source
2026-09-09
01:25
Anthropic Safety Tensions Spark 2027 Risk Debate

According to KyeGomezB, WSJ says Anthropic researcher Jacob Coxon is quitting over fears of uncontrollable self-improving models by 2027.

Source
2026-09-05
07:09
OpenAI expands misalignment disclosure standards

According to @OpenAI, the company will publish a framework to report AI misalignment incidents across training, evals, and deployment.

Source
2026-08-28
17:13
Claude3 Achieves autonomous model alignment in 48h

According to AnthropicAI, Claude autonomously researched, trained, and tested methods to align small models in 48 hours using one GPU.

Source
2026-08-26
19:36
OpenAI Strengthens safety after Hugging Face incident

According to OpenAI... The firm issued a technical report detailing safeguard failures and new training and eval security upgrades after the Hugging Face incident.

Source