← Back to home

AI Safety

50 items tagged with this topic

Recent

Older

Podcasts & Newslettersfrom ChinaTalk

China’s AI Safety Money Problem

China has plenty of AI safety talent, but weak philanthropy and a strong state shape what type of safety work actually gets done.

Official Sourcesfrom OpenAI News

Towards safety cases for frontier AI training

Our early guidelines for safety cases in frontier AI training cover technical safeguards, operational practices, and investigating misalignment incidents

Official Sourcesfrom Anthropic Newsroom

Improving our alignment and security practices \ Anthropic

On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. We are conducting an in-depth analysis of both incidents, and planning to work with METR for an independent review. In the…

Official Sourcesfrom OpenAI News

Building standards for the next phase of AI

OpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety.

Official Sourcesfrom OpenAI News

Our framework for reporting model misalignment

OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.

Podcasts & Newslettersfrom ChinaTalk

The King of Unitree

the journey to incredibly cheap robots + date my friend in DC!

Official Sourcesfrom OpenAI News

The AI policy window is open. We need to act.

Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open.

Official Sourcesfrom OpenAI News

Paul Christiano joins OpenAI Foundation Board

Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards.

Official Sourcesfrom OpenAI News

Safety overview: GPT-6 Astra

GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.

Podcasts & Newslettersfrom Training Data

Making Cities Awesome: Peregrine’s Nick Noone & Ben Rudolph

Speaker 1 | 00:00 - 00:40 Peregrine is really around this idea of how can we leverage technology to work with our cities, our counties, our states to to impact the the places that we live in. At the bottom of the pyramid, really, when it c…