AI Safety
50 items tagged with this topic
Recent
2026 Usage Policy update \ Anthropic
We’re publishing a new version of our Usage Policy. In this post, we summarize the changes we’ve made.
Introducing the Anthropic Cyber Mission \ Anthropic
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Older
[AINews] not much happened today
a quiet day.
China’s AI Safety Money Problem
China has plenty of AI safety talent, but weak philanthropy and a strong state shape what type of safety work actually gets done.
Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to H…
Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to Haiku 4.5, with far fewer instances of misaligned behavior.
Logan Wright on Broken China
.5% growth...how did it get so bad?
I am very uncomfortable about people trying to ascribe religious force or a surrender of human…
I am very uncomfortable about people trying to ascribe religious force or a surrender of human judgment to AI models, and think it is a real safety issue.
It’s entirely plausible that the AI industry can deal with safety and security through a set of…
It’s entirely plausible that the AI industry can deal with safety and security through a set of shared standards and practices for the foreseeable future. At some point in capability progress there will inevitably be gr…
Towards safety cases for frontier AI training
Our early guidelines for safety cases in frontier AI training cover technical safeguards, operational practices, and investigating misalignment incidents
Xi Descends on DC
With Julian Gewirtz
Sam Altman’s remarks at the United Nations Security Council
OpenAI CEO Sam Altman discusses AI safety, human control, and international cooperation in remarks to the United Nations Security Council.
Introducing the Life Sciences Verification Program \ Anthropic
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Developing Enterprise Frontier Safeguards with our customers \ Anthropic
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Improving our alignment and security practices \ Anthropic
On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. We are conducting an in-depth analysis of both incidents, and planning to work with METR for an independent review. In the…
Safety and alignment of "models" is the raging topic today. But the original "safety" debate fo…
Safety and alignment of "models" is the raging topic today. But the original "safety" debate for AI/MachineLearning was actually in autonomous vehicles! My first Waymo ride truly felt like a religious experience. It was…
Priorities and principles for effective third party assessments
OpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.
Building standards for the next phase of AI
OpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety.
Introducing the Australian Youth Safety Blueprint
OpenAI introduces the Australian Youth Safety Blueprint, a six-pillar roadmap for safer AI experiences that protect and empower young people.
Be careful what you wish for This is why we need alignment to humankind vs any other goal https…
Be careful what you wish for This is why we need alignment to humankind vs any other goal https://t.co/E3dxymT1aV
This is the right way to solve alignment and safety. @GoodfireAI is the leading non-frontier-la…
This is the right way to solve alignment and safety. @GoodfireAI is the leading non-frontier-lab company that is working on this. This is generationally important. https://t.co/BkufCYX1Qb
Safety and security are * features * of your AI product and model - not guardrails that need to…
Safety and security are * features * of your AI product and model - not guardrails that need to be imposed on you from the outside. https://t.co/jtN3rzdZV8 https://t.co/J8dM9vHsoN
We're seeing extraordinary results from @typesafeai. Default mode in 𝚏𝚡 is auto, with a safet…
We're seeing extraordinary results from @typesafeai. Default mode in 𝚏𝚡 is auto, with a safety reviewer analyzing every command. That reviewer runs on GPT Luna today. Jev is up to 18x faster (p95) *and* more accurate.…
Incredibly exciting that there are entire universes of AI innovation that still exist that were…
Incredibly exciting that there are entire universes of AI innovation that still exist that weren’t even on most of our radars. Being able to process information insanely quickly, at crazy low costs, with high levels of…
Expanding our support for scientists \ Anthropic
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
[AINews] AEF-1 standard emerges for Third Party Evaluators, as Xai, OpenAI, and Anthropic all cosign
Pacing gathers pace.
The AI naming curse strikes again. “AI Safety” firm made AI unsafe. “Effective Altruists” are b…
The AI naming curse strikes again. “AI Safety” firm made AI unsafe. “Effective Altruists” are both ineffective and enabling criminal activity. “Irregular” is regularly incompetent. https://t.co/vaOgS5lyKf
There’s a massive chasm between the power of AI models and the ultimate workflows that enterpri…
There’s a massive chasm between the power of AI models and the ultimate workflows that enterprises are trying to automate. This gap is the opportunity for the applied AI layer to fill. You need to connect the intelligen…
Our framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
The King of Unitree
the journey to incredibly cheap robots + date my friend in DC!
Can Taiwan Build Drones?
procurement, politics, and the arsenal
There are two ways AI progress could go very badly and that we must avoid. First, we could lose…
There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve peopl…
The world deserves confidence that American companies developing increasingly capable AI will a…
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there i…
[AINews] not much happened today
a quiet day
US-China Biorisk Cooperation: Yes, It’s Possible
The U.S. AI CEOs all want to screen dangerous biological orders. Coming to an agreement with China is the next step.
Before we solve AI alignment, we have a pretty serious human alignment problem. It is jarring t…
Before we solve AI alignment, we have a pretty serious human alignment problem. It is jarring to see the bad-faith reaction to Dario's piece from so many sides. I don't agree with every point of his essay but do believe…
How do we encourage true innovation in CA? How do we enable a two-party state? How should techn…
How do we encourage true innovation in CA? How do we enable a two-party state? How should technology be regulated? We're hosting conversations about tech policy with decision makers at @spc this fall. @SteveHiltonx on S…
Welcome, Paul. Grateful you are doing this, and all you have done for AI safety. Excited to wor…
Welcome, Paul. Grateful you are doing this, and all you have done for AI safety. Excited to work together again. https://t.co/k6NUqN6Lsd
How Trump and Xi Can Do AI Safety
it may be possible!
The AI policy window is open. We need to act.
Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open.
Paul Christiano joins OpenAI Foundation Board
Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards.
Funding grants for new research into AI and teen development
Apply now for OpenAI’s $5 million grant program supporting independent research on how generative AI affects teen development, well-being, and safety.
Safety overview: GPT-6 Astra
GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.
Over the summer, we have been sprinting on safety priorities; it's more important than ever for…
Over the summer, we have been sprinting on safety priorities; it's more important than ever for capabilities and safeguards to advance together. We have more to do but have made a lot of progress. We are also going to b…
Making Cities Awesome: Peregrine’s Nick Noone & Ben Rudolph
Speaker 1 | 00:00 - 00:40 Peregrine is really around this idea of how can we leverage technology to work with our cities, our counties, our states to to impact the the places that we live in. At the bottom of the pyramid, really, when it c…
How law firm Gilbert + Tobin governs and scales AI with OpenAI
See how Gilbert + Tobin combines CEO-led commitment, rigorous governance, and human accountability to scale ChatGPT Enterprise and Codex across the firm.
OpenAI supports California’s bill to advance youth AI safety
OpenAI supports California SB 1119, advancing strong, age-appropriate AI safeguards for teens while preserving opportunities to learn, create, and explore.
The lesson from the Hugging Face Incident should be that RL with verifiable rewards is an incre…
The lesson from the Hugging Face Incident should be that RL with verifiable rewards is an incredibly powerful optimization algorithm that will produce increasingly weird and surprising behavior from LLMs. The obvious mi…
California high speed rail is pretty obviously hijacked by public sector unions bilking the hap…
California high speed rail is pretty obviously hijacked by public sector unions bilking the hapless state and NIMBY landowners abusing CEQA These are policy disasters and can be fixed with better policies but so far Cal…
How AI Becomes a Political Crisis
The political economy of AI — things are getting weird, it seems.