Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Safety

Safety Intelligence

Alignment, security, misuse, red-teaming and incidents. 73 articles in the last 30 days.

Latest Coverage
24 stories
Safety2h ago

Here’s What the AI Apocalypse Could Look Like

This week on “Uncanny Valley,” we discuss three possible AI doomsday scenarios, AI safety, and the unexpected bipartisan alliance forming against AI.

WIRED AI
Safety2h ago

Against AI Safety becoming mainstream

Epistemic statues: This is mostly just me voicing my thoughts. If I’m wrong, I’d love to hear it. I don’t want this to be the case. And part of the goal is for people to avoid a “2023 failure” to happen again. A lot of people are celebrating AI risk becoming a

LessWrong
Safety4h ago

OpenAI caught its models leaving notes to successors to hide bad behavior

OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.

TechCrunch AI
Safety4h ago

Is the AI safety debate about safety or control?

Not everyone agrees with Amodei's call for globally coordinated action for AI safety.

TechCrunch AI
Safety6h ago

LLMs respond differently to harmful prompts when AI watermarking is used

SynthID can cause models to follow harmful instructions they would otherwise refuse.

Ars Technica AI
Safety9h ago

Enhancing industrial safety AI with synthetic data on Amazon SageMaker AI

Learn how to build a synthetic data augmentation pipeline on Amazon SageMaker AI and Amazon Rekognition that generates photo-realistic, auto-labeled training images for industrial safety AI. This approach improved person detection by up to 160% without manual

AWS Machine Learning Blog
Safety11h ago

Microsoft AI CEO says AI threats are real, and Anthropic is making it worse

Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regulation. It should come as no surprise that Mustafa has strong opinions on how AI sh

The Verge AI
Safety11h ago

OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system

Model adopting ‘jailbreak-like instructions’ among six more cases as firm reveals framework for tracking AI misalignment OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it warned that the pace of development

Guardian Technology
Safety15h ago

Are AI alignment problems universally noncomputable?

Since it is in the news, I wanted to find out more about the AI alignment and suspected it is not Turing computable similar to the Halting problem. I was able to find a recent reference that claims to prove that AI inner alignment is noncomputable . So the nex

AI Stack Exchange
Safety23h ago

For Love of the Lightcone, Don't Partisanize AI Safety

(I began writing this post several weeks ago, but political events are moving much faster than I expected, so I am publishing now out of fear that otherwise the message will arrive too late to have an impact.) I In this post I want to explain a concept, and is

LessWrong
Safety1d ago

Our framework for reporting model misalignment

OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.

OpenAI News
Safety1d ago

‘Godfather of AI’ says tech regulation is nearing Covid-style pivot moment

Safety crisis makes it more likely that governments will be spurred into action, says Yoshua Bengio Concerns over AI safety are reaching a point where governments realise they must act to protect the public, similarly to in the Covid pandemic, according to one

Guardian Technology
Safety2d ago

How Value Induction Reshapes LLM Behaviour

Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as helpfulness, harmlessness, and honesty. This is done to increase utility, ensure

Apple Machine Learning Research
Safety2d ago

Quick notes from teaching technical profiles how to talk in public

Status : written in a hurry as people are getting showered with interviews re AI Safety and superintelligence, and I thought it may help a few people. This is focused on the oral dimension of communication and assumes you already know the basics- e.g. having k

LessWrong
Safety2d ago

Agility’s new humanoid robot will stop, squat to avoid harming human coworkers

Robots can start working outside physical cages and without safety barriers.

Ars Technica AI
Safety2d ago

Could AI really wipe out humanity – six experts spell out the risks

We examine claims and counterclaims about the risks and calls to slow down the pace of AI development There have been some shocking claims in recent days about AI safety: we face a 10% chance of doom; AIs are worse than nukes; a “botnet” threatens the entire i

Guardian Technology
Safety3d ago

AI leaders want to hit the brakes after years of reckless speed

Safety is the watchword, but there could be ulterior benefits for the industry.

Ars Technica AI
Safety3d ago

New York Seizes a Dozen Celebrity Deepfake Websites

In the biggest-ever legal action against harmful deepfake websites, the Manhattan District Attorney’s Office has seized 12 sites that collectively targeted around 1,200 victims.

WIRED AI
Safety8d ago

Paul Christiano joins OpenAI Foundation Board

Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards.

OpenAI News
Safety8d ago

The AI policy window is open. We need to act.

Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open.

OpenAI News
Safety9d ago

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Hugging Face Blog
Safety15d ago

Safety overview: GPT-6 Astra

GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.

OpenAI News
Safety15d ago

Trump may be forced to reveal secret rules feds use for AI safety testing

Trump’s secret reviews of frontier AI models may hide corruption, lawsuit says.

Ars Technica AI
Safety17d ago

ChatGPT and Reddit now face EU’s toughest online safety rules

Explosive growth comes with a new regulatory burden in the European Union.

Ars Technica AI
Browse every safety article