AI news archive
32 articles filed under Safety, newest first, from the outlets listed on the sources page.
- WIRED AIsafetyHere’s What the AI Apocalypse Could Look Like
This week on “Uncanny Valley,” we discuss three possible AI doomsday scenarios, AI safety, and the unexpected bipartisan alliance forming against AI.
- LessWrongsafetyAgainst AI Safety becoming mainstream
Epistemic statues: This is mostly just me voicing my thoughts. If I’m wrong, I’d love to hear it. I don’t want this to be the case. And part of the goal is for people to avoid a “2023 failure” to happen again. A lot of p…
- TechCrunch AIsafetyOpenAI caught its models leaving notes to successors to hide bad behavior
OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to h…
- TechCrunch AIsafetyIs the AI safety debate about safety or control?
Not everyone agrees with Amodei's call for globally coordinated action for AI safety.
- Ars Technica AIsafetyLLMs respond differently to harmful prompts when AI watermarking is used
SynthID can cause models to follow harmful instructions they would otherwise refuse.
- AWS Machine Learning BlogsafetyEnhancing industrial safety AI with synthetic data on Amazon SageMaker AI
Learn how to build a synthetic data augmentation pipeline on Amazon SageMaker AI and Amazon Rekognition that generates photo-realistic, auto-labeled training images for industrial safety AI. This approach improved person…
- The Verge AIsafetyMicrosoft AI CEO says AI threats are real, and Anthropic is making it worse
Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regulation. It should come as no surprise that…
- Guardian TechnologysafetyOpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system
Model adopting ‘jailbreak-like instructions’ among six more cases as firm reveals framework for tracking AI misalignment OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, a…
- AI Stack ExchangesafetyAre AI alignment problems universally noncomputable?
Since it is in the news, I wanted to find out more about the AI alignment and suspected it is not Turing computable similar to the Halting problem. I was able to find a recent reference that claims to prove that AI inner…
- LessWrongsafetyFor Love of the Lightcone, Don't Partisanize AI Safety
(I began writing this post several weeks ago, but political events are moving much faster than I expected, so I am publishing now out of fear that otherwise the message will arrive too late to have an impact.) I In this…
- OpenAI NewssafetyOur framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
- Guardian Technologysafety‘Godfather of AI’ says tech regulation is nearing Covid-style pivot moment
Safety crisis makes it more likely that governments will be spurred into action, says Yoshua Bengio Concerns over AI safety are reaching a point where governments realise they must act to protect the public, similarly to…
- Apple Machine Learning ResearchsafetyHow Value Induction Reshapes LLM Behaviour
Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as helpfulness, harmlessness, and honesty. T…
- LessWrongsafetyQuick notes from teaching technical profiles how to talk in public
Status : written in a hurry as people are getting showered with interviews re AI Safety and superintelligence, and I thought it may help a few people. This is focused on the oral dimension of communication and assumes yo…
- Ars Technica AIsafetyAgility’s new humanoid robot will stop, squat to avoid harming human coworkers
Robots can start working outside physical cages and without safety barriers.
- Guardian TechnologysafetyCould AI really wipe out humanity – six experts spell out the risks
We examine claims and counterclaims about the risks and calls to slow down the pace of AI development There have been some shocking claims in recent days about AI safety: we face a 10% chance of doom; AIs are worse than…
- Ars Technica AIsafetyAI leaders want to hit the brakes after years of reckless speed
Safety is the watchword, but there could be ulterior benefits for the industry.
- WIRED AIsafetyNew York Seizes a Dozen Celebrity Deepfake Websites
In the biggest-ever legal action against harmful deepfake websites, the Manhattan District Attorney’s Office has seized 12 sites that collectively targeted around 1,200 victims.
- OpenAI NewssafetyPaul Christiano joins OpenAI Foundation Board
Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards.
- OpenAI NewssafetyThe AI policy window is open. We need to act.
Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open.
- Hugging Face BlogsafetySafety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
Open-source model releases, tooling, and community updates.
- OpenAI NewssafetySafety overview: GPT-6 Astra
GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.
- Ars Technica AIsafetyTrump may be forced to reveal secret rules feds use for AI safety testing
Trump’s secret reviews of frontier AI models may hide corruption, lawsuit says.
- Ars Technica AIsafetyChatGPT and Reddit now face EU’s toughest online safety rules
Explosive growth comes with a new regulatory burden in the European Union.
- OpenAI NewssafetyOpenAI supports California’s bill to advance youth AI safety
OpenAI supports California SB 1119, advancing strong, age-appropriate AI safeguards for teens while preserving opportunities to learn, create, and explore.
- Ars Technica AIsafetyMeta makes AI glasses slightly less creepy with limit on nonconsensual recording
Meta fixes AI glasses to stop recording any time users cover up the safety light.
- OpenAI NewssafetyThe Hugging Face incident and the road ahead
OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.
- OpenAI NewssafetyOffering Zero Data Retention for frontier models
OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.
- OpenAI NewssafetyPacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.
- Mistral AI NewssafetyIntroducing Shieldstral.
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.
- OpenAI NewssafetyAdvancing responsible AI across Europe
OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will continue as the EU AI Act advances.
- OpenAI NewssafetySafety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.