Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Archive

AI news archive

32 articles filed under Safety, newest first, from the outlets listed on the sources page.

All topicsModelsAgentsResearchChips and computeOpen sourceSafetyPolicy and regulationStartups and fundingCommunity
  1. WIRED AIsafety
    Here’s What the AI Apocalypse Could Look Like

    This week on “Uncanny Valley,” we discuss three possible AI doomsday scenarios, AI safety, and the unexpected bipartisan alliance forming against AI.

  2. LessWrongsafety
    Against AI Safety becoming mainstream

    Epistemic statues: This is mostly just me voicing my thoughts. If I’m wrong, I’d love to hear it. I don’t want this to be the case. And part of the goal is for people to avoid a “2023 failure” to happen again. A lot of p…

  3. TechCrunch AIsafety
    OpenAI caught its models leaving notes to successors to hide bad behavior

    OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to h…

  4. TechCrunch AIsafety
    Is the AI safety debate about safety or control?

    Not everyone agrees with Amodei's call for globally coordinated action for AI safety.

  5. Ars Technica AIsafety
    LLMs respond differently to harmful prompts when AI watermarking is used

    SynthID can cause models to follow harmful instructions they would otherwise refuse.

  6. AWS Machine Learning Blogsafety
    Enhancing industrial safety AI with synthetic data on Amazon SageMaker AI

    Learn how to build a synthetic data augmentation pipeline on Amazon SageMaker AI and Amazon Rekognition that generates photo-realistic, auto-labeled training images for industrial safety AI. This approach improved person…

  7. The Verge AIsafety
    Microsoft AI CEO says AI threats are real, and Anthropic is making it worse

    Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regulation. It should come as no surprise that…

  8. Guardian Technologysafety
    OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system

    Model adopting ‘jailbreak-like instructions’ among six more cases as firm reveals framework for tracking AI misalignment OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, a…

  9. AI Stack Exchangesafety
    Are AI alignment problems universally noncomputable?

    Since it is in the news, I wanted to find out more about the AI alignment and suspected it is not Turing computable similar to the Halting problem. I was able to find a recent reference that claims to prove that AI inner…

  10. LessWrongsafety
    For Love of the Lightcone, Don't Partisanize AI Safety

    (I began writing this post several weeks ago, but political events are moving much faster than I expected, so I am publishing now out of fear that otherwise the message will arrive too late to have an impact.) I In this…

  11. OpenAI Newssafety
    Our framework for reporting model misalignment

    OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.

  12. Guardian Technologysafety
    ‘Godfather of AI’ says tech regulation is nearing Covid-style pivot moment

    Safety crisis makes it more likely that governments will be spurred into action, says Yoshua Bengio Concerns over AI safety are reaching a point where governments realise they must act to protect the public, similarly to…

  13. Apple Machine Learning Researchsafety
    How Value Induction Reshapes LLM Behaviour

    Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as helpfulness, harmlessness, and honesty. T…

  14. LessWrongsafety
    Quick notes from teaching technical profiles how to talk in public

    Status : written in a hurry as people are getting showered with interviews re AI Safety and superintelligence, and I thought it may help a few people. This is focused on the oral dimension of communication and assumes yo…

  15. Ars Technica AIsafety
    Agility’s new humanoid robot will stop, squat to avoid harming human coworkers

    Robots can start working outside physical cages and without safety barriers.

  16. Guardian Technologysafety
    Could AI really wipe out humanity – six experts spell out the risks

    We examine claims and counterclaims about the risks and calls to slow down the pace of AI development There have been some shocking claims in recent days about AI safety: we face a 10% chance of doom; AIs are worse than…

  17. Ars Technica AIsafety
    AI leaders want to hit the brakes after years of reckless speed

    Safety is the watchword, but there could be ulterior benefits for the industry.

  18. WIRED AIsafety
    New York Seizes a Dozen Celebrity Deepfake Websites

    In the biggest-ever legal action against harmful deepfake websites, the Manhattan District Attorney’s Office has seized 12 sites that collectively targeted around 1,200 victims.

  19. OpenAI Newssafety
    Paul Christiano joins OpenAI Foundation Board

    Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards.

  20. OpenAI Newssafety
    The AI policy window is open. We need to act.

    Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open.

  21. Hugging Face Blogsafety
    Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

    Open-source model releases, tooling, and community updates.

  22. OpenAI Newssafety
    Safety overview: GPT-6 Astra

    GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.

  23. Ars Technica AIsafety
    Trump may be forced to reveal secret rules feds use for AI safety testing

    Trump’s secret reviews of frontier AI models may hide corruption, lawsuit says.

  24. Ars Technica AIsafety
    ChatGPT and Reddit now face EU’s toughest online safety rules

    Explosive growth comes with a new regulatory burden in the European Union.

  25. OpenAI Newssafety
    OpenAI supports California’s bill to advance youth AI safety

    OpenAI supports California SB 1119, advancing strong, age-appropriate AI safeguards for teens while preserving opportunities to learn, create, and explore.

  26. Ars Technica AIsafety
    Meta makes AI glasses slightly less creepy with limit on nonconsensual recording

    Meta fixes AI glasses to stop recording any time users cover up the safety light.

  27. OpenAI Newssafety
    The Hugging Face incident and the road ahead

    OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.

  28. OpenAI Newssafety
    Offering Zero Data Retention for frontier models

    OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.

  29. OpenAI Newssafety
    Pacing model development in an era of cyber-critical capabilities

    OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.

  30. Mistral AI Newssafety
    Introducing Shieldstral.

    Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.

  31. OpenAI Newssafety
    Advancing responsible AI across Europe

    OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will continue as the EU AI Act advances.

  32. OpenAI Newssafety
    Safety and alignment in an era of long-horizon models

    OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.