Skip to content
Safety

Why I'm scared of RL.

LessWrongAnalysis or commentary··Updated just now
AI brief

A post on LessWrong argues that reinforcement learning is a black-box source of agency, which raises classic misalignment concerns, particularly when compared with agency created through scaffolding.

Why it matters: It frames reinforcement learning as a source of misalignment risk relative to alternative approaches to building agentic systems.

Written by AI from LessWrong's published text. Read the original for full details.

Source

Interpretation or community commentary rather than straight reporting.

Read original story ↗