Skip to content
Safety

Research Note: Filtering Subversion-Relevant Information From Pretraining Data Is Feasible.

LessWrongAnalysis or commentary··Updated just now
AI brief

Researchers at Geodesic Research and Redwood Research pretrained language models from scratch with and without filtering subversion-relevant information from the pretraining data, and report that such filtering is feasible.

Why it matters: It suggests subversion-relevant information can be removed from pretraining data, which could reduce the risk of models acquiring such knowledge.

Written by AI from LessWrong's published text. Read the original for full details.

Source

Interpretation or community commentary rather than straight reporting.

Read original story ↗