AI brief
Researchers at Geodesic Research and Redwood Research pretrained language models from scratch with and without filtering subversion-relevant information from the pretraining data, and report that such filtering is feasible.
Why it matters: It suggests subversion-relevant information can be removed from pretraining data, which could reduce the risk of models acquiring such knowledge.
Written by AI from LessWrong's published text. Read the original for full details.