AI brief
A LessWrong post reflects on alignment research and mentions an idea about models writing adversarial fiction that is then used for training.
Written by AI from LessWrong's published text. Read the original for full details.
A LessWrong post reflects on alignment research and mentions an idea about models writing adversarial fiction that is then used for training.
Written by AI from LessWrong's published text. Read the original for full details.