AI brief
A post describes Jev, a new model format that outputs certainties over a defined set of options instead of text, and applies it to monitoring harmful thought traces.
Why it matters: The post claims this classification-based monitoring approach is faster and cheaper, which could make monitoring harmful reasoning traces more practical.
Written by AI from LessWrong's published text. Read the original for full details.