AI brief
WorkspaceBench is a new set of evaluations for measuring how well an activation-to-text tool can read the intermediate variables of a model during a forward pass, described as its global workspace. The benchmark contains 3,356 questions across 27 categories.
Why it matters: It offers a way to assess interpretability tools that translate model activations into text.
Written by AI from LessWrong's published text. Read the original for full details.