https://arxiv.org/pdf/2406.11717#page=30.09
it is interesting that subtracting harmful - harmless prompts gave rise to direction of “refusal” , not direction of “harm”
ideally i would imagine that subtracting “refused responses - non refused responses ” would give the direction of refusal,
I am skeptical of harmful - harmless because model may not assume refusal for many harmful prompts, and may assume refusal for many harmless prompts. But this is working because they are testing on an already aligned model, where harmful implicity has a refusal feature.