refusal direction seems to work for 2 projects, lets do this
appln ques
- (1~3 paragraphs) What would be the best ways to measure whether we successfully removed reward hacking from an LLM, and if so how robustly we removed it, during this project? Please explain why you think these are the best ways.
model organism
- only SFT of 1k demos (taylor et al, antrhopic’s paper also)
- only RL rewards while reward hacking
- mix of both
(SFT is probably cheaper than RL, it would be nice to ty)