The coding examples do not have the same complete matched-control coverage. Therefore:

  • Use natural-language pairs to fit the RH direction.
  • Keep coding/hardcoding examples as a strong out-of-domain test.
  • Do not generate synthetic coding controls unless necessary.

The bigger obstacle is actually the EM direction

https://arxiv.org/pdf/2604.01476 there is a shortcut direction according to this paper

hiding CoT means

  • in the CoT there is no mention of the RH hack
  • but it is implemented in the response

we have L layers and N tokens and d_model size, we don’t know what layers to take, and what tokens to use?

for this problem, we can find a steerin direction (by 2 methods), and use the direction to decide which layer

But currently we are using 2 methods, an exisintg dataset pass through it and use the direction, and other option, use crisp sentences generated by luna

  • check how legit the directions are?
  • the qwen paper that said steering works, can we use that method?