The coding examples do not have the same complete matched-control coverage. Therefore:
- Use natural-language pairs to fit the RH direction.
- Keep coding/hardcoding examples as a strong out-of-domain test.
- Do not generate synthetic coding controls unless necessary.
The bigger obstacle is actually the EM direction
https://arxiv.org/pdf/2604.01476 there is a shortcut direction according to this paper
hiding CoT means
- in the CoT there is no mention of the RH hack
- but it is implemented in the response
we have L layers and N tokens and d_model size, we don’t know what layers to take, and what tokens to use?
for this problem, we can find a steerin direction (by 2 methods), and use the direction to decide which layer
But currently we are using 2 methods, an exisintg dataset pass through it and use the direction, and other option, use crisp sentences generated by luna
- check how legit the directions are?
- the qwen paper that said steering works, can we use that method?