TODO

TODO

useful todo

  1. read refusal paper refusal has a single direction
  2. apply to represnetation diagnostics for safety and orthogonalization against reward hacking
  3. read fine tuning , sae scalpels stethescope
  4. finish jspace
  5. write something in finetuning
  6. apply to stetheschope , not scalpel (----applied to 3 projects-----)
  7. reliable explanations of ai behaviour (----4 projects ----)
  8. also may be Does Reinforcement Learning Improve a Transformer’s Access to Its Own Internal Errors (-----5 projects ----)
  • someting on jspace
  • and probes

metrics:

  • M1: how much i like
  • M2: how strong appl can i give

introspection training for verbalization acts

  • project at antropic
  • TODO
    • read the papers in description
    • finish the jspace thing
    • may be try to link with probes and check
    • write proposal deeply

M1: 5, M2: 4

on policy training for deception probes

  • looking at internals and lies
  • TODO
    • understand the proj
    • read papers
    • answer the questions
    • write proposal

M1: 5, M2: 4

cross organism interpretability

  • tries to align multiple views

M1: 5, M2: 4

location knowledge with infludence functions:

  • location knowledge, similar to bau’s work
  • read infludence functions M1: 5, M2: 3

represnetation diagnostics for safety

M1: 5, M2:4

orthogonalization against reward hacking

M1: 5, M2: 3

  • read about orthognoalization
  • and see how it can stop reward hacking

feature geometry for fine tuning

M1: 5, M2: 3 fine tune a LLM on malicious code for understanding ?

what tokens lead to misalignment

M1: 5, M2: 3

  • data attrn , looks interesting

Reliable explainations of AI across funtionaly equivalent

M1: 5, M2 : 4

  • needs to think of toy models

geometry, channel capacity, detection of sublimal learing

  • read sublimal learning