TODO
- signatures of scheming: https://sparai.org/projects/f26/rec8h4EE5tioWPoGA/ -(similar: https://sparai.org/projects/f26/recMpvlAH5wPRp4Mj/)
- lottery ticket hypothesis: https://sparai.org/projects/f26/recArOjpEPNym0XZD/
- measuring grader awareness: https://sparai.org/projects/f26/recYFIyEAEjTCHuC2/
TODO
-
last question of representation diagnostics for sagety
-
last question also of reward hacking orthoganiation
-
project: Model psychology & neuroscience Explain behavior on the circuit-level
-
(6 projects)
-
wat tokens lead to misalignment: https://sparai.org/projects/f26/recn72ZNYRMuOGcV4
-
(7th project)
-
project: 2nd looks research: reproduce scalpel paper?
useful todo
read refusal paper refusal has a single directionapply to represnetation diagnostics for safety and orthogonalization against reward hackingread fine tuning , sae scalpels stethescope- finish jspace
- write something in finetuning
apply to stetheschope , not scalpel(----applied to 3 projects-----)reliable explanations of ai behaviour(----4 projects ----)also may be Does Reinforcement Learning Improve a Transformer’s Access to Its Own Internal Errors(-----5 projects ----)
- someting on jspace
- and probes
metrics:
- M1: how much i like
- M2: how strong appl can i give
introspection training for verbalization acts
- project at antropic
- TODO
- read the papers in description
- finish the jspace thing
- may be try to link with probes and check
- write proposal deeply
M1: 5, M2: 4
on policy training for deception probes
- looking at internals and lies
- TODO
- understand the proj
- read papers
- answer the questions
- write proposal
M1: 5, M2: 4
cross organism interpretability
- tries to align multiple views
M1: 5, M2: 4
location knowledge with infludence functions:
- location knowledge, similar to bau’s work
- read infludence functions M1: 5, M2: 3
represnetation diagnostics for safety
M1: 5, M2:4
orthogonalization against reward hacking
M1: 5, M2: 3
- read about orthognoalization
- and see how it can stop reward hacking
feature geometry for fine tuning
M1: 5, M2: 3 fine tune a LLM on malicious code for understanding ?
what tokens lead to misalignment
M1: 5, M2: 3
- data attrn , looks interesting
Reliable explainations of AI across funtionaly equivalent
M1: 5, M2 : 4
- needs to think of toy models
geometry, channel capacity, detection of sublimal learing
- read sublimal learning