Current priorities and next steps
Working snapshot — 14 August 2026
1. Finish the SPAR application
Outcome: submit the strongest possible applications around circuit-level model behavior and representation diagnostics.
Next steps
- Finish the final question for represnetation diagnostics for safety.
- Finish the final question for orthogonalization against reward hacking.
- Complete the J-space/Jacobian Lens explanation and turn it into a concrete research proposal.
- Add a clear role for probes: what they measure, how they complement causal interventions, and what result would falsify the hypothesis.
- Rank the remaining project choices using the existing criteria: how much I like the project and how strong an application I can make.
- Do a final pass that makes the research question, proposed experiment, expected result, alternatives, and my relevant experience explicit.
2. Turn circuit-level behavior explanation into a concrete project
Outcome: move from understanding attribution graphs to a testable project about why a model produces a behavior.
Next steps
- Write the core research question in one sentence.
- Use the Shanghai–Beijing attribution graph as a worked example: identify the path from place evidence through country and capital concepts to the answer.
- Separate descriptive evidence from causal evidence. A readable graph is a hypothesis; intervention results are the test.
- Design matched comparisons where source and target concepts play the same causal role and occur at comparable layers.
- Define success before experimenting: semantic replacement, behavioral suppression, or merely output disruption.
3. Close the J-space/Jacobian Lens experiment loop
Outcome: produce one clean, interpretable intervention result and document its limitations.
Next steps
- Resolve the intervention math and the layer-normalization question in jspace, jlens paper.
- Start with one controlled single-token example; locate the concept across layers before manipulating it.
- Compare the clean run, source suppression, target activation, and the combined swap.
- Check whether the intervention changes the intended intermediate concept rather than only changing or degrading the final output.
- Save the result, code path, plots, negative results, and interpretation together.
4. Keep learning focused on the applications
Outcome: turn reading into proposal-ready claims and experiments instead of expanding the reading queue.
Next steps
- Continue the BlueDot technical AI safety material, but prioritize sections relevant to the chosen SPAR projects.
- Read only the papers needed to answer a live application question or design an experiment.
- After each paper, record: the main claim, the strongest evidence, one limitation, and one experiment it suggests.
- Convert useful notes into concise text that can be reused in the application or project plan.
Suggested order
- Finish the two incomplete SPAR answers.
- Choose and sharpen the strongest proposal.
- Run one small J-space intervention that directly supports the proposal.
- Use the result—positive or negative—to strengthen the application narrative.
- Return to broader reading after the application is ready.
Maintenance — not a current focus
- Obsidian/Quartz sync and index generation are already automated; intervene only if a run fails or the published result is wrong.
- Codex remote/mobile setup is separate operational work; finish phone pairing only when remote access is immediately useful.
End-of-day check
- What concrete application text did I finish?
- What uncertainty did I reduce with evidence?
- What is the single next action for tomorrow?