Current priorities and next steps

Working snapshot — 14 August 2026

1. Finish the SPAR application

Outcome: submit the strongest possible applications around circuit-level model behavior and representation diagnostics.

Next steps

  • Finish the final question for represnetation diagnostics for safety.
  • Finish the final question for orthogonalization against reward hacking.
  • Complete the J-space/Jacobian Lens explanation and turn it into a concrete research proposal.
  • Add a clear role for probes: what they measure, how they complement causal interventions, and what result would falsify the hypothesis.
  • Rank the remaining project choices using the existing criteria: how much I like the project and how strong an application I can make.
  • Do a final pass that makes the research question, proposed experiment, expected result, alternatives, and my relevant experience explicit.

2. Turn circuit-level behavior explanation into a concrete project

Outcome: move from understanding attribution graphs to a testable project about why a model produces a behavior.

Next steps

  • Write the core research question in one sentence.
  • Use the Shanghai–Beijing attribution graph as a worked example: identify the path from place evidence through country and capital concepts to the answer.
  • Separate descriptive evidence from causal evidence. A readable graph is a hypothesis; intervention results are the test.
  • Design matched comparisons where source and target concepts play the same causal role and occur at comparable layers.
  • Define success before experimenting: semantic replacement, behavioral suppression, or merely output disruption.

3. Close the J-space/Jacobian Lens experiment loop

Outcome: produce one clean, interpretable intervention result and document its limitations.

Next steps

  • Resolve the intervention math and the layer-normalization question in jspace, jlens paper.
  • Start with one controlled single-token example; locate the concept across layers before manipulating it.
  • Compare the clean run, source suppression, target activation, and the combined swap.
  • Check whether the intervention changes the intended intermediate concept rather than only changing or degrading the final output.
  • Save the result, code path, plots, negative results, and interpretation together.

4. Keep learning focused on the applications

Outcome: turn reading into proposal-ready claims and experiments instead of expanding the reading queue.

Next steps

  • Continue the BlueDot technical AI safety material, but prioritize sections relevant to the chosen SPAR projects.
  • Read only the papers needed to answer a live application question or design an experiment.
  • After each paper, record: the main claim, the strongest evidence, one limitation, and one experiment it suggests.
  • Convert useful notes into concise text that can be reused in the application or project plan.

Suggested order

  1. Finish the two incomplete SPAR answers.
  2. Choose and sharpen the strongest proposal.
  3. Run one small J-space intervention that directly supports the proposal.
  4. Use the result—positive or negative—to strengthen the application narrative.
  5. Return to broader reading after the application is ready.

Maintenance — not a current focus

  • Obsidian/Quartz sync and index generation are already automated; intervene only if a run fails or the published result is wrong.
  • Codex remote/mobile setup is separate operational work; finish phone pairing only when remote access is immediately useful.

End-of-day check

  • What concrete application text did I finish?
  • What uncertainty did I reduce with evidence?
  • What is the single next action for tomorrow?