Jack Hopkins AI Safety Researcher

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets (arXiv)

Our paper on amplifying reasoning task vectors to surface hidden information is on arXiv, and accepted at ICML 2026. With Rowan Wang, Dipika Khullar, and Fabien Roger. More context on the project page.