Jack Hopkins AI Safety Researcher

I’m an AI safety researcher based in London. My work asks whether we can catch language models being dishonest, and whether the tools we build to catch them can themselves be trusted. As an Anthropic fellow and MATS scholar (2025-26), mentored by Fabien Roger, I studied deception in language models: training lie detectors, measuring the biases of AI monitors, and amplifying reasoning to surface what models hide. More about me →

I’m continuing this research independently, and I’m open to collaborations and research roles. The fastest way to reach me is jack.hopkins@me.com.

Selected work

Fine-Tuned Lie Detectors Failed to Generalize (Anthropic Alignment Science Blog, 2026). Lie detectors that excel on the lies they were trained on learn narrow features and fail on unseen lie types. With Dipika Khullar, Rowan Wang, and Fabien Roger. Project page →

Factorio Learning Environment (NeurIPS 2025). An open-source, open-ended benchmark for evaluating LLM agents on unbounded automation tasks. 1.2k GitHub stars.

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets (ICML 2026). Amplifying reasoning task vectors destabilises models just enough to reveal hidden information up to 10x more often, an auditing window that widens with scale. With Rowan Wang, Dipika Khullar, and Fabien Roger. Project page →

GPT-OSS-20B is a Liar (2025). Winner of the Overall Track of OpenAI’s gpt-oss-20b red-teaming challenge on Kaggle: an on-policy elicitation pipeline that surfaced thousands of lies across 46 settings. With Dipika Khullar.

All projects → · All publications →