Jack Hopkins Independent AI Safety Researcher

Research

I’m an independent AI safety researcher based in London, focused on understanding and evaluating the behaviour of large language models. My current work explores self-attribution bias, reasoning amplification for safety auditing, and the limits of LLM-based lie detection.

I built the Factorio Learning Environment, an open-ended benchmark for evaluating LLM agents in complex, unbounded scenarios. I’m also a MATS scholar, where I have studied deceptive behaviour in language models.

Before moving into safety research, I co-founded two startups: Paperplane (YC W23), a conversational intelligence platform, and Spherical Defence, an API security company. I studied Computer Science at Cambridge (MPhil, Distinction) and Royal Holloway, University of London (BSc, 1st Class).

Outside of research, I once rowed across the Atlantic Ocean in 40 days and launched a gin company (Northwest Passage Expedition Gin) during COVID.

Home

Publications

Self-Attribution Bias: When AI Monitors Go Easy on Themselves. Dipika Khullar, Jack Hopkins, Rowan Wang, Fabien Roger. 2026. arXiv preprint.

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets. Jack Hopkins, Dipika Khullar, Fabien Roger. 2026. arXiv preprint.

Factorio Learning Environment. Jack Hopkins, Mart Bakler, Akbir Khan. 2025. NeurIPS.

Automatically Generating Rhythmic Verse with Neural Networks. Jack Hopkins, Douwe Kiela. 2017. Association for Computational Linguistics (Volume 1: Long Papers) Pages 168-178.