Jack Hopkins AI Safety Researcher

About

I’m an independent AI safety researcher based in London, focused on understanding and evaluating the behaviour of large language models. My current work explores self-attribution bias, reasoning amplification for safety auditing, and the limits of LLM-based lie detection.

I built the Factorio Learning Environment, an open-ended benchmark for evaluating LLM agents in complex, unbounded scenarios. I was an Anthropic fellow and MATS scholar (2025-26), where I studied deceptive behaviour in language models.

Before moving into safety research, I co-founded two startups: Paperplane (YC W23), a conversational intelligence platform, and Spherical Defence, an API security company. I left because startup life felt a little solipsistic (how can I build something, so that I raise money, so that I make money), and I never found an idea with enough impact to be worth doing even if it made no money. I studied Computer Science at Cambridge (MPhil, Distinction) and Royal Holloway, University of London (BSc, 1st Class).

Outside of research, I once rowed across the Atlantic Ocean in 40 days and launched a gin company (Northwest Passage Expedition Gin) during COVID. I have also been teaching a language model to translate ancient Akkadian.

Jack and Rosie
We got married!