Research Scientist, Anthropic
Collaborator on LLM evaluation and AI safety research. MATS mentor.
Research Scientist
Co-author on the Factorio Learning Environment and open-ended agent evaluation.
Co-author on the Factorio Learning Environment.
Collaborator on AI safety research at Anthropic.
AI Safety Researcher
Co-author on overthinking in reasoning models.
Co-author on self-attribution bias in large language models.
Co-author on lie detection generalisation in LLMs.