| Does lie detection transfer across lie types? | MATS/Anthropic |
A reliable lie detector for language models would be transformative for AI safety: if we could detect when a model is lying, we could detect any misalignment simply by asking the model about it. This research investigates whether lie detectors trained on one type of lie transfer to detecting other types.
The work was conducted as part of the MATS (ML Alignment Theory Scholars) programme, under the mentorship of Fabien Roger.
We fine-tuned models to detect one type of lie (or half of the lie types in our dataset) and evaluated them on the remaining types. We found minimal transfer — detectors that perform well on the lie types they were trained on fail to generalise to new types. This is a concerning result for two reasons:
[Generalisation results figure placeholder]
Forthcoming as an Anthropic blog post, 2026.