Jack Hopkins AI Safety Researcher

Fine-Tuned Lie Detectors Failed to Generalize (Anthropic Alignment Science Blog)

Our lie detection research is out on the Anthropic Alignment Science Blog. With Dipika Khullar, Rowan Wang, and Fabien Roger. More context on the project page.