Jack Hopkins AI Safety Researcher

Self-Attribution Bias: When AI Monitors Go Easy on Themselves (arXiv)

Our paper showing that AI monitors judge their own actions as safer than identical actions by others is on arXiv. Led by Dipika Khullar, with Rowan Wang and Fabien Roger.