Open-ended LLM agent benchmark built in Factorio
An open-source environment for evaluating LLM agents on complex, unbounded automation tasks with exponential complexity scaling.
NeurIPS 2025 935+ GitHub stars Open-source
40-day unsupported Atlantic crossing with a team of 4
Rowed 3,000 miles across the Atlantic Ocean in the 2018-19 Talisker Whisky Atlantic Challenge, finishing 5th and raising £25k for Multiple Sclerosis charities.
40 days at sea 5th place finish £25k raised for MS charities
Do fine-tuned lie detectors generalise across domains?
Research investigating whether LLM-based lie detectors trained on one domain can detect deception in others, with implications for AI safety.
MATS/Anthropic fellowship Forthcoming 2026
Amplifying reasoning weights to extract learned secrets
Using reasoning task vector arithmetic to amplify models' propensity to think out loud, surfacing hidden information up to 10× more frequently for pre-deployment safety auditing.
ICML 2026 submission Up to 10× improvement 2B-32B scale
Neural machine translation of ancient cuneiform
Applying test-time compute to ancient language translation. Distilled reasoning chains from Claude Opus nearly double BLEU (63 vs 37) on Akkadian-to-English, filtering 140k ORACC lines down to 26k high-quality training pairs.
63 BLEU (val) 140k → 26k filtered pairs