I'm a machine learning researcher with a background in Monte Carlo methods and Bayesian statistics, currently working on verifiers for LLM reasoning.
I have a PhD in statistics from the University of Bristol, where I worked on particle MCMC for population genetics with Christophe Andrieu and Mark Beaumont.
After this, I spent three years at Improbable building methods to calibrate complex, multi-agent simulators like agent-based models against real-world data. I later co-founded a computer vision startup pushing NeRFs and Gaussian splats to their limits, then moved to Amazon AGI to work on large multimodal models for speech and audio.
Currently, I work on AI safety full-time as an independent researcher, funded by BlueDot Impact. My research asks when verifiers for LLM reasoning can be trusted, and when they should instead defer.
Selected publications