Introducing SWE-bench Verified
A human-validated benchmark for more reliable evaluation of AI coding capabilities.

I'm a Member of Technical Staff at Thinking Machines Lab.
I work on AI safety and alignment. Previously, I was a founding engineer at Transluce and a safety researcher at OpenAI.
Showing 4 featured publications.
A human-validated benchmark for more reliable evaluation of AI coding capabilities.
A new red-teaming objective for discovering rare, harmful behaviors in language models.
Testing AI agents on 75 Kaggle machine learning competitions.
A public catalog of unexpected behaviors in frontier open-weight models with 175,000+ annotated transcripts.