Neil Chowdhury

I'm a Member of Technical Staff at Thinking Machines Lab.

I work on AI safety and alignment. Previously, I was a founding engineer at Transluce and a safety researcher at OpenAI.

Publications

Showing 4 featured publications.

  • 2024

    Introducing SWE-bench Verified

    A human-validated benchmark for more reliable evaluation of AI coding capabilities.

    Neil Chowdhury*, James Aung*, Chan Jun Shern*, Oliver Jaffe*, Dane Sherburn*, Giulio Starace*, Evan Mays*, Rachel Dias, Marwan Aljubeh, Mia Glaese, Carlos E. Jimenez, John Yang, Kevin Liu, Aleksander Madry

  • 2025

    Surfacing Pathological Behaviors in Language Models

    A new red-teaming objective for discovering rare, harmful behaviors in language models.

    Neil Chowdhury, Sarah Schwettmann, Jacob Steinhardt, Daniel D. Johnson

  • 2025

    MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

    Testing AI agents on 75 Kaggle machine learning competitions.

    Jun Shern Chan*, Neil Chowdhury*, Oliver Jaffe*, James Aung*, Dane Sherburn*, Evan Mays*, Giulio Starace*, Kevin Liu, Leon Maksin, Tejal Patwardhan, Lilian Weng, Aleksander Madry

  • 2026

    WeirdChat: A catalog of unexpected AI behaviors, discovered automatically

    A public catalog of unexpected behaviors in frontier open-weight models with 175,000+ annotated transcripts.

    Neil Chowdhury, Cassidy Laidlaw, Kaiying Hou, Daniel Johnson, Sarah Schwettmann, Jacob Steinhardt