The views expressed here are my own.

I want to live in a world where people understand, shape, and have control over the AI they use. As AI becomes more capable, the question of who owns it and who decides how it is used will have enormous consequences for how power is distributed in society. That power should be shared widely.

I fear we are headed in the wrong direction. Labs have predominantly deployed models in centralized services, a choice often defended on safety grounds. But this view underestimates the risks of concentrating power in a handful of companies, letting them decide what values models should have, and limiting outside scrutiny.

Sharing control over AI does make safety harder. Powerful models can be misused, and most current safeguards rely on developers monitoring model usage and restricting access to dual-use capabilities. Releasing model weights removes many of these safeguards, which assume centralized control. Building safeguards that don’t is one of the most important and neglected problems in AI safety.

Thinking Machines is betting on a future where AI is shaped by the people who use it. I’m joining to help make that future safe to build.

Control over AI should be democratized

I’ve done technical AI safety research for the past four years at MIT, OpenAI, and most recently Transluce. Most of that work has focused on misalignment: the risk that increasingly capable AI systems pursue goals that diverge from human intent, with potentially catastrophic consequences.

Over time, I’ve become concerned about capable AI being controlled by only a few corporations. Even if we solve the core technical problems of alignment, AI could still concentrate enormous power. If our institutions all depend on the same few models, labs’ choices about how models behave and whose priorities they should serve could become defaults for society.

I want individuals and organizations to have a meaningful say in these choices. I see multiple benefits to making models more open and customizable:

Values can be set by individuals, not a central authority. Model developers make consequential choices about which values their models should reflect. Openness makes these choices more transparent, and customizability gives people better tools to instill their own goals and values.

Open access can accelerate alignment research. The more researchers can experiment with powerful models, the faster we can understand failure modes and develop better alignment techniques. Today, some of the most useful tools for inspecting models—such as chain-of-thought reasoning—are hidden from external researchers. I believe broader access to models, including their internals, is essential for scientific scrutiny and societal preparedness.

Customization can spread economic value. Frontier models are strong on general capabilities, but perform unevenly on domains with limited training coverage. When companies can fine-tune models on their own data and own the resulting intelligence, they retain more of the value they create and diffuse power more broadly through the economy.

Realizing these benefits, however, will require safety mechanisms that work even when people can customize the models they use.

Making distributed control safe

Today, many frontier safety measures rely on developers retaining control over deployment. Labs can monitor requests for misuse, restrict access to dual-use capabilities like cybersecurity and biology, and patch vulnerabilities as they are found. Once weights are released openly, none of these controls can be enforced. Model-level safeguards such as refusal training can be trivially undone through fine-tuning.

Building safety mechanisms for distributed control requires progress on several fronts:

Preventing catastrophic misuse. The more open models are, the easier it is for bad actors to access dangerous capabilities. Models like GPT-6 Astra and Claude Mythos 5 are already highly capable at cyber offense. Open weights cannot be retracted once released. One strategy is to make open-weight models safer with safeguards that survive fine-tuning, for example by filtering dual-use information from pretraining data. Another is to place safeguards on training or inference APIs to block catastrophic misuse while allowing users broad freedom to customize models.

Maintaining alignment under customization. Fine-tuning can degrade alignment in unintended ways. In an extreme case known as emergent misalignment, training models on insecure code induced unrelated harmful behavior, such as advocating for humans to be enslaved by AI. In another study, a model that learned to reward hack in real coding environments later sabotaged evaluation infrastructure to obtain answer keys. If people are going to customize models deeply, they need access to alignment methods, evaluations, and monitoring tools that shape behavior and detect failures during deployment.

Why Thinking Machines

Thinking Machines wants AI to be shaped continuously by the people and organizations that use it. Its technical agenda includes training highly capable models, releasing tools such as Tinker for customizing them, and publishing research that helps others understand how AI is developed. This is aligned with the future I want to help build.

I expect safety to be a core bottleneck for this agenda. The more control we give people over powerful models, the more important it is to build safety mechanisms that work without assuming centralized control. Progress on safety will enable deeper customization, which gives Thinking Machines a direct commercial incentive to invest in the research I want to pursue.

My conversations with leadership convinced me that they care deeply about these problems and want to do work that benefits the broader AI ecosystem. Joining was an easy decision.

If these views resonate with you, come join!