In December 2020, DeepMind announced that AlphaFold had essentially solved protein folding—a problem that had resisted solution for half a century. The system predicted protein structures with near-atomic accuracy, rivaling experimental methods. By 2022, it had mapped the structures of roughly 200 million proteins. In 2024, Demis Hassabis and John Jumper were awarded the Nobel Prize in Chemistry for this work.
AlphaFold is now the canonical example of AI producing a genuine scientific breakthrough.
The natural question is: when will we see the equivalent in physics?
Machine learning systems are increasingly powerful, but they remain stochastic, opaque, and prone to subtle failure. Fields like theoretical physics and pure mathematics place a premium on rigor, derivation, and conceptual clarity. So how should AI be deployed in domains where explanation matters as much as prediction?
Over the past few weeks, I’ve been reading and talking with researchers about what it would actually take for AI to produce a physics breakthrough on the scale of AlphaFold.
Signs of Movement
It would be wrong to say nothing is happening.
OpenAI recently reported results showing GPT-5.2 contributed to progress in theoretical physics. In particular, it demonstrated that under constrainted conditions, the scattering amplitudes of “single-minus” gluons were non-zero. These were previously believed to be zero. The team worked with an internally scaffolded version of GPT-5.2. The model operated for roughly 12 hours and a team of physicists from Harvard, Cambridge, and the Institute for Advanced Study checked the results. This may not represent a breakthrough in theoretical physics, but it is nevertheless a major milestone. The AI partnered with human experts and uncovered a previously unknown result that may have remained hidden for decades otherwise.
Similarly, Google DeepMind has reported substantial progress with Gemini Deep Think. Building on IMO gold-medal performance in 2025, the system has contributed to professional research across mathematics, physics, and computer science: solutions to open problems from the Erdős conjectures database, progress on long-stalled CS problems such as Max-Cut and Steiner Tree, and analytical solutions for gravitational radiation from cosmic strings. As with the OpenAI result, the emphasis is on orchestration—agentic workflows with generation, verification, and revision, guided by expert researchers—rather than autonomous capability alone.
These are early signals, not paradigm shifts. But they suggest that the “LLM as research assistant” phase is evolving toward something more structured and self-directed.
Still, the breakthroughs are incremental. We have not yet seen a result in physics analogous to AlphaFold’s clean, decisive resolution of a longstanding benchmark.
What Would “AlphaFold for Physics” Even Mean?
Before asking when it will happen, we should clarify what we are expecting.
Progress in physics could take several forms:
- Hypothesis generation: proposing viable theoretical structures or effective field theories.
- Formal derivation and proof assistance: resolving conjectures or deriving new relationships.
- Simulation breakthroughs: dramatically compressing the time and cost required for high-fidelity modeling.
- Autonomous experimentation: systems that design and execute experiments in closed-loop environments.
- Physics-native foundation models: models trained directly on physical data, simulations, and formal mathematical objects—not just on language.
Reconciliation of quantum mechanics and general relativity would qualify. But that is an extreme case. More realistically, an “AlphaFold moment” in physics may first emerge where prediction, simulation, or inverse design unlocks new regimes of discovery.
The question is less “Can AI solve quantum gravity?” and more “Where are the structural conditions similar to protein folding circa 2020?”
Why Biology Was First
Protein folding had several properties that made it uniquely tractable:
- Data abundance — Decades of structural biology produced large, high-quality datasets.
- Clear benchmarks — The CASP competition provided a precise evaluation metric.
- Massive economic incentives — Predicting protein structure unlocks drug discovery. The commercial upside is obvious.
A generic chatbot could not have made progress against the challenge of protein folding and drug design, even one that had been trained on biology textbooks and the scientific literature. Such a language model might be very useful to researchers, but not for structure prediction. It had to be a more task-specific model.
The breakthrough was domain-native.
The Physical Sciences Are Structurally Different
Physics has abundant data—collider outputs, astronomical surveys, condensed matter experiments—but the relationship between data and deep theoretical unification is often indirect.
Many of the hardest problems in physics lack:
- A single agreed-upon benchmark.
- Dense, labeled supervision.
- Clear short-term commercial payoff.
The bottleneck in theoretical physics is not merely pattern recognition. It is conceptual synthesis.
Today’s progress largely takes a centaur form: human researchers augmented by AI tools for literature review, coding, data analysis, and drafting. Systems like GPT-5.2 and Gemini Deep Think push that boundary further by integrating longer-horizon reasoning and tool use. But they remain components in a larger human-directed loop.
To move beyond augmentation toward independent scientific agency, something else is required.
The Foundation Model Hypothesis
We are in a Cambrian explosion of foundation models.
A foundation model follows a “train once, adapt broadly” paradigm. Earlier machine learning systems were trained for a single task with curated labels. Foundation models are trained at scale—often self-supervised—and adapted downstream.
Large language models are one species. But we now see foundation models for vision, robotics, geospatial data, time series, graphs, biology, and increasingly, world simulation.
The biology breakthrough came from a foundation model trained on protein data—not from a general-purpose language model alone.
By analogy, physics breakthroughs are unlikely to emerge from prompting an LLM to “derive string theory.” They are more likely to come from:
- Physics-informed world models trained on large simulation corpora and experimental data.
- Neural surrogates that approximate expensive solvers.
- Agentic systems that design and execute large simulation campaigns.
- Hybrid architectures that combine conjecture generation with formal verification.
Language models in this picture are orchestration layers: proposing hypotheses, structuring experiments, managing toolchains, and documenting outcomes. The physical reasoning itself will likely live in specialized models trained directly on physical structure and experimental data.
Why World Models Matter
This is where efforts like World Labs become strategically important.
The core thesis behind world-model research is straightforward: intelligence requires structured internal models of physical reality. A system that understands geometry, dynamics, occlusion, and causality in three dimensions is qualitatively different from one that merely predicts the next token in text.
If we want AI systems that can meaningfully participate in physics—designing experiments, reasoning about constraints, manipulating physical abstractions—we need models grounded in space, time, and invariance. Physical reasoning is not just symbolic manipulation; it is constrained by geometry and dynamics.
World Labs and related efforts aim to build foundation models of 3D environments and physical structure. Even if their near-term applications are robotics or embodied AI, the underlying representations—persistent object identity, geometric consistency, physical priors—are precisely the substrate one would want for AI-native physical reasoning.
They are not “doing physics” in the academic sense. They are building the infrastructure upon which physics-capable systems might eventually run.
The most serious commercial efforts in this direction are not aimed at quantum gravity. They are concentrated in three economically grounded domains:
- Industrial simulation, exemplified by PhysicsX, where machine learning is used to accelerate or surrogate computational fluid dynamics, finite-element analysis, and engineering optimization.
- Physics-informed materials discovery, as pursued by companies such as Periodic Labs, where generative and predictive models operate over constrained physical design spaces.
- Embodied and spatial world models, pursued by organizations like World Labs, which aim to build persistent 3D representations of physical environments—capabilities that robotics companies such as Figure AI will ultimately depend on.
These sectors—aerospace, energy, semiconductors, advanced manufacturing, and humanoid robotics—represent trillions of dollars in downstream economic value. That scale of opportunity justifies training large physics-grounded foundation models and building scalable simulation stacks.
The industrial use case funds the world model.
If fundamental physics advances as a consequence, it will likely do so as a second-order effect: tools built to optimize aircraft components or robotic manipulation may later be repurposed to explore regimes of plasma confinement, condensed matter systems, or cosmological structure formation.
Biology’s breakthrough emerged where data, architecture, and incentives aligned. In physics, that alignment may first occur in industry.
Only later might it migrate upstream.
Simulation as the Likely First Lever
If there is a near-term path to “AlphaFold for physics,” it likely runs through simulation.
High-fidelity numerical modeling—CFD, finite element methods, plasma simulations, cosmological N-body codes—is computationally expensive. In many fields, wall-clock simulation time is the binding constraint.
AI offers two accelerants:
- Surrogate modeling: replacing expensive solvers with learned approximations.
- Search over design spaces: generative exploration of parameter regimes that brute-force methods cannot reach.
The immediate payoff is industrial. But once the infrastructure exists, it diffuses.
Physics may not experience a single dramatic announcement. Instead, we may see:
- A 100-1000× compression in simulation time for certain regimes.
- AI-generated materials validated experimentally.
- A hybrid symbolic-neural system resolving a longstanding mathematical structure.
- An autonomous lab that closes the loop from hypothesis to experimental validation.
Only retrospectively will we identify the inflection point.
When Will Physics Have Its AlphaFold Moment?
The answer may depend less on scaling language models and more on building physics-native foundation models and world models with strong inductive biases toward physical law.
Biology’s breakthrough came when data, incentives, benchmarks, and architecture aligned.
Physics will require a similar alignment:
- Clear intermediate targets.
- Large-scale, structured training corpora.
- Economic incentives that justify training expensive models.
- Hybrid systems that combine neural intuition with formal verification.
We are beginning to see the scaffolding: GPT-5.2 contributing within structured toolchains; Gemini Deep Think extending reasoning horizons; world-model startups building spatially grounded representations; simulation companies compressing industrial experimentation.
The first AlphaFold moment in physics may not unify quantum gravity. It may instead quietly remove a computational bottleneck that has constrained a field for decades.
And only in hindsight will it appear obvious.