The instruction that started the run was a single sentence: compute the nine-loop, six-particle hexagon amplitude. What followed was closer to a standing order than a research program. “I am going to sleep,” the operator wrote. “Keep going. Give me an update every four to six hours.”
Anthropic published the results on September 25 in a guest post on its research site. Two of its physicists had set Claude running for several days with almost no scientific supervision, and the model produced the six-gluon scattering amplitude in planar N=4 super-Yang-Mills theory at nine loops, a calculation that sits near the edge of what the field attempts.
The quantity involved describes how particles scatter. Each additional loop in the calculation adds a layer of complexity, and the resulting expressions grow so quickly that physicists rely on symmetry, clever notation and years of accumulated technique to keep them manageable. Nine loops is far beyond what anyone computes by hand.
The cost was modest by the standards of frontier AI work. The total run came to roughly $1,000 to $2,000, according to Anthropic, with about $100 of that spent on a bootstrap computation that occupied 96 CPUs for a week. The rest went to model inference and orchestration.
The system was a model Anthropic identifies as Fable 5.1, paired with a Claude Science toolchain built for physics workflows. Two independent routes through the calculation produced matching answers, and the run also reproduced a previously published eight-loop result, which gave the team a way to check that the machinery was working before trusting the new output.
Verification came from outside the company. Lance Dixon, a physicist at SLAC National Accelerator Laboratory and Stanford University, checked the result independently. His own group had spent about two years pursuing the same target. “A machine solved a problem I thought could not be done directly,” he said.
The claim is narrower than the headline suggests, and the authors said so. Matthew von Hippel, a co-author of the post, wrote that the computation is not a new law of physics. A team at the Chinese Academy of Sciences led by He Song had already obtained most of the same results using a separate AI-assisted method, he noted, which places the work in an existing line of research rather than at its origin.
What is different is the amount of human attention involved. The operators described the run as a sequence of check-ins rather than a collaboration in the usual sense. They set the objective, went to sleep, and reviewed updates when they woke. The model chose intermediate steps that the physicists said they had not specified.
That distinction matters to how the field judges such results. A calculation done by a model over days cannot be explained step by step the way a paper by a graduate student can, and the reasoning path is not fully legible even to the people who built the system. Independent replication of the final number is the check that remains available.
The physics payoff is not immediate. Amplitudes at high loop order feed into precision predictions for collider experiments, where theory and measurement are compared at fine resolution. The nine-loop result is unlikely to change any experimental conclusion on its own, but it extends the range of what tools can reach, and the methods may travel to other theories where the same obstacles appear.
The result also fits a run of machine progress in mathematics and theoretical physics over the past two years. Competition benchmarks have been cleared, olympiad problems have fallen, and several groups have reported model-assisted proofs. Each of those claims has been met with the same question about who deserves credit and how the work should be checked.
Anthropic published the work as a guest post rather than a formal paper, which reflects how results reach this field now. There was no peer review before the announcement and no preprint carrying a full derivation. Readers received a description of the method, the computed object, and the name of an outside physicist who checked the answer against his own expectations.
The toolchain matters as much as the model. Claude Science is built to run long jobs, keep state across sessions and hand intermediate results to conventional computer algebra systems, which is what lets a calculation continue while no one is watching it.
Anthropic has not said how much of the computation would have been feasible with a smaller model or a shorter run, which is the comparison that would show how much of the outcome came from scale. The company also has not released the full reasoning trace, only the result and the methods used to verify it.
Whether a model that works through the night on a hard calculation becomes a standard instrument in theoretical physics, or remains a demonstration that needs a physicist standing next to it, is the question the field has not yet answered.


