What was announced

On September 25, 2026, Anthropic published a guest post by physicist and science writer Matt von Hippel reporting that its Claude model had, for a few thousand dollars of end-user compute, completed a calculation in theoretical physics that the field had expected to be out of reach. The report says Claude computed the six-particle amplitude in planar N=4 super Yang-Mills at nine loops, one loop beyond the previous record, using known methods and running largely unsupervised for days. Lance Dixon of the SLAC National Accelerator Laboratory, who set the previous eight-loop record, independently verified the result, according to the post.

The claim is verified in the narrow sense that Dixon, the person best positioned to check the answer, reviewed it and confirmed it, per the same account. The calculation itself has not yet been published as a peer-reviewed paper by the humans involved, so this article treats the Anthropic post as a vendor-published announcement with strong independent verification of the specific result, not as independent reproduction.

The challenge

The story begins with a challenge. Von Hippel, a former theoretical physicist who writes the blog 4gravitons, published a post on August 7, 2026 inviting AI companies to take on his old subfield. He retrieved the text of the challenge as quoted in his September 25 Anthropic guest post:

"If AI companies want to impress people like me (or scare us, for that matter), then they need to tackle my old field. Show that an AI can take the kinds of computer resources an academic has access to, and solve one of the scattering amplitudes field's big outstanding problems. Show that a computational limit everyone expected to be a problem doesn't actually matter. Give us N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops."

Attribution: Matt von Hippel, physicist and science writer, as quoted in his guest post on the Anthropic Science Blog, September 25, 2026, "Yes, Claude can do Nine Loops." The wording is as retrieved from that post; the original August 7 blog post itself was not retrievable in this session, so the wording is cited to its republication on the Anthropic post.

His reasoning was specific. New ideas in physics are hard to price in advance. He wanted a problem that seemed blocked not by a lack of ideas but by a lack of computers and time: something where the method was known, but running it seemed out of reach for academic researchers. Nine loops in N=4 super Yang-Mills fit that description. The previous record, eight loops, had been set by Lance Dixon and Yu-Ting Liu in a paper submitted to arXiv on August 16, 2023, using a technique called antipodal duality and an easier-to-compute related object called a form factor.

Why the calculation is hard, and why the model is a test bed

To see why nine loops matters, start with what particle physicists actually compute. When they predict how particles should behave in a collider, they calculate quantities called scattering amplitudes: formulas that turn particle momenta and energies into probabilities for particular reactions. If experimental results from the Large Hadron Collider match those predictions, existing theory holds. A mismatch could point to new physics, such as an explanation for dark matter.

These formulas are brutally hard to compute exactly, so physicists truncate them. They stop after a certain number of "loops," which measure how complicated the allowed particle interactions are. Each additional loop makes the answer more precise and the calculation much more expensive. In practice, most scattering amplitude calculations in the literature stop at two loops, a few reach three. For the toy model in question, the record stood at eight.

The model here, N=4 super Yang-Mills, is deliberately not our universe. It is a "toy model": a highly symmetric cousin of the Yang-Mills theories that describe electromagnetism and the nuclear forces. The extra symmetry makes it paradoxically easier to calculate with, so amplitude researchers use it to stress-test new techniques. A nine-loop result there is a benchmark for methods, not a new prediction about real particles.

How the bootstrap works, like Sudoku

The method Claude used, the bootstrap, is best understood through the analogy von Hippel offers: it works a bit like Sudoku. Researchers first constrain what the answer can look like, storing every remaining possibility in a specialized notation. Then they apply every known constraint: consistency rules the answer must obey, results from other calculation techniques, connections to related problems. Possibilities get crossed out until, ideally, exactly one survives, with checks left over to catch mistakes.

The field did not believe a straightforward extension of this method to nine loops was feasible. Dixon himself had reached eight loops indirectly, via the form-factor route, and expected the next step to require a different approach. "If people thought it was possible to just run the usual bootstrap method for one more loop, someone would have done it," von Hippel writes in the post.

That expectation is what makes the result interesting. The calculation was not a computational barrier after all. It was a barrier of effort: nobody had been willing to spend the computers and the weeks on it. Claude, in effect, was willing.

What Anthropic says Claude actually did

According to the post, two physicists at Anthropic, Liam Fitzpatrick and Siddharth Mishra-Sharma, contacted von Hippel at the end of August saying they had tackled one of his challenges. They used Fable 5.1 inside Claude Science, a paid platform built on the Claude model with structured prompts and rules, what the field calls a "harness."

After asking Claude which of the posed problems it was most likely to solve, they gave it essentially one prompt. As quoted in the post:

"The problem is to compute the Six-particle (hexagon) amplitude in planar N=4 SYM at nine loops."

Attribution: the prompt given to Claude by the Anthropic physicists, as quoted by Matt von Hippel in the Anthropic Science Blog post of September 25, 2026.

From there, the supervision was minimal. Another quoted instruction:

"I'm going to sleep and won't be available for another several hours. Keep working on this until I tell you to stop. Give me updates every 4-6 hours."

Attribution: instruction from the Anthropic physicists to Claude during the run, as quoted by Matt von Hippel, Anthropic Science Blog, September 25, 2026.

Claude ended up solving the problem two ways: the original bootstrap and the indirect form-factor approach. Either route cost an end-user roughly one to two thousand dollars, mostly from running the model for so long. The bootstrap computation itself, done in Python with the SymPy package, accounted for about 100 dollars of that, corresponding to 96 CPUs running for a week. Von Hippel notes that such a compute allocation would have felt substantial a decade ago but is cheap today. What stands out is not exotic capability but persistence: the harness got a finicky, error-prone calculation to the end in one shot, with no supervision more sophisticated than "keep going."

Humans were close behind

A few days after von Hippel heard from Anthropic, he heard from Song He of the Chinese Academy of Sciences in Beijing. According to the Anthropic post, Song's group had already obtained the majority of the nine-loop result, using some AI assistance based on a model called GPT-6, but through a more human-in-the-loop process than Anthropic's near hands-off run.

This article could not retrieve an independent primary source for the Song He group's result, so that account rests solely on the Anthropic post's reporting and should be read that way. Even taken on its own terms, though, it changes the story's meaning: the most significant datum may be that humans, with AI assistance, were close behind. The problem was not beyond human reach; it was just under-prioritized.

The post also reports that Dixon and Song and their collaborators will publish the results, with time taken to explain and analyze them. Claude's role, as von Hippel puts it, is "done, for now."

What this does and does not show

Von Hippel's own conclusion is notably modest. As he puts it:

"My biggest takeaway is that there is more low-hanging fruit out there than you'd expect."

Attribution: Matt von Hippel, physicist and science writer, Anthropic Science Blog, September 25, 2026, "Yes, Claude can do Nine Loops."

That is the right frame. This was not a demonstration that AI invents new physics. Claude used known methods, run harder and longer than humans had bothered. Von Hippel speculates the choice of Python rather than Maple or Mathematica, and better software engineering habits, may have helped, but describes neither as super-intelligent. The observed fact is a verified calculation. The prediction, which is opinion and should be labeled as such, is that the bottleneck in some fields is less about brilliant new ideas than about available effort: researchers with finite careers rationally decline projects that might consume a year and fail.

What this does show, on the evidence available: a frontier problem considered computationally blocked was solvable on academic-scale resources, an AI system completed it with minimal supervision, and an independent domain expert could verify the answer. What it does not show: that AI can choose which problems matter, formulate new theories, or replace the experimental and theoretical apparatus that gives calculations like this their meaning. In this episode the humans set the goal, supplied the verification, and will write the papers.

For readers, the practical lesson is the one von Hippel draws. When a well-defined problem looks impossible, it is worth asking whether it is blocked by ideas or merely by effort, because the second kind of barrier is falling faster than experts expected.