A halved 70-billion-parameter model that keeps its scores, according to its authors
The simplest way to make a large language model cheaper to run is also the bluntest: delete some of its internal building blocks. Each transformer block is a layer of processing, and removing them makes the model literally shorter, which cuts both memory use and the time each inference takes. The catch is that block removal, or depth pruning, tends to wreck the model if the wrong blocks are chosen, and the blocks interact with each other in ways that most existing methods ignore.
A team at Multiverse Computing, writing on the Hugging Face blog on September 21, 2026, argues that this selection problem is really a physics problem. Their paper, 'LLM Compression by Block Removal with Constrained Binary Optimization', reformulates the choice of which blocks to delete as what physicists call an Ising glass: a disordered system of interacting spins, the same mathematics used to describe certain magnetic materials. Once reframed that way, the authors say, a cheap 'energy' score can rank enormous numbers of candidate prunings without ever running the model on a benchmark.
The headline self-reported result concerns Llama-3.3-70B-Instruct, a model with 80 blocks, cut in half. With 40 of the 80 blocks removed and no retraining, the authors report the method holds MMLU, a broad knowledge benchmark, at 76.9, versus 54.0 for 'block influence', the strongest competing block-removal method they compared against. That is a gap of almost 23 percentage points. The original, unpruned model scored 82.2 on the same benchmark according to their table. These are the authors' own numbers and have not been independently verified.
Why cutting blocks one at a time is the wrong mental model
To see why the physics framing matters, consider how block removal is usually done. Most existing methods score each block individually, using heuristics such as magnitude, sensitivity or 'block influence', and then delete the blocks that look least important. The blog calls these mean-field methods, borrowing a term from physics: they treat each block as if its importance were independent of the others.
The authors argue that this independence assumption is wrong. Whether removing block 20 hurts the model depends on whether block 19 or block 24 was also removed. In physics language, the decisions are coupled, like neighboring spins in a magnet. Ignoring the couplings, they write, 'leaves quality on the table, especially when you want to remove a lot of blocks at once.'
The interactions also make the problem combinatorially huge. Each block either stays or goes, so the number of possible configurations grows exponentially with model depth. Brute force over all of them looks hopeless for a large model, which is why simple ranking heuristics have dominated. This is precisely the regime, exponentially large configuration spaces with pairwise couplings, where the statistical physics toolkit was built to operate.
Turning block selection into an energy-minimization problem
The reformulation works like this. Give each transformer block a binary variable: 0 means keep, 1 means remove, just like a spin pointing up or down. Then approximate how the model's loss changes as blocks are removed, using a second-order Taylor expansion that produces a Hessian matrix. The diagonal of that matrix captures how much each block matters on its own; the off-diagonal entries capture the pairwise couplings between blocks.
The optimization task is then to find the set of blocks whose removal minimizes an energy expression, subject to removing an exact number of blocks. The blog describes this as 'a constrained binary optimization (CBO) problem that maps directly onto an Ising glass, a disordered spin system with all-to-all interactions and a fixed number of "up" spins.' In this mapping, the fixed number of removed blocks plays the role of a fixed total spin.
The key claim, and the load-bearing one, is that this energy is a strong proxy for downstream quality: low-energy states of the spin system correspond to high-performing pruned models. If that holds, minimizing energy and maximizing benchmark score become the same search, and candidate configurations can be ranked without benchmarking any of them.
The practical payoff is cost. The Hessian, meaning the full set of couplings, is computed just once, from forward and backward passes on a small calibration dataset. After that, evaluating any candidate configuration is a single cheap energy calculation. Because the couplings do not depend on the compression target, the same Hessian can be reused to solve for many different numbers of removed blocks.
Solving it: brute force when possible, tabu search when not
For moderate cases, the authors brute-force the search on a single GPU, checking up to tens of billions of configurations. A few million take seconds; the hardest case they report, removing 8 of Llama-3.3-70B's 80 blocks, means about 29 billion configurations and took roughly two days on one GPU.
Beyond that scale, the exact approach breaks down, and this is where the Ising formulation pays off a second time. In its equivalent QUBO form, with the constraint absorbed into a penalty term, the same task can be handed to specialized solvers built for this class of problem: quantum annealing, QAOA, tabu search and branch-and-bound. The blog reports that 'an open-source tabu solver reliably reaches the lowest-energy states in seconds', even on the hardest cases they could verify against brute force.
A clarification worth emphasizing for general readers: no quantum computer performed the pruning. The solver doing the heavy lifting here is classical tabu search, an established combinatorial optimization technique. Quantum and quantum-inspired solvers are mentioned as part of the same solver family the company works with, and the mapping makes the problem available to them, but the reported results come from classical computation.
The authors also note a counterintuitive point about what counts as success. Conventional solvers are judged on finding the true ground state, the single lowest-energy configuration. They argue that is unnecessary: what practitioners need is a handful of good low-energy states, a much easier bar, which is why lightweight solvers suffice.
The best answer is often not the best answer
The energy proxy is strong but not perfect, so the single lowest-energy state is not always the best model. The authors present this as a feature rather than a flaw. Once the Hamiltonian is set up, reading off the ground state and the low-lying excited states, the next-best configurations in the energy ranking, costs essentially nothing, giving engineers a spectrum of candidate prunings to try.
They give a concrete example. For Llama-3.1-8B-Instruct with 16 of 32 blocks removed, most of the top-ranked states cut blocks toward the end of the model, which matches what prior work would predict. But the 17th excited state is the first to propose removing a block near the beginning, and after light retraining that configuration outperformed the ground state on several benchmarks. The blog concludes: 'The best model is an excited state, not the ground state.'
This also undercuts a common assumption in pruning practice, that the best removal is one consecutive chunk of middle or late blocks. The authors say this example 'directly disproves' that assumption, and attribute the win to respecting the full many-body structure of the problem.
Results beyond the flagship number
The authors report results across Llama-3.1-8B-Instruct, Qwen3-14B and Llama-3.3-70B-Instruct, claiming the method is on par with or better than state-of-the-art block-removal baselines, with the gap widening as compression gets more aggressive. Up to 24 of Llama-3.3-70B's 80 blocks removed, they say the method is roughly on par with block influence; at 32 and 40 removed, it pulls decisively ahead. For Qwen3-14B with 12 of 40 removed, they report about a 10-point MMLU lead.
To test generality, the team applied the method to NVIDIA-Nemotron-3-Nano-30B-A3B-FP8, a hybrid architecture that interleaves Mamba2, attention and mixture-of-experts layers in a non-uniform pattern. The Ising formulation does not care what kind of block sits at each site, so the method transfers without retraining. The authors report that removing 2 to 3 MoE layers, or 2 attention layers, produced configurations beating block influence on AIME25 and GPQA. They also observe that redundancy in these hybrid models is real but unevenly distributed, with some expert layers far more disposable than others.
The method is designed to compose with other compression techniques, including quantization, low-rank compression, width pruning and knowledge-distillation healing, rather than competing with them. The code has been open-sourced under the CompactifAI organization on GitHub.
What is verified, and what is still the vendor's word
Several caveats are essential here, and the assignment of confidence should be explicit.
First, every performance number in this article comes from the paper and blog written by the method's own authors. No independent replication of the MMLU results was found in the sources reviewed. Independent confirmation would require third parties running the open-sourced code on the same models and benchmarks, and no such verification is cited in the primary materials.
Second, the comparison depends on the chosen baselines. The almost 23-point gap is measured against block influence, the best competing method in their comparison set. Other compression approaches, such as quantization alone or combined pipelines, are described as complementary rather than directly compared at this setting.
Third, the headline result is for a model pruned without retraining. The authors note the method also works with a retraining or healing stage, and the excited-state example above required light retraining to beat the ground state, so the practical recipe varies by setting.
Finally, the energy proxy is approximate by construction: it comes from a second-order Taylor expansion of the loss. The claim that low energy reliably predicts high benchmark quality is the paper's central empirical assertion, not a mathematical guarantee, and it is exactly the claim most worth independent testing.
What it means if it holds
If the self-reported results hold up under independent replication, the practical implication is straightforward: smaller AI models that run on ordinary hardware with less loss of capability. A 70-billion-parameter model cut to half its depth, retaining nearly all of its measured knowledge performance, changes what can be served locally, on modest accelerators or in cost-sensitive deployments.
The deeper interest of the work is methodological. It shows a mature scientific framework, the statistical mechanics of disordered spin systems, being applied to a mundane engineering question: which parts of a neural network are safe to delete. The cross-disciplinary move is not new; physics tools have appeared throughout machine learning research. What is notable here is the specific payoff, that a one-time Hessian computation converts a search that seemed to require benchmarking everything into one where a cheap energy score ranks the candidates.
Whether this approach becomes standard practice depends on verification the community has not yet performed. The code is public, the claims are specific, and the benchmarks are standard. Those are the right conditions for an independent check, and until one arrives, the 23-point figure should be read as a claim, not a fact.
