More than 11,000 clicks from an integrated sampler
A Shanghai Jiao Tong University and TuringQ team reported on September 10 that its photonic processor registered as many as 11,059 detector clicks in a single one-millisecond sampling window. The version 1 preprint describes a Gaussian boson sampler called Zhiyuan 3.0 that combines spatial interferometers, electro-optic modulation and optical delay lines on a thin-film lithium-niobate chip. It operated at a reported clock rate of 4 gigahertz and fit with its source, detectors and controls inside one rack.
The number is impressive, but it needs exact language. The experiment counted detector clicks, which are the binary responses of its superconducting nanowire detectors across many optical modes. A click does not resolve how many photons reached the same detector in the same mode. The paper sometimes describes the result as exceeding 10,000 photons, yet the directly reported high-scale observable is a click count. Nor did the researchers calculate the probability of every possible 11,059-click output. The useful advance is the scale and integration of the sampling hardware, not full certification of its largest event.
Space and time expanded one chip
Gaussian boson sampling begins with squeezed light, a quantum optical state whose fluctuations are redistributed in a controlled way. An interferometer mixes that light among many modes, and detectors record where photons appear. Calculating the probabilities of large output patterns can become extremely demanding for classical computers. Earlier systems expanded by adding physical optical paths or by reusing a smaller network across successive time bins. Zhiyuan 3.0 combines both approaches.
Four spatial channels pass through networks of Mach-Zehnder interferometers, while delay loops cause light from different time bins to interact. Thermo-optic elements set relatively static phases, and faster electro-optic modulators change operations on short timescales. This creates a larger effective network without fabricating a separate path for every mode. The team reports a 67-gigahertz-plus modulation bandwidth, a half-wave voltage of 3.18 volts and temperature stabilization within 0.01 degrees Celsius. These component measurements support the 4-gigahertz operating rate, though they do not by themselves verify the complete output distribution.
Wafer processing is part of the result
The processor was fabricated on a commercial six-inch wafer carrying a 400-nanometer layer of crystalline lithium niobate above insulating oxide. The researchers patterned waveguides with deep-ultraviolet lithography, added titanium-nitride heaters and gold electrodes, then attached fiber arrays and wire bonds. They report an average coupling loss of 0.9 decibels per facet around 1,550 nanometers and interferometer extinction above 35 decibels.
That manufacturing route matters because large free-space optical experiments demand extensive alignment and can drift as their many components move. A monolithic circuit fixes more of the geometry and permits multiple devices to be made through the same wafer process. It does not make the entire system a single chip. The paper says squeezed-light generation and photon detection still include off-chip components, and it identifies integration of sources and detectors plus lower waveguide and coupling losses as future work. Rack-scale packaging is progress toward engineered hardware, not proof of inexpensive mass production.
Validation covered tractable slices
The authors checked the sampler at scales where detailed comparison remained possible. Across 32 time steps and 128 modes, measured photon correlations were compared with predictions for the intended squeezed state and three alternatives: coherent light, thermal light and a classical squashed-state model. Few-photon output distributions produced reported similarity coefficients of 0.9978 for single-photon events, 0.9962 for two-photon bunching and 0.9911 for two-photon non-bunching. A Bayesian test on a 30-photon subsystem favored the target Gaussian-boson-sampling model over the tested spoofers.
These checks show consistency with the intended behavior in selected regimes. They cannot establish every probability in the largest samples because the state space grows too quickly. Validation against several named alternatives also does not exclude every conceivable classical model. The experiment therefore combines direct small-scale agreement, subsystem tests and a computational-cost estimate. That is stronger than reporting a large count alone, but it remains different from independently reproducing the full high-scale distribution.
The advantage claim rests on a cost model
For its classical-simulation comparison, the team modeled a 40,000-mode problem containing 10,000 time steps. It assumed two decibels of loss per on-chip delay, two decibels in a fiber delay and five decibels of spatial loss. Using a matrix-product-state simulation framework, the authors calculated an effective photon number of 117.9 and estimated that an NVIDIA A100-class processor would require at least 9.7 million years to simulate a 2.5-microsecond sampling task at the specified truncation setting.
This is an author-reported complexity estimate, not a timed race in which the optical machine and a classical supercomputer completed identical jobs and their answers were compared. The conclusion depends on the assumed loss model, bond dimension, allowed truncation error and choice of classical method. Classical algorithms for boson sampling continue to improve, and loss can make some distributions easier to approximate. The paper offers evidence for computational intractability under its stated model, but the preprint has not been peer reviewed or independently replicated.
The same hardware predicted one fluid sequence
The researchers also reprogrammed the photonic network as a reservoir for forecasting pressure in a simulated Kármán vortex street. They compressed each 780-by-780 pressure field into five regional averages, encoded those values as optical squeezing inputs and used measured photonic features to predict the next five values. A classical decoder then reconstructed a two-dimensional field. Training used 1,500 snapshots, validation used 400 and the final 100 frames formed one chronological test segment.
On that segment, the photonic system produced a one-step mean-squared error of 0.006537. The mean for the largest tested classical echo-state network, with 320 nodes and 100 random initializations, was 0.006781 plus or minus 0.000091. The paper reports a 3.60 percent reduction relative to that mean. Its photonic readout used 255 trained coefficients, 84.1 percent fewer than the 1,605 trained readout coefficients counted for that classical model. Fixed coefficients and circuit settings were tabulated separately, while the encoder and decoder common to both approaches were excluded.
Integration is firmer than general AI advantage
The fluid result demonstrates that the chip can do more than emit hard-to-simulate samples. Its delay loops retain information from earlier inputs, interference transforms that history and photon counts provide features for a simple trained readout. Reconfiguration is valuable because specialized boson samplers have often struggled to connect scale with a useful task.
Still, one-step prediction on one prepared trajectory is not evidence that photonic hardware generally outperforms conventional machine learning. The comparison covered one echo-state-network family, used an encoder that reduced each field to five values and did not provide an end-to-end accounting of optical equipment, data acquisition, training time or energy. The 3.60 percent margin is much narrower than the headline sampling scale. The sound conclusion is that a wafer-fabricated space-time photonic network reached a notable operating regime and supported two distinct demonstrations. Whether it delivers reproducible computational or economic advantage beyond those tests remains open.
