Thousands of quantum calculations were deliberately thrown away before a single accepted result counted, yet that costly strategy allowed an experimental quantum computation to do something that has long been difficult to combine: perform a classically hard sampling task while also providing direct evidence that the final quantum state remained highly faithful despite the presence of hardware errors.
For years, one of the central challenges in quantum computing has not simply been carrying out calculations that overwhelm classical computers, but demonstrating that those calculations were actually performed correctly. As quantum circuits become larger and more complex, unavoidable hardware errors accumulate. At the same time, the very complexity that makes these computations difficult for classical machines also makes them increasingly difficult to verify.
The new research tackles both problems together by designing a family of quantum circuits that can detect many of their own errors while remaining theoretically hard for classical computers to simulate.
Rather than relying only on statistical benchmarks or assumptions about how hardware behaves, the approach uses the internal structure of the quantum circuit itself to place a lower bound on how faithfully the computation was executed.
Why hard quantum sampling has remained difficult to trust
Sampling problems have become one of the leading candidates for demonstrating quantum advantage—the point at which a quantum computer performs a task beyond the practical reach of classical supercomputers.
These experiments do not necessarily solve practical applications. Instead, they generate samples from complicated quantum probability distributions that, according to complexity theory, should be extraordinarily difficult for classical algorithms to reproduce.
But demonstrating computational difficulty alone is not enough.
According to the authors, a scalable demonstration of quantum advantage also requires increasing circuit size without allowing hardware noise to overwhelm the computation. Equally important, researchers need a practical way to verify that the quantum processor actually produced the intended quantum state with high fidelity.
Existing approaches have struggled to satisfy both requirements simultaneously.
Some sampling proposals map efficiently onto today’s hardware but are not well suited to error suppression through quantum encoding. Others possess useful mathematical structure but become difficult to implement on near-term quantum processors in the regimes where classical simulation is expected to be hardest.
Verification presents another obstacle. Directly measuring fidelity for these complicated quantum states often requires prohibitively many samples. Alternative benchmark methods are generally more efficient but depend on assumptions about noise that may not hold in realistic hardware.
The new work was designed specifically to bridge these gaps.
Building a difficult computation around an error-detecting code
The researchers based their protocol on what they call doped Clifford sampling.
The starting point is a highly entangling random Clifford circuit. Clifford circuits possess mathematical properties that make their output states efficient to characterize, allowing researchers to directly estimate state fidelity without making assumptions about the underlying hardware noise.
The team then encoded this circuit using spacetime codes.
Instead of correcting every possible error, the code detects many faults that occur during execution. Measurements of additional ancilla qubits reveal whether an error has been detected. Any computation producing a non-zero syndrome is discarded.
This means that many experimental runs never contribute to the final dataset. Although this dramatically lowers the effective sampling rate, it projects accepted runs into a much lower-error subspace, increasing the fidelity of the surviving quantum states.
The final ingredient introduces the computational complexity needed for quantum advantage.
The researchers inserted large numbers of T gates—non-Clifford operations that supply the “magic” required for universal quantum computation—but only at carefully selected locations where they preserve the structure of the error-detecting code.
Because these T gates commute with the measured stabilizers, the same syndrome measurements remain valid before and after the circuit is doped.
That property becomes essential for verifying the final computation.
A circuit that becomes harder without losing its certificate
The researchers began with a Clifford circuit whose fidelity could be measured efficiently using direct fidelity estimation.
They then transformed this circuit into a much harder computational problem by inserting T gates while preserving the code’s stabilizers.
Although the final doped circuit cannot itself be directly verified through the same efficient methods, the relationship between the original and doped circuits allows the researchers to derive a lower bound on the fidelity of the harder computation.
Their analysis shows that both circuits are projected into the same encoded subspace after syndrome post-selection.
The principal way fidelity could decrease after doping is if faults that were previously harmless become harmful because of the added T gates.
The authors argue that such harmless faults become asymptotically negligible in random Clifford circuits and can be efficiently estimated for a specific circuit instance.
Importantly, the protocol assumes a general Pauli noise model, which the authors state can be obtained through randomized gate twirling. Unlike several existing fidelity proxy methods, it does not require weak, spatially uniform, temporally independent, or gate-independent noise assumptions.
A 97-qubit demonstration
To test the protocol experimentally, the researchers implemented it on a superconducting quantum processor.
The logical computation consisted of a 70-qubit, depth-70 Clifford circuit encoded across a total of 97 physical qubits using 27 ancilla qubits dedicated to syndrome measurements.
The encoded circuit was heavily modified with 468 T gates while preserving the code structure.
The complete physical circuit contained 2,869 controlled-Z (CZ) gates. Of these, 2,415 belonged to the depth-70 computation itself, while another 454 implemented syndrome extraction.
The researchers optimized several aspects of the experiment to reduce hardware noise. They selected qubits with relatively low error rates and favorable ancilla connectivity, optimized the spacetime Pauli checks, calibrated gates and readout for the specific circuit, discarded runs exhibiting detected non-Markovian errors, and applied Pauli twirling to tailor the noise toward a stochastic Pauli channel.
Rejecting most computations to improve the remaining ones
The error-detecting strategy comes with an obvious cost.
Only runs with zero measured syndrome are accepted.
For the encoded Clifford circuit, the measured post-selection rate was approximately 5.90 × 10⁻⁴. In other words, only a very small fraction of all experimental executions survived the filtering process.
The tradeoff, however, was substantially improved fidelity.
Using direct fidelity estimation based on 80 randomly selected stabilizers and 250,000 measurement shots for each stabilizer, the researchers measured a fidelity of 0.32(1) for the post-selected stabilizer state.
After applying a readout mitigation strategy, they estimated the corresponding state fidelity to be 0.57(2).
The researchers also report that encoding increased state fidelity by a factor of 29 compared with the unencoded circuit, although this improvement came alongside an approximately 860-fold reduction in effective sampling rate.
They note that this slowdown remains practical for superconducting quantum hardware.
Estimating the fidelity of the harder computation
The central challenge was determining the fidelity of the final T-doped circuit.
Because direct fidelity estimation is no longer efficient once large numbers of T gates are introduced, the researchers instead estimated the maximum possible fidelity loss caused by doping.
They performed Monte Carlo simulations under various Pauli noise strengths and polarizations, classifying simulated faults according to whether they were detected by the code and whether they stabilized the Clifford circuit.
From these simulations they estimated that doping could reduce fidelity by at most 0.013(1).
Combining this estimate with the experimentally measured Clifford fidelity allowed them to derive a conservative lower bound.
With 95% confidence, they concluded that the fidelity of the classically hard doped quantum state was at least 0.284.
The syndrome acceptance probabilities measured before and after doping were statistically indistinguishable, consistent with the expectation that adding virtual T gates did not itself introduce measurable additional noise.
Testing whether the bound really holds
The researchers did not rely solely on the final large-scale experiment.
They also designed several independent validation tests covering different amounts and types of doping.
With only five inserted T gates, direct fidelity estimation remained computationally tractable because the quantum state still possessed a sparse Pauli description. Measurements in this low-magic regime agreed with the predicted fidelity bound.
They also replaced T gates with S gates in additional validation experiments. Since S gates preserve the Clifford structure, direct fidelity estimation could still be applied, again producing results consistent with the predicted lower bound.
For intermediate circuits containing roughly 75 T gates, the researchers turned to cross-entropy benchmarking.
According to the paper, circuits in this regime begin to exhibit output distributions whose statistical moments approach those expected for Haar-random quantum states, making cross-entropy benchmarking a calibrated fidelity proxy under the stated assumptions. Although this validation required considerably more classical computation and stronger assumptions than the primary protocol, it again supported the fidelity bound.
Across all validation experiments, the syndrome probability distributions remained highly consistent regardless of doping strategy.
Making the computation difficult for classical algorithms
The researchers also analyzed whether existing classical simulation techniques could realistically reproduce the demonstrated computation.
They argue that the circuit combines two independent sources of computational difficulty.
Its deep Clifford backbone generates substantial entanglement, while the 468 T gates contribute a large amount of quantum “magic.”
The paper discusses several leading simulation strategies—including tensor-network methods, matrix product state approaches, stabilizer-based techniques, hybrid algorithms, and noisy approximations—and explains why each faces substantial challenges for the demonstrated circuit.
The authors emphasize that improvements in classical algorithms remain possible and explicitly leave open the possibility that future simulation methods could perform better than those available today.
Nevertheless, they report that extensive numerical investigations found current leading approaches intractable for the experimental instance.
They also estimate that syndrome post-selection effectively reduced the logical controlled-Z gate error rate to approximately 1.8 × 10⁻⁴, representing about a tenfold improvement over the underlying physical gate error rate.
What remains unresolved
The work does not claim to solve all remaining challenges for scalable quantum advantage.
Error detection through post-selection improves fidelity but cannot scale indefinitely because rejecting increasing numbers of experimental runs eventually becomes impractical.
The authors suggest that as hardware improves, related spacetime codes or other coding strategies could eventually support fault-tolerant demonstrations rather than relying solely on post-selection.
Verification also remains incomplete.
Although the protocol substantially weakens the noise assumptions required to certify fidelity, it is still device dependent. The authors distinguish this from fully device-independent verification, which would require verifying only the computational output itself without making quantum-specific assumptions about the hardware.
They note that recent efforts toward such verification have focused on embedding hidden secrets within quantum circuits that classical verifiers can later check.
For now, the researchers present their approach as a way to bring together three goals that have often remained separate: constructing sampling problems with strong theoretical hardness, suppressing hardware errors through encoding, and experimentally certifying the fidelity of computations that can no longer be directly simulated.
Publication details
Simon Martiel et al, Sampling hard circuits with verifiably high fidelity, arXiv (2026). DOI: 10.48550/arxiv.2607.25941






