Verified Quantum Advantage Record - empirical scout thread
Date checked: 2026-09-21 [verified-at-source: user-supplied session date]. Scope: empirical, source-fetched claims about quantum advantage where verification/checkability matters.
Bottom line
Random circuit sampling record, not soundly self-verifying: the strongest current public RCS claim I found remains Google's Willow demonstration: 103 qubits [verified-at-source], depth 40 [verified-at-source], XEB fidelity = 0.1% [verified-at-source], quantum runtime 5 minutes [verified-at-source], claimed classical runtime 10^25 years [verified-at-source] on a Frontier/exascale-class comparison [verified-at-source]. Source quote from Google's Willow spec sheet: "Application performance 103 qubits, depth 40, XEB fidelity = 0.1% Estimated time on Willow vs. classical supercomputer 5 minutes vs. 10^25 years" (Google Willow spec sheet PDF, fetched 2026-09-21). Google's public blog says Willow "performed a standard benchmark computation in under five minutes that would take one of today's fastest supercomputers 10 septillion (that is, 10^25) years" and identifies the benchmark: "As a measure of Willow's performance, we used the random circuit sampling (RCS) benchmark." Source: https://blog.google/innovation-and-ai/technology/research/google-willow-quantum-chip/
I found no newer public 2025/2026 RCS demonstration that overtakes Willow by claimed classical runtime. I did find a later peer-reviewed superconducting RCS result, Zuchongzhi 3.0: 105-qubit processor [verified-at-source], 83-qubit [verified-at-source], 32-cycle [verified-at-source] RCS, 4.1 x 10^8 collected bitstrings [verified-at-source], XEB/fidelity 0.025% [verified-at-source], 8.4 x 10^33 FLOPs for 1,000,000 noisy classical samples [verified-at-source], claimed 6.4 x 10^9 years on Frontier [verified-at-source]. That is newer/peer-reviewed but below Willow's 10^25-year claim. Source quote: "Our experiments with an 83-qubit, 32-cycle random circuit sampling on Zuchongzhi 3.0 highlight its superior performance, achieving one million samples in just a few hundred seconds." Source quote: "Frontier, which would require approximately 6.4 x 10^9 years to replicate the task." Source: arXiv:2412.11924 / Phys. Rev. Lett. 134, 090601 (2025), DOI 10.1103/PhysRevLett.134.090601 [verified-at-source via APS/arXiv fetch].
The exact Nature citation in the prompt is not the Willow RCS paper. Nature 638, pages 920-926 (2025) [verified-at-source] is "Quantum error correction below the surface code threshold," DOI 10.1038/s41586-024-08449-y [verified-at-source]. It reports surface-code memory on Willow, not the RCS benchmark.
Verified/certified advantage record is different from RCS advantage. The strongest actually-checkable result I found is the JPMorganChase/Quantinuum certified-randomness experiment: 71,313 bits of certified entropy [verified-at-source], 71,273 extracted bits [verified-at-source], 56 x 30,010 raw bits [verified-at-source], 30,010 valid samples [verified-at-source], 60,952 circuits submitted [verified-at-source], 1.1 exaFLOPS combined classical verification [verified-at-source], and a restricted adversarial/security model [verified-at-source]. Source quote: "This type of protocol allows a classical client to verify randomness using only remote access to an untrusted quantum server." Source quote: "Frontier and Summit were used at full-machine scale" and achieved "a combined performance of 1.1 exaFLOPS." Source: Minzhao Liu et al., "Certified randomness using a trapped-ion quantum processor," Nature 640, 343-348 (2025), DOI 10.1038/s41586-025-08737-1 [verified-at-source].
Shor's algorithm on hardware is still tiny. The largest Shor-family hardware factoring result I verified is N = 21 [verified-at-source], but the source itself calls it a "two-photon compiled algorithm." If "genuine, non-cheating" excludes answer-dependent or problem-specific compilation, I cannot certify any larger-than-toy hardware factoring result as a clean Shor run. The 143 [verified-at-source] and 56153 [verified-at-source] claims are adiabatic/SAT-style factoring, not circuit Shor. Smolin, Smith & Vargo's criticism remains the key warning: "Previous experimental implementations have used simplifications dependent on knowing the factors in advance" and "Valid implementations should not make use of the answer sought." Source: Nature 499, 163-165 (2013), DOI 10.1038/nature12290 [verified-at-source].
1. Current RCS record
Google Willow
Source: Google Quantum AI public blog, 2024-12-09 [verified-at-source], updated 2025-06-12 [verified-at-source], and Google Willow spec sheet PDF fetched from google/quantumai.
Numbers:
- Qubits: 103 [verified-at-source].
- Circuit depth: 40 [verified-at-source].
- XEB fidelity: 0.1% [verified-at-source].
- Quantum runtime: 5 minutes [verified-at-source].
- Claimed classical runtime: 10^25 years [verified-at-source].
- Classical hardware comparison: "one of today's fastest supercomputers" [verified-at-source]; Google blog notes estimates involving Frontier and says: "we assumed full access to secondary storage, i.e., hard drives, without any bandwidth overhead -- a generous and unrealistic allowance for Frontier." [verified-at-source]
- Method/cost caveat: Google says: "Computational costs are heavily influenced by available memory. Our estimates therefore consider a range of scenarios, from an ideal situation with unlimited memory ... to a more practical, embarrassingly parallelizable implementation on GPUs." [verified-at-source]
Independent context from Scott Aaronson, 2024-12 [verified-at-source]: "Google has also announced a new quantum supremacy experiment on its 105-qubit chip, based on Random Circuit Sampling with 40 layers of gates." Aaronson's cost summary: "if you use the best currently-known simulation algorithms (based on Johnnie Gray's optimized tensor network contraction), as well as an exascale supercomputer, their new experiment would take ~300 million years to simulate classically if memory is not an issue, or ~10^25 years if memory is an issue" [all numbers verified-at-source]. His verification caveat is direct: "for the exact same reason why ... this quantum computation would take ~10^25 years for a classical computer to simulate, it would also take ~10^25 years for a classical computer to directly verify the quantum computer's results!!" and "all validation of Google's new supremacy experiment is indirect, based on extrapolations from smaller circuits." Source: https://scottaaronson.blog/?p=8525
Verdict: Willow is the current public RCS record by claimed classical runtime that I found. It is not directly verified from inside by recomputing probabilities for the full circuit; validation is indirect/extrapolative.
Later/nearby RCS result: Zuchongzhi 3.0
Source: Dongxin Gao, Daojin Fan, Chen Zha, Jiahao Bei, et al., "Establishing a New Benchmark in Quantum Computational Advantage with 105-qubit Zuchongzhi 3.0 Processor," arXiv:2412.11924 [verified-at-source], Phys. Rev. Lett. 134, 090601 (2025) [verified-at-source], DOI 10.1103/PhysRevLett.134.090601 [verified-at-source].
Numbers and quotes:
- Processor: "105 qubits" [verified-at-source]. Quote: "This superconducting quantum computer prototype, comprising 105 qubits, achieves high operational fidelities, with single-qubit gates, two-qubit gates, and readout fidelity at 99.90%, 99.62% and 99.18%, respectively." [all numbers verified-at-source]
- RCS circuit: "83-qubit, 32-cycle" [verified-at-source]. Quote: "Our experiments with an 83-qubit, 32-cycle random circuit sampling on Zuchongzhi 3.0 highlight its superior performance, achieving one million samples in just a few hundred seconds." [all numbers verified-at-source]
- Samples: "For the largest full circuit featuring 83 qubits and 32 cycles, we have collected a total of approximately 4.1 x 10^8 bitstrings." [verified-at-source]
- XEB/fidelity: "fidelity of 0.025%" [verified-at-source].
- Classical cost: "The estimated number of floating-point operations required to generate a million uncorrelated bitstrings with a fidelity of 0.025% from an 83-qubit, 32-cycle random circuit using a classical computer is 8.4 x 10^33." [all numbers verified-at-source]
- Frontier runtime: table reports "8.4 x 10^33" FLOPs and "6.4 x 10^9 yr" under the 9.2 PB memory case, and "7.5 x 10^31" FLOPs and "5.7 x 10^7 yr" under a 762.2 PB storage-assisted case [all numbers verified-at-source]. The paper states: "The Frontier supercomputer boasts a theoretical peak performance of 1.685 x 10^18 FLOPS. In our estimations, we presume a 20% FLOP efficiency" [all numbers verified-at-source].
Verdict: important later RCS result, but not a larger claimed gap than Willow.
Post-Willow search status
Searches for 2025/2026 [verified-at-source: web searches run 2026-09-21] RCS hardware records found no public Google successor or non-Google result exceeding Willow's 103-qubit/depth-40/10^25-year claim. This is a null search result, not proof of nonexistence.
2. How weak is XEB as verification?
XEB is not a sound certificate by itself
Aaronson & Gunn, "On the Classical Hardness of Spoofing Linear Cross-Entropy Benchmarking," arXiv:1910.12085 [verified-at-source], arXiv DOI 10.48550/arXiv.1910.12085 [verified-at-source]. Key quote: "This raises a theoretical question: how hard is it for a classical computer to spoof the results of the Linear XEB test?" Their positive result is conditional, not an unconditional certificate: "we show that the problem is classically hard, assuming that there is no efficient classical algorithm that... estimates the probability of C outputting a specific output string... with variance even slightly better than that of the trivial estimator" [verified-at-source].
Gao, Kalinowski, Chou, Lukin, Barak & Choi, "Limitations of Linear Cross-Entropy as a Measure for Quantum Advantage," arXiv:2112.01657 [verified-at-source], PRX Quantum 5, 010334 (2024) [verified-at-source], DOI 10.1103/PRXQuantum.5.010334 [verified-at-source]. Quotes: "achieving relatively high XEB values does not imply faithful simulation of quantum dynamics"; their classical algorithm "with 1 GPU within 2s, yields high XEB values, namely 2-12% of those obtained in experiments" [all numbers verified-at-source]; and "the XEB alone has limited utility as a benchmark for quantum advantage." [verified-at-source]
Classical XEB/spoofing scorecard
Original Sycamore 53-qubit/20-cycle instance [verified-at-source]: Pan Zhang et al., "Solving the sampling problem of the Sycamore quantum circuits," arXiv:2111.03011 [verified-at-source], Phys. Rev. Lett. 129, 090502 (2022) [verified-at-source], DOI 10.1103/PhysRevLett.129.090502 [verified-at-source]. Quote: "For the Sycamore quantum supremacy circuit with 53 qubits and 20 cycles, we have generated one million uncorrelated bitstrings ... approximate state has fidelity F~0.0037. The whole computation has cost about 15 hours on a computational cluster with 512 GPUs." [all numbers verified-at-source] This is the largest classical XEB/fidelity I verified for the full 53-qubit/20-cycle Sycamore instance: 0.0037, i.e. 0.37% [verified-at-source].
Same Sycamore class, bounded-fidelity classical sampling: Kalachev et al., "Classical Sampling of Random Quantum Circuits with Bounded Fidelity," arXiv:2112.15083 [verified-at-source]. Quote: "classically produced 1 million samples with the fidelity bounded by 0.2%, based on the 20-cycle circuit of the Sycamore 53-qubit quantum chip" and "took about 14.5 days ... 32 GPUs" [all numbers verified-at-source].
Faster Sycamore replay: "Leapfrogging Sycamore," National Science Review/PMC source fetched [verified-at-source]. Quote: "using 1432 GPUs to simulate quantum random circuit sampling that generates uncorrelated samples with a higher linear cross-entropy score and is 7 faster than the Sycamore 53-qubit experiment" [all numbers verified-at-source]. The paper states Google obtained "one (three) million uncorrelated samples in 200 (600) s, with a linear cross-entropy (XEB) of 0.2%" and reports "1432 NVIDIA A100 GPUs" producing "three million uncorrelated samples" in "86.4 s" with "13.7 kWh" [all numbers verified-at-source].
Energetic-superiority preprint: Fu, Su, Zhong, Zhang, Pan Zhang, Jian-Wei Pan et al., arXiv:2407.00769 [verified-at-source]. Quote: "we have achieved a time-to-solution of 14.22 seconds with energy consumption of 2.39 kWh which achieved fidelity of 0.002 and our most remarkable result is a time-to-solution of 17.18 seconds, with energy consumption of only 0.29 kWh which achieved a XEB of 0.002 after post-processing" [all numbers verified-at-source].
Largest raw XEB score I found, but on an easier shallower instance: arXiv:2512.07311 [verified-at-source] reports "simulating the 53-qubit, 14-cycle Sycamore circuit and achieving a linear cross-entropy benchmarking (XEB) score of 0.549, exceeding the published XEB score of 0.002 from Google's reference data" [all numbers verified-at-source]. This is not comparable to the 53-qubit/20-cycle Sycamore supremacy instance.
2026 theoretical cat-and-mouse claim: arXiv:2607.04054 [verified-at-source], "Frozen-Tree Sampling Refutes Quantum Advantage of Random Circuit Sampling." Quotes: it "draws bitstrings of n qubits in O(n) time per sample" and "no statistical test acting on samples alone can distinguish the classical frozen-tree sampler from a quantum random circuit" [verified-at-source]. This is a 2026 preprint/theoretical claim; I do not treat it as a settled empirical overthrow of Willow.
Verdict on XEB: classical simulation/spoofing overtook the original Sycamore benchmark. For Willow-scale RCS, quantum hardware remains ahead by published cost estimates, but the verification gap is exactly the problem: full XEB verification for the record instance is classically out of reach, so the full claim rests on extrapolated validation, model trust, and anti-spoofing assumptions rather than a compact sound certificate.
3. Certified / verifiable quantum advantage
Aaronson-Hung protocol
Source: Scott Aaronson & Shih-Han Hung, "Certified Randomness from Quantum Supremacy," arXiv:2303.01625 [verified-at-source], arXiv DOI 10.48550/arXiv.2303.01625 [verified-at-source], STOC 2023 [verified-at-source], ACM DOI 10.1145/3564246.3585145 [verified-at-source via ACM/search result], pages 933-944 [verified-at-source via Nature reference].
Quotes: the paper studies "generating cryptographically certified random bits" [verified-at-source] and argues that when RCS outputs pass LXEB, "under plausible hardness assumptions they necessarily contain Omega(n) min-entropy" [verified-at-source]. Its bottleneck is explicit: "Currently, the central drawback of our protocol is the exponential cost of verification, which in practice will limit its implementation to at most n~60 qubits" [verified-at-source].
JPMorganChase / Quantinuum / ORNL certified randomness
Source: Minzhao Liu, Ruslan Shaydulin, Pradeep Niroula, Matthew DeCross, Shih-Han Hung, Scott Aaronson, Marco Pistoia et al., "Certified randomness using a trapped-ion quantum processor," Nature 640, 343-348 (2025) [verified-at-source], DOI 10.1038/s41586-025-08737-1 [verified-at-source].
Exact numbers and quotes:
- Device: Quantinuum H2-1 trapped-ion processor [verified-at-source]. Quote: "We demonstrate our protocol using the Quantinuum H2-1 trapped-ion quantum processor accessed remotely over the Internet." [verified-at-source]
- Qubit/raw-bit size: 56 x 30,010 raw bits [verified-at-source], implying 56 measured bits per valid sample [verified-at-source]. Quote: "feed the 56 x 30,010 raw bits into a Toeplitz randomness extractor and extract 71,273 bits." [all numbers verified-at-source]
- Challenge circuits: "fixed arrangement of 10 layers of entangling UZZ gates, each sandwiched between layers of pseudorandomly generated SU(2) gates on all qubits" [numbers verified-at-source].
- Thresholds and response timing: expected fidelity "phi >= 0.3 or better on depth-10 circuits" [verified-at-source]; t_threshold = 2.2 s [verified-at-source]; chi = 0.3 [verified-at-source]; quantum device time t_QC = 2.154 s per sample [verified-at-source].
- Run size: b = 15 and b = 20 [verified-at-source]; 1,993 batches [verified-at-source]; 60,952 circuits [verified-at-source]; M = 30,010 valid samples [verified-at-source]; 984 successful batches [verified-at-source]; cumulative device time 64,652 s [verified-at-source].
- Verification cost: exact simulation time "100.3 s per circuit when using the entire [Frontier] supercomputer at a numerical efficiency of 45%" [all numbers verified-at-source]. "Frontier and Summit were used at full-machine scale" with "sustained peak performance of 897 petaFLOPS and 228 petaFLOPS" and "a combined performance of 1.1 exaFLOPS" [all numbers verified-at-source].
- Verification sample: XEB score for m = 1,522 circuit-sample pairs [verified-at-source], XEB_test = 0.32 [verified-at-source].
- Certified output: "at epsilon_sou=10^-6, we have Q_min=1,297, corresponding to H_min=71,313 against an adversary four times more powerful than Frontier" [all numbers verified-at-source]. Extracted bits: 71,273 [verified-at-source]. Input randomness: "only 32 bits" [verified-at-source].
- Security caveat: source says security is in a "restricted adversarial model" [verified-at-source]. Aaronson's public commentary says "about 70,000 certified random bits were generated over 18 hours" [verified-at-source] and that the parameters are "not yet good enough for my and Shih-Han's formal security reduction" but support "practical security" [verified-at-source].
Verdict: this is the cleanest empirical example of classically checkable/certified quantum advantage I found. It certifies randomness/entropy under assumptions and a restricted adversarial model; it is not universal delegated quantum computation.
BCMVV / Mahadev-style classically verifiable proofs
Source: Brakerski, Christiano, Mahadev, Vazirani & Vidick, "A Cryptographic Test of Quantumness and Certifiable Randomness from a Single Quantum Device," arXiv:1804.00640 [verified-at-source], arXiv DOI 10.48550/arXiv.1804.00640 [verified-at-source]. Quote: it gives "a protocol for efficient classical verification that the untrusted device is 'truly quantum,' and a protocol for producing certifiable randomness from a single untrusted quantum device" [verified-at-source]. Assumption: post-quantum/LWE-style trapdoor claw-free functions [verified-at-source].
Source: Urmila Mahadev, "Classical Verification of Quantum Computations," arXiv:1804.01082 [verified-at-source], arXiv DOI 10.48550/arXiv.1804.01082 [verified-at-source]. Quote: "We present the first protocol allowing a classical computer to interactively verify the result of an efficient quantum computation" [verified-at-source]. Soundness depends on "the assumption that the learning with errors problem is computationally intractable for efficient quantum machines" [verified-at-source].
Hardware status: I found a proof-of-principle, not a full-scale sound implementation. Roman Stricker et al., "Towards experimental classical verification of quantum computation," arXiv:2203.07395 [verified-at-source], Quantum Science and Technology DOI 10.1088/2058-9565/ad2986 [verified-at-source via IOP search result]. Quote: "first, proof-of-principle experiment a verification protocol using only classical means on a small trapped-ion quantum processor" [verified-at-source]. Quote: "We implement the protocol experimentally on an eight-qubit trapped-ion quantum processor" [verified-at-source]. The authors explicitly flag relaxed/security-limited status: "We show how to verify existing quantum processors under relaxed security constraints, while the most stringent variant of such a protocol remains too demanding for current quantum hardware" [verified-at-source]. They also say the full protocol requires a "very large range" of trapdoor functions and "many auxiliary qubits" and is "not feasible on current devices" [verified-at-source].
Related efficiently verifiable advantage proposal: Kahanamoku-Meyer, Choi, Vazirani & Yao, "Classically verifiable quantum advantage from a computational Bell test," Nature Physics 18, 918-924 (2022) [verified-at-source], DOI 10.1038/s41567-022-01643-7 [verified-at-source]. Quote: "Sampling-based protocols... correctness ... is exponentially difficult to verify" [verified-at-source]. Their proposed test has quantum success "~85%" and classical bound "75%" [verified-at-source], but needs about "10^3 qubits" and "gate depth ~10^5" [verified-at-source]. I found no hardware run at that scale.
Verdict: cryptographic verification exists theoretically; hardware has proof-of-principle demonstrations and the 2025 certified-randomness result, but not general Mahadev verification of a useful quantum computation at scale.
4. Shor's algorithm on real hardware
Largest Shor-family factoring result I verified
Source: Anthony Laing, Thomas Lawson, Roberto Alvarez, Xiao-Qi Zhou, Jeremy L. O'Brien et al., "Experimental realization of Shor's quantum factoring algorithm using qubit recycling," Nature Photonics 6, 773-776 (2012) [verified-at-source], DOI 10.1038/nphoton.2012.259 [verified-at-source].
Quote: "Encoding the work register in higher-dimensional states, we implement a two-photon compiled algorithm to factor N = 21." [verified-at-source] The article says this followed "four small-scale demonstrations" [verified-at-source].
Interpretation: N = 21 [verified-at-source] is the largest Shor-family hardware factorization I verified, but the word "compiled" matters. It is not evidence of cryptographically meaningful Shor scaling.
Factor 15 baseline
Source: IBM retrospective and Nature paper page for Vandersypen et al., "Experimental realization of Shor's quantum factoring algorithm using nuclear magnetic resonance," Nature 414, 883-887 (2001) [verified-at-source], DOI 10.1038/414883a [verified-at-source]. IBM quote: the 2001 experiment "successfully factor[ed] the number 15" [verified-at-source]. IBM also quotes Isaac Chuang: "There are different ways to write the factor 15 algorithm, and if you simplify it sufficiently, well, then you can run it, but it's not terribly meaningful" [verified-at-source]. Vandersypen quote from IBM: "we still have a long way to go from toy problems to relevant applications. After all, we knew what the answer would be when we set out to factor the number 15." [verified-at-source]
Why 143 and 56153 do not count as Shor records
Source: Nanyang Xu, Jing Zhu, Dawei Lu, Xianyi Zhou, Xinhua Peng & Jiangfeng Du, "Quantum Factorization of 143 on a Dipolar-Coupling Nuclear Magnetic Resonance System," Phys. Rev. Lett. 108, 130501 (2012) [verified-at-source], DOI 10.1103/PhysRevLett.108.130501 [verified-at-source], arXiv:1111.3726 [verified-at-source]. Quote: "Adiabatic quantum computation for this is an alternative approach other than Shor's algorithm" and they report factoring "the number 143" [verified-at-source]. Therefore it is not a Shor record.
Source: Nikesh S. Dattani & Nathaniel Bryans, "Quantum factorization of 56153 with only 4 qubits," arXiv:1411.6758 [verified-at-source], arXiv DOI 10.48550/arXiv.1411.6758 [verified-at-source]. Quote: the 143 computation "actually also factored much larger numbers such as 3599, 11663, and 56153" [all numbers verified-at-source], but the same abstract distinguishes it from Shor: "unlike the implementations of Shor's algorithm performed thus far" [verified-at-source], and concedes: "because they only use 4 qubits, these factorizations can also be performed trivially on classical computers" [verified-at-source].
Source: Sebastian Verschoor, "Factoring semi-primes with (quantum) SAT-solvers," arXiv:1902.01448 [verified-at-source], arXiv DOI 10.48550/arXiv.1902.01448 [verified-at-source]. Quote: "Shor's quantum factoring algorithm factors any integer in polynomial time, although large-scale fault-tolerant quantum computers capable of implementing Shor's algorithm are not yet available, so relevant benchmarking experiments for factoring via Shor's algorithm are not yet possible." [verified-at-source] On quantum/SAT/annealing factoring: "We find no evidence that this is a viable path toward factoring large numbers" [verified-at-source].
Source: John A. Smolin, Graeme Smith & Alexander Vargo, "Oversimplifying quantum factoring," Nature 499, 163-165 (2013) [verified-at-source], DOI 10.1038/nature12290 [verified-at-source]. Quote: "Previous experimental implementations have used simplifications dependent on knowing the factors in advance." Quote: "all composite numbers admit simplification of the algorithm to a circuit equivalent to flipping coins." Quote: "The difficulty of a particular experiment therefore depends on the level of simplification chosen, not the size of the number factored." Quote: "Valid implementations should not make use of the answer sought." [all verified-at-source]
Verdict: as of the sources checked on 2026-09-21 [verified-at-source: search date], nothing has changed in the meaningful Shor record. N = 21 is the largest Shor-family hardware demo I verified; no larger clean, non-answer-dependent Shor run was found.
5. Cost numbers for the current RCS record
Willow cost numbers found
- Quantum runtime: 5 minutes [verified-at-source].
- Classical runtime: 10^25 years [verified-at-source].
- Alternative independent estimate: ~300 million years if memory is not an issue [verified-at-source], ~10^25 years if memory is an issue [verified-at-source].
- Classical method named by Aaronson: "Johnnie Gray's optimized tensor network contraction" [verified-at-source].
- Hardware class: exascale/Frontier-class supercomputer [verified-at-source].
- Raw FLOP count: not found in the fetched Google/Aaronson sources [not verified].
- Dollar cost: not found in the fetched sources [not verified].
- Core-hours: not found in the fetched sources [not verified].
The best exact cost statement I can support for Willow is the public runtime claim and memory-conditional tensor-network estimate, not a raw FLOP bill.
Exact RCS cost numbers from Zuchongzhi 3.0, for comparison
- Method: tensor-network classical simulation/contraction [verified-at-source from paper context].
- Full circuit: 83 qubits [verified-at-source], 32 cycles [verified-at-source], fidelity 0.025% [verified-at-source], 1,000,000 samples [verified-at-source].
- Cost: 8.4 x 10^33 FLOPs [verified-at-source], 6.4 x 10^9 years on Frontier with 9.2 PB memory [all verified-at-source].
- More storage-rich scenario: 7.5 x 10^31 FLOPs [verified-at-source], 5.7 x 10^7 years with 762.2 PB storage [all verified-at-source].
- Frontier assumption: 1.685 x 10^18 FLOPS peak and 20% efficiency [verified-at-source].
Exact verification cost from certified randomness, for contrast
- Verification simulations: 100.3 s per circuit on full Frontier at 45% numerical efficiency [all verified-at-source].
- Combined verification supercomputing: Frontier + Summit at 897 petaFLOPS + 228 petaFLOPS = 1.1 exaFLOPS [all verified-at-source].
- Verified sample subset: 1,522 circuit-sample pairs [verified-at-source].
- Certified entropy: 71,313 bits [verified-at-source]; extracted randomness: 71,273 bits [verified-at-source].
Evidence class map
- Established: Willow public RCS benchmark numbers as a claimed hardware/classical-runtime demonstration; Zuchongzhi 3.0 peer-reviewed RCS numbers; Sycamore classical replay/spoofing numbers; Quantinuum/JPMC certified-randomness numbers; Shor N = 21 compiled record and Smolin et al. critique.
- Serious speculation / assumption-dependent: Aaronson-Hung security reduction; Aaronson-Gunn spoofing hardness; Willow's full-scale extrapolated verification; 2026 frozen-tree preprint.
- Null result: no post-Willow 2025/2026 public RCS record exceeding Willow was found in searches run 2026-09-21.
- Not verified: Willow raw FLOPs/core-hours/dollars; a full-scale hardware implementation of Mahadev's classical verification protocol; any clean larger-than-21 Shor run.
View exactly as delivered (raw text)
# Verified Quantum Advantage Record - empirical scout thread
Date checked: 2026-09-21 [verified-at-source: user-supplied session date]. Scope: empirical, source-fetched claims about quantum advantage where verification/checkability matters.
## Bottom line
1. **Random circuit sampling record, not soundly self-verifying:** the strongest current public RCS claim I found remains Google's Willow demonstration: 103 qubits [verified-at-source], depth 40 [verified-at-source], XEB fidelity = 0.1% [verified-at-source], quantum runtime 5 minutes [verified-at-source], claimed classical runtime 10^25 years [verified-at-source] on a Frontier/exascale-class comparison [verified-at-source]. Source quote from Google's Willow spec sheet: "Application performance 103 qubits, depth 40, XEB fidelity = 0.1% Estimated time on Willow vs. classical supercomputer 5 minutes vs. 10^25 years" (Google Willow spec sheet PDF, fetched 2026-09-21). Google's public blog says Willow "performed a standard benchmark computation in under five minutes that would take one of today's fastest supercomputers 10 septillion (that is, 10^25) years" and identifies the benchmark: "As a measure of Willow's performance, we used the random circuit sampling (RCS) benchmark." Source: https://blog.google/innovation-and-ai/technology/research/google-willow-quantum-chip/
2. **I found no newer public 2025/2026 RCS demonstration that overtakes Willow by claimed classical runtime.** I did find a later peer-reviewed superconducting RCS result, Zuchongzhi 3.0: 105-qubit processor [verified-at-source], 83-qubit [verified-at-source], 32-cycle [verified-at-source] RCS, 4.1 x 10^8 collected bitstrings [verified-at-source], XEB/fidelity 0.025% [verified-at-source], 8.4 x 10^33 FLOPs for 1,000,000 noisy classical samples [verified-at-source], claimed 6.4 x 10^9 years on Frontier [verified-at-source]. That is newer/peer-reviewed but below Willow's 10^25-year claim. Source quote: "Our experiments with an 83-qubit, 32-cycle random circuit sampling on Zuchongzhi 3.0 highlight its superior performance, achieving one million samples in just a few hundred seconds." Source quote: "Frontier, which would require approximately 6.4 x 10^9 years to replicate the task." Source: arXiv:2412.11924 / Phys. Rev. Lett. 134, 090601 (2025), DOI 10.1103/PhysRevLett.134.090601 [verified-at-source via APS/arXiv fetch].
3. **The exact Nature citation in the prompt is not the Willow RCS paper.** Nature 638, pages 920-926 (2025) [verified-at-source] is "Quantum error correction below the surface code threshold," DOI 10.1038/s41586-024-08449-y [verified-at-source]. It reports surface-code memory on Willow, not the RCS benchmark.
4. **Verified/certified advantage record is different from RCS advantage.** The strongest actually-checkable result I found is the JPMorganChase/Quantinuum certified-randomness experiment: 71,313 bits of certified entropy [verified-at-source], 71,273 extracted bits [verified-at-source], 56 x 30,010 raw bits [verified-at-source], 30,010 valid samples [verified-at-source], 60,952 circuits submitted [verified-at-source], 1.1 exaFLOPS combined classical verification [verified-at-source], and a restricted adversarial/security model [verified-at-source]. Source quote: "This type of protocol allows a classical client to verify randomness using only remote access to an untrusted quantum server." Source quote: "Frontier and Summit were used at full-machine scale" and achieved "a combined performance of 1.1 exaFLOPS." Source: Minzhao Liu et al., "Certified randomness using a trapped-ion quantum processor," Nature 640, 343-348 (2025), DOI 10.1038/s41586-025-08737-1 [verified-at-source].
5. **Shor's algorithm on hardware is still tiny.** The largest Shor-family hardware factoring result I verified is N = 21 [verified-at-source], but the source itself calls it a "two-photon compiled algorithm." If "genuine, non-cheating" excludes answer-dependent or problem-specific compilation, I cannot certify any larger-than-toy hardware factoring result as a clean Shor run. The 143 [verified-at-source] and 56153 [verified-at-source] claims are adiabatic/SAT-style factoring, not circuit Shor. Smolin, Smith & Vargo's criticism remains the key warning: "Previous experimental implementations have used simplifications dependent on knowing the factors in advance" and "Valid implementations should not make use of the answer sought." Source: Nature 499, 163-165 (2013), DOI 10.1038/nature12290 [verified-at-source].
## 1. Current RCS record
### Google Willow
Source: Google Quantum AI public blog, 2024-12-09 [verified-at-source], updated 2025-06-12 [verified-at-source], and Google Willow spec sheet PDF fetched from google/quantumai.
Numbers:
- Qubits: 103 [verified-at-source].
- Circuit depth: 40 [verified-at-source].
- XEB fidelity: 0.1% [verified-at-source].
- Quantum runtime: 5 minutes [verified-at-source].
- Claimed classical runtime: 10^25 years [verified-at-source].
- Classical hardware comparison: "one of today's fastest supercomputers" [verified-at-source]; Google blog notes estimates involving Frontier and says: "we assumed full access to secondary storage, i.e., hard drives, without any bandwidth overhead -- a generous and unrealistic allowance for Frontier." [verified-at-source]
- Method/cost caveat: Google says: "Computational costs are heavily influenced by available memory. Our estimates therefore consider a range of scenarios, from an ideal situation with unlimited memory ... to a more practical, embarrassingly parallelizable implementation on GPUs." [verified-at-source]
Independent context from Scott Aaronson, 2024-12 [verified-at-source]: "Google has also announced a new quantum supremacy experiment on its 105-qubit chip, based on Random Circuit Sampling with 40 layers of gates." Aaronson's cost summary: "if you use the best currently-known simulation algorithms (based on Johnnie Gray's optimized tensor network contraction), as well as an exascale supercomputer, their new experiment would take ~300 million years to simulate classically if memory is not an issue, or ~10^25 years if memory is an issue" [all numbers verified-at-source]. His verification caveat is direct: "for the exact same reason why ... this quantum computation would take ~10^25 years for a classical computer to simulate, it would also take ~10^25 years for a classical computer to directly verify the quantum computer's results!!" and "all validation of Google's new supremacy experiment is indirect, based on extrapolations from smaller circuits." Source: https://scottaaronson.blog/?p=8525
Verdict: Willow is the current public RCS record by claimed classical runtime that I found. It is not directly verified from inside by recomputing probabilities for the full circuit; validation is indirect/extrapolative.
### Later/nearby RCS result: Zuchongzhi 3.0
Source: Dongxin Gao, Daojin Fan, Chen Zha, Jiahao Bei, et al., "Establishing a New Benchmark in Quantum Computational Advantage with 105-qubit Zuchongzhi 3.0 Processor," arXiv:2412.11924 [verified-at-source], Phys. Rev. Lett. 134, 090601 (2025) [verified-at-source], DOI 10.1103/PhysRevLett.134.090601 [verified-at-source].
Numbers and quotes:
- Processor: "105 qubits" [verified-at-source]. Quote: "This superconducting quantum computer prototype, comprising 105 qubits, achieves high operational fidelities, with single-qubit gates, two-qubit gates, and readout fidelity at 99.90%, 99.62% and 99.18%, respectively." [all numbers verified-at-source]
- RCS circuit: "83-qubit, 32-cycle" [verified-at-source]. Quote: "Our experiments with an 83-qubit, 32-cycle random circuit sampling on Zuchongzhi 3.0 highlight its superior performance, achieving one million samples in just a few hundred seconds." [all numbers verified-at-source]
- Samples: "For the largest full circuit featuring 83 qubits and 32 cycles, we have collected a total of approximately 4.1 x 10^8 bitstrings." [verified-at-source]
- XEB/fidelity: "fidelity of 0.025%" [verified-at-source].
- Classical cost: "The estimated number of floating-point operations required to generate a million uncorrelated bitstrings with a fidelity of 0.025% from an 83-qubit, 32-cycle random circuit using a classical computer is 8.4 x 10^33." [all numbers verified-at-source]
- Frontier runtime: table reports "8.4 x 10^33" FLOPs and "6.4 x 10^9 yr" under the 9.2 PB memory case, and "7.5 x 10^31" FLOPs and "5.7 x 10^7 yr" under a 762.2 PB storage-assisted case [all numbers verified-at-source]. The paper states: "The Frontier supercomputer boasts a theoretical peak performance of 1.685 x 10^18 FLOPS. In our estimations, we presume a 20% FLOP efficiency" [all numbers verified-at-source].
Verdict: important later RCS result, but not a larger claimed gap than Willow.
### Post-Willow search status
Searches for 2025/2026 [verified-at-source: web searches run 2026-09-21] RCS hardware records found no public Google successor or non-Google result exceeding Willow's 103-qubit/depth-40/10^25-year claim. This is a null search result, not proof of nonexistence.
## 2. How weak is XEB as verification?
### XEB is not a sound certificate by itself
Aaronson & Gunn, "On the Classical Hardness of Spoofing Linear Cross-Entropy Benchmarking," arXiv:1910.12085 [verified-at-source], arXiv DOI 10.48550/arXiv.1910.12085 [verified-at-source]. Key quote: "This raises a theoretical question: how hard is it for a classical computer to spoof the results of the Linear XEB test?" Their positive result is conditional, not an unconditional certificate: "we show that the problem is classically hard, assuming that there is no efficient classical algorithm that... estimates the probability of C outputting a specific output string... with variance even slightly better than that of the trivial estimator" [verified-at-source].
Gao, Kalinowski, Chou, Lukin, Barak & Choi, "Limitations of Linear Cross-Entropy as a Measure for Quantum Advantage," arXiv:2112.01657 [verified-at-source], PRX Quantum 5, 010334 (2024) [verified-at-source], DOI 10.1103/PRXQuantum.5.010334 [verified-at-source]. Quotes: "achieving relatively high XEB values does not imply faithful simulation of quantum dynamics"; their classical algorithm "with 1 GPU within 2s, yields high XEB values, namely 2-12% of those obtained in experiments" [all numbers verified-at-source]; and "the XEB alone has limited utility as a benchmark for quantum advantage." [verified-at-source]
### Classical XEB/spoofing scorecard
- Original Sycamore 53-qubit/20-cycle instance [verified-at-source]: Pan Zhang et al., "Solving the sampling problem of the Sycamore quantum circuits," arXiv:2111.03011 [verified-at-source], Phys. Rev. Lett. 129, 090502 (2022) [verified-at-source], DOI 10.1103/PhysRevLett.129.090502 [verified-at-source]. Quote: "For the Sycamore quantum supremacy circuit with 53 qubits and 20 cycles, we have generated one million uncorrelated bitstrings ... approximate state has fidelity F~0.0037. The whole computation has cost about 15 hours on a computational cluster with 512 GPUs." [all numbers verified-at-source] This is the largest classical XEB/fidelity I verified for the full 53-qubit/20-cycle Sycamore instance: 0.0037, i.e. 0.37% [verified-at-source].
- Same Sycamore class, bounded-fidelity classical sampling: Kalachev et al., "Classical Sampling of Random Quantum Circuits with Bounded Fidelity," arXiv:2112.15083 [verified-at-source]. Quote: "classically produced 1 million samples with the fidelity bounded by 0.2%, based on the 20-cycle circuit of the Sycamore 53-qubit quantum chip" and "took about 14.5 days ... 32 GPUs" [all numbers verified-at-source].
- Faster Sycamore replay: "Leapfrogging Sycamore," National Science Review/PMC source fetched [verified-at-source]. Quote: "using 1432 GPUs to simulate quantum random circuit sampling that generates uncorrelated samples with a higher linear cross-entropy score and is 7 faster than the Sycamore 53-qubit experiment" [all numbers verified-at-source]. The paper states Google obtained "one (three) million uncorrelated samples in 200 (600) s, with a linear cross-entropy (XEB) of 0.2%" and reports "1432 NVIDIA A100 GPUs" producing "three million uncorrelated samples" in "86.4 s" with "13.7 kWh" [all numbers verified-at-source].
- Energetic-superiority preprint: Fu, Su, Zhong, Zhang, Pan Zhang, Jian-Wei Pan et al., arXiv:2407.00769 [verified-at-source]. Quote: "we have achieved a time-to-solution of 14.22 seconds with energy consumption of 2.39 kWh which achieved fidelity of 0.002 and our most remarkable result is a time-to-solution of 17.18 seconds, with energy consumption of only 0.29 kWh which achieved a XEB of 0.002 after post-processing" [all numbers verified-at-source].
- Largest raw XEB score I found, but on an easier shallower instance: arXiv:2512.07311 [verified-at-source] reports "simulating the 53-qubit, 14-cycle Sycamore circuit and achieving a linear cross-entropy benchmarking (XEB) score of 0.549, exceeding the published XEB score of 0.002 from Google's reference data" [all numbers verified-at-source]. This is not comparable to the 53-qubit/20-cycle Sycamore supremacy instance.
- 2026 theoretical cat-and-mouse claim: arXiv:2607.04054 [verified-at-source], "Frozen-Tree Sampling Refutes Quantum Advantage of Random Circuit Sampling." Quotes: it "draws bitstrings of n qubits in O(n) time per sample" and "no statistical test acting on samples alone can distinguish the classical frozen-tree sampler from a quantum random circuit" [verified-at-source]. This is a 2026 preprint/theoretical claim; I do not treat it as a settled empirical overthrow of Willow.
Verdict on XEB: classical simulation/spoofing overtook the original Sycamore benchmark. For Willow-scale RCS, quantum hardware remains ahead by published cost estimates, but the verification gap is exactly the problem: full XEB verification for the record instance is classically out of reach, so the full claim rests on extrapolated validation, model trust, and anti-spoofing assumptions rather than a compact sound certificate.
## 3. Certified / verifiable quantum advantage
### Aaronson-Hung protocol
Source: Scott Aaronson & Shih-Han Hung, "Certified Randomness from Quantum Supremacy," arXiv:2303.01625 [verified-at-source], arXiv DOI 10.48550/arXiv.2303.01625 [verified-at-source], STOC 2023 [verified-at-source], ACM DOI 10.1145/3564246.3585145 [verified-at-source via ACM/search result], pages 933-944 [verified-at-source via Nature reference].
Quotes: the paper studies "generating cryptographically certified random bits" [verified-at-source] and argues that when RCS outputs pass LXEB, "under plausible hardness assumptions they necessarily contain Omega(n) min-entropy" [verified-at-source]. Its bottleneck is explicit: "Currently, the central drawback of our protocol is the exponential cost of verification, which in practice will limit its implementation to at most n~60 qubits" [verified-at-source].
### JPMorganChase / Quantinuum / ORNL certified randomness
Source: Minzhao Liu, Ruslan Shaydulin, Pradeep Niroula, Matthew DeCross, Shih-Han Hung, Scott Aaronson, Marco Pistoia et al., "Certified randomness using a trapped-ion quantum processor," Nature 640, 343-348 (2025) [verified-at-source], DOI 10.1038/s41586-025-08737-1 [verified-at-source].
Exact numbers and quotes:
- Device: Quantinuum H2-1 trapped-ion processor [verified-at-source]. Quote: "We demonstrate our protocol using the Quantinuum H2-1 trapped-ion quantum processor accessed remotely over the Internet." [verified-at-source]
- Qubit/raw-bit size: 56 x 30,010 raw bits [verified-at-source], implying 56 measured bits per valid sample [verified-at-source]. Quote: "feed the 56 x 30,010 raw bits into a Toeplitz randomness extractor and extract 71,273 bits." [all numbers verified-at-source]
- Challenge circuits: "fixed arrangement of 10 layers of entangling UZZ gates, each sandwiched between layers of pseudorandomly generated SU(2) gates on all qubits" [numbers verified-at-source].
- Thresholds and response timing: expected fidelity "phi >= 0.3 or better on depth-10 circuits" [verified-at-source]; t_threshold = 2.2 s [verified-at-source]; chi = 0.3 [verified-at-source]; quantum device time t_QC = 2.154 s per sample [verified-at-source].
- Run size: b = 15 and b = 20 [verified-at-source]; 1,993 batches [verified-at-source]; 60,952 circuits [verified-at-source]; M = 30,010 valid samples [verified-at-source]; 984 successful batches [verified-at-source]; cumulative device time 64,652 s [verified-at-source].
- Verification cost: exact simulation time "100.3 s per circuit when using the entire [Frontier] supercomputer at a numerical efficiency of 45%" [all numbers verified-at-source]. "Frontier and Summit were used at full-machine scale" with "sustained peak performance of 897 petaFLOPS and 228 petaFLOPS" and "a combined performance of 1.1 exaFLOPS" [all numbers verified-at-source].
- Verification sample: XEB score for m = 1,522 circuit-sample pairs [verified-at-source], XEB_test = 0.32 [verified-at-source].
- Certified output: "at epsilon_sou=10^-6, we have Q_min=1,297, corresponding to H_min=71,313 against an adversary four times more powerful than Frontier" [all numbers verified-at-source]. Extracted bits: 71,273 [verified-at-source]. Input randomness: "only 32 bits" [verified-at-source].
- Security caveat: source says security is in a "restricted adversarial model" [verified-at-source]. Aaronson's public commentary says "about 70,000 certified random bits were generated over 18 hours" [verified-at-source] and that the parameters are "not yet good enough for my and Shih-Han's formal security reduction" but support "practical security" [verified-at-source].
Verdict: this is the cleanest empirical example of classically checkable/certified quantum advantage I found. It certifies randomness/entropy under assumptions and a restricted adversarial model; it is not universal delegated quantum computation.
### BCMVV / Mahadev-style classically verifiable proofs
Source: Brakerski, Christiano, Mahadev, Vazirani & Vidick, "A Cryptographic Test of Quantumness and Certifiable Randomness from a Single Quantum Device," arXiv:1804.00640 [verified-at-source], arXiv DOI 10.48550/arXiv.1804.00640 [verified-at-source]. Quote: it gives "a protocol for efficient classical verification that the untrusted device is 'truly quantum,' and a protocol for producing certifiable randomness from a single untrusted quantum device" [verified-at-source]. Assumption: post-quantum/LWE-style trapdoor claw-free functions [verified-at-source].
Source: Urmila Mahadev, "Classical Verification of Quantum Computations," arXiv:1804.01082 [verified-at-source], arXiv DOI 10.48550/arXiv.1804.01082 [verified-at-source]. Quote: "We present the first protocol allowing a classical computer to interactively verify the result of an efficient quantum computation" [verified-at-source]. Soundness depends on "the assumption that the learning with errors problem is computationally intractable for efficient quantum machines" [verified-at-source].
Hardware status: I found a proof-of-principle, not a full-scale sound implementation. Roman Stricker et al., "Towards experimental classical verification of quantum computation," arXiv:2203.07395 [verified-at-source], Quantum Science and Technology DOI 10.1088/2058-9565/ad2986 [verified-at-source via IOP search result]. Quote: "first, proof-of-principle experiment a verification protocol using only classical means on a small trapped-ion quantum processor" [verified-at-source]. Quote: "We implement the protocol experimentally on an eight-qubit trapped-ion quantum processor" [verified-at-source]. The authors explicitly flag relaxed/security-limited status: "We show how to verify existing quantum processors under relaxed security constraints, while the most stringent variant of such a protocol remains too demanding for current quantum hardware" [verified-at-source]. They also say the full protocol requires a "very large range" of trapdoor functions and "many auxiliary qubits" and is "not feasible on current devices" [verified-at-source].
Related efficiently verifiable advantage proposal: Kahanamoku-Meyer, Choi, Vazirani & Yao, "Classically verifiable quantum advantage from a computational Bell test," Nature Physics 18, 918-924 (2022) [verified-at-source], DOI 10.1038/s41567-022-01643-7 [verified-at-source]. Quote: "Sampling-based protocols... correctness ... is exponentially difficult to verify" [verified-at-source]. Their proposed test has quantum success "~85%" and classical bound "75%" [verified-at-source], but needs about "10^3 qubits" and "gate depth ~10^5" [verified-at-source]. I found no hardware run at that scale.
Verdict: cryptographic verification exists theoretically; hardware has proof-of-principle demonstrations and the 2025 certified-randomness result, but not general Mahadev verification of a useful quantum computation at scale.
## 4. Shor's algorithm on real hardware
### Largest Shor-family factoring result I verified
Source: Anthony Laing, Thomas Lawson, Roberto Alvarez, Xiao-Qi Zhou, Jeremy L. O'Brien et al., "Experimental realization of Shor's quantum factoring algorithm using qubit recycling," Nature Photonics 6, 773-776 (2012) [verified-at-source], DOI 10.1038/nphoton.2012.259 [verified-at-source].
Quote: "Encoding the work register in higher-dimensional states, we implement a two-photon compiled algorithm to factor N = 21." [verified-at-source] The article says this followed "four small-scale demonstrations" [verified-at-source].
Interpretation: N = 21 [verified-at-source] is the largest Shor-family hardware factorization I verified, but the word "compiled" matters. It is not evidence of cryptographically meaningful Shor scaling.
### Factor 15 baseline
Source: IBM retrospective and Nature paper page for Vandersypen et al., "Experimental realization of Shor's quantum factoring algorithm using nuclear magnetic resonance," Nature 414, 883-887 (2001) [verified-at-source], DOI 10.1038/414883a [verified-at-source]. IBM quote: the 2001 experiment "successfully factor[ed] the number 15" [verified-at-source]. IBM also quotes Isaac Chuang: "There are different ways to write the factor 15 algorithm, and if you simplify it sufficiently, well, then you can run it, but it's not terribly meaningful" [verified-at-source]. Vandersypen quote from IBM: "we still have a long way to go from toy problems to relevant applications. After all, we knew what the answer would be when we set out to factor the number 15." [verified-at-source]
### Why 143 and 56153 do not count as Shor records
Source: Nanyang Xu, Jing Zhu, Dawei Lu, Xianyi Zhou, Xinhua Peng & Jiangfeng Du, "Quantum Factorization of 143 on a Dipolar-Coupling Nuclear Magnetic Resonance System," Phys. Rev. Lett. 108, 130501 (2012) [verified-at-source], DOI 10.1103/PhysRevLett.108.130501 [verified-at-source], arXiv:1111.3726 [verified-at-source]. Quote: "Adiabatic quantum computation for this is an alternative approach other than Shor's algorithm" and they report factoring "the number 143" [verified-at-source]. Therefore it is not a Shor record.
Source: Nikesh S. Dattani & Nathaniel Bryans, "Quantum factorization of 56153 with only 4 qubits," arXiv:1411.6758 [verified-at-source], arXiv DOI 10.48550/arXiv.1411.6758 [verified-at-source]. Quote: the 143 computation "actually also factored much larger numbers such as 3599, 11663, and 56153" [all numbers verified-at-source], but the same abstract distinguishes it from Shor: "unlike the implementations of Shor's algorithm performed thus far" [verified-at-source], and concedes: "because they only use 4 qubits, these factorizations can also be performed trivially on classical computers" [verified-at-source].
Source: Sebastian Verschoor, "Factoring semi-primes with (quantum) SAT-solvers," arXiv:1902.01448 [verified-at-source], arXiv DOI 10.48550/arXiv.1902.01448 [verified-at-source]. Quote: "Shor's quantum factoring algorithm factors any integer in polynomial time, although large-scale fault-tolerant quantum computers capable of implementing Shor's algorithm are not yet available, so relevant benchmarking experiments for factoring via Shor's algorithm are not yet possible." [verified-at-source] On quantum/SAT/annealing factoring: "We find no evidence that this is a viable path toward factoring large numbers" [verified-at-source].
Source: John A. Smolin, Graeme Smith & Alexander Vargo, "Oversimplifying quantum factoring," Nature 499, 163-165 (2013) [verified-at-source], DOI 10.1038/nature12290 [verified-at-source]. Quote: "Previous experimental implementations have used simplifications dependent on knowing the factors in advance." Quote: "all composite numbers admit simplification of the algorithm to a circuit equivalent to flipping coins." Quote: "The difficulty of a particular experiment therefore depends on the level of simplification chosen, not the size of the number factored." Quote: "Valid implementations should not make use of the answer sought." [all verified-at-source]
Verdict: as of the sources checked on 2026-09-21 [verified-at-source: search date], nothing has changed in the meaningful Shor record. N = 21 is the largest Shor-family hardware demo I verified; no larger clean, non-answer-dependent Shor run was found.
## 5. Cost numbers for the current RCS record
### Willow cost numbers found
- Quantum runtime: 5 minutes [verified-at-source].
- Classical runtime: 10^25 years [verified-at-source].
- Alternative independent estimate: ~300 million years if memory is not an issue [verified-at-source], ~10^25 years if memory is an issue [verified-at-source].
- Classical method named by Aaronson: "Johnnie Gray's optimized tensor network contraction" [verified-at-source].
- Hardware class: exascale/Frontier-class supercomputer [verified-at-source].
- Raw FLOP count: not found in the fetched Google/Aaronson sources [not verified].
- Dollar cost: not found in the fetched sources [not verified].
- Core-hours: not found in the fetched sources [not verified].
The best exact cost statement I can support for Willow is the public runtime claim and memory-conditional tensor-network estimate, not a raw FLOP bill.
### Exact RCS cost numbers from Zuchongzhi 3.0, for comparison
- Method: tensor-network classical simulation/contraction [verified-at-source from paper context].
- Full circuit: 83 qubits [verified-at-source], 32 cycles [verified-at-source], fidelity 0.025% [verified-at-source], 1,000,000 samples [verified-at-source].
- Cost: 8.4 x 10^33 FLOPs [verified-at-source], 6.4 x 10^9 years on Frontier with 9.2 PB memory [all verified-at-source].
- More storage-rich scenario: 7.5 x 10^31 FLOPs [verified-at-source], 5.7 x 10^7 years with 762.2 PB storage [all verified-at-source].
- Frontier assumption: 1.685 x 10^18 FLOPS peak and 20% efficiency [verified-at-source].
### Exact verification cost from certified randomness, for contrast
- Verification simulations: 100.3 s per circuit on full Frontier at 45% numerical efficiency [all verified-at-source].
- Combined verification supercomputing: Frontier + Summit at 897 petaFLOPS + 228 petaFLOPS = 1.1 exaFLOPS [all verified-at-source].
- Verified sample subset: 1,522 circuit-sample pairs [verified-at-source].
- Certified entropy: 71,313 bits [verified-at-source]; extracted randomness: 71,273 bits [verified-at-source].
## Evidence class map
- Established: Willow public RCS benchmark numbers as a claimed hardware/classical-runtime demonstration; Zuchongzhi 3.0 peer-reviewed RCS numbers; Sycamore classical replay/spoofing numbers; Quantinuum/JPMC certified-randomness numbers; Shor N = 21 compiled record and Smolin et al. critique.
- Serious speculation / assumption-dependent: Aaronson-Hung security reduction; Aaronson-Gunn spoofing hardness; Willow's full-scale extrapolated verification; 2026 frozen-tree preprint.
- Null result: no post-Willow 2025/2026 public RCS record exceeding Willow was found in searches run 2026-09-21.
- Not verified: Willow raw FLOPs/core-hours/dollars; a full-scale hardware implementation of Mahadev's classical verification protocol; any clean larger-than-21 Shor run.