Quantum chemistry, from molecular data to real hardware and back
Q-Alchemy helps pharmaceutical, materials and chemical research teams take molecular states through a complete quantum experiment: prepare them on a quantum processor, collect measurements, and reconstruct compact models for further analysis.
Both ends of that journey are demanding. A molecular description cannot simply be copied into a quantum computer, and the resulting quantum state cannot simply be downloaded. Q-Alchemy connects the methods needed to cross both boundaries.
The key is structure. Our QTucker approach (research paper (external site), published chapter (external site)) exploits patterns in molecular states to construct their preparation circuits. On the return journey, classical-shadow-based measurement information guides the reconstruction of a compact state model. This makes it possible to work with structured states without expanding and measuring every part of a completely general quantum state.
- Prepare molecular states on real hardware. Turn classical molecular calculations into quantum circuits that exploit the state’s structure.
- Recover a model you can work with. Combine selected measurements with a Tucker representation to reconstruct an explicit molecular-state model.
- Evaluate the complete experiment. Follow the state through preparation, hardware execution and reconstruction, then compare the result with your reference.
Demonstrated on IBM quantum hardware
We completed this workflow across twelve molecular systems. For twenty-qubit oxygen and nitrogen, reconstruction from 399 observables per molecule achieved fidelities of 0.929 and 0.895 against the reference states.
These results bring preparation and reconstruction together in a working hardware pipeline. The detailed campaign below shows the methods, results and limits encountered along the way.
For an accessible introduction, read From Molecular Data to Real Quantum Hardware and Back.
Running Molecular States on IBM Quantum Hardware: An End-to-End Pipeline
Over the past few months we have been building QTucker (arXiv preprint (external site), Springer chapter (external site)), a framework for preparing, measuring and reconstructing molecular quantum states using low-rank tensor structure. Recently we took the framework from simulation and onto real IBM hardware. This post describes what that campaign involved, focusing less on the mathematics and more on the engineering: the shape of the pipeline, how information flows in and out of a quantum processor, and what the hardware told us.
This work was funded by the German Federal Ministry of Research, Technology and Space (BMFTR) under grant 13N17157 (project QROM). The experiments were made possible by quantum credits awarded to data cybernetics through the IBM Quantum Startup Program. Access to real devices is the constraint that separates a plausible method from a validated one, and we are grateful for it.
The Pipeline in Three Stages
The workflow is deliberately split into three stages that communicate only through files on disk. No state is shared in memory. Each stage can be inspected, re-run and audited independently.
Stage one — state preparation. We begin classically. Molecular electronic structure is computed with PySCF at the CISD level, and the resulting wavefunction is converted into a sparse representation rather than a dense state vector. That sparse object is then passed to our iterative Tucker initialisation routine, which constructs a quantum circuit that approximately prepares the state. Each circuit is exported as an OpenQASM 3 file, alongside a JSON that defines how the observable set for that molecule is to be constructed.
Stage two — acquisition. The exported circuits are submitted to IBM hardware. Our Tucker state-estimation tooling provides the set of Pauli observables to be measured and handles the mapping between logical and physical qubits required for hardware execution. The resulting expectation values are written back out as JSON, each accompanied by its reported standard error.
Stage three — reconstruction. Entirely offline, Tucker state estimation is used to fit a compressed low-rank surrogate state to the measured expectation values, and the result is compared against the original reference wavefunction.
Twelve molecules were carried through this pipeline, ranging from hydrogen at four qubits to dichromium at seventy-two.
The value of this separation is practical, as hardware runs are expensive and so the measured data must be treated as a durable asset rather than a transient intermediate result. Configuration lives in versioned JSON schemas rather than in code, meaning observable budgets or hyperparameters can be retuned per molecule without touching the pipeline itself, and without re-running the quantum stage.
The Quantum Input and Output
Quantum input and output here means the two data objects that define the interface between the classical and quantum halves of the workflow: the description of the state to be prepared, which is sent to the device, and the set of measured expectation values, which is what comes back. Everything the quantum processor contributes to this experiment passes through those two objects, so it is worth being precise about their form.
Going in: An OpenQASM 3 circuit restricted to a single-qubit rotation gate and the two-qubit CX gate, together with a list of Pauli strings to be measured. Nothing vendor-specific, nothing serialised from a particular SDK. We enforce this strictly — the export step fails loudly if any gate outside that basis survives decomposition. The result is an artefact that is portable, human-readable and independently checkable.
Coming out: a list of triples. Each entry is an observable, its measured expectation value, and its standard error. That is the entire quantum contribution to the experiment. Everything before it is electronic structure theory and circuit synthesis; everything after it is optimisation.
Two engineering details in this exchange proved important.
The first is that we never allow a dense state vector to exist anywhere in the pipeline. For seventy-two qubits, the dense representation would require roughly 4.7 × 10²¹ complex amplitudes. Our reference state for dichromium instead holds 54,167 non-zero terms, and the fitted surrogate is described by 453 complex parameters. Sparsity and low-rank structure are not optimisations here; they are the only reason the problem is expressible at all.
The second is bookkeeping between logical and physical qubits, handled by the same state-estimation tooling that defines the observable set. Hardware execution requires circuits transpiled to the device’s actual topology, and the observables must be mapped onto the physical layout alongside them. A fourteen-qubit molecular circuit becomes a circuit across the full device width. We therefore map observables forward for submission and map the returned values back onto the original logical observables before saving, so that the reconstruction stage recovers a molecular state rather than a machine-shaped one. This is unglamorous work, and it is exactly the kind of detail that quietly invalidates results when it is handled loosely.
The IBM Hardware
Execution took place on ibm_berlin, a 120-qubit device accessed through IBM’s Frankfurt region. Measurement used the Estimator primitive at 8192 shots per observable with the platform’s built-in error mitigation enabled. Transpilation was performed locally at optimisation level one, so that the submitted circuit was known and recorded rather than inferred.
The full campaign ran for approximately six and a half hours of wall-clock time. Across twelve molecules it measured roughly 21,000 observables, corresponding to on the order of 170 million individual shots.
One qualification applies to everything that follows. These figures come from a single pass over the twelve molecules — one execution of each circuit, on one device, on one day. We did not repeat runs, vary the qubit layout, or average across sessions, so nothing here carries run-to-run error bars. Device calibration drifts, and a second campaign would not reproduce these numbers exactly. They are reported as they were obtained, and they should be read as a first measurement rather than a characterisation.
The headline result is that the pipeline worked on real hardware, at every scale we attempted, and produced reconstructions of genuine chemical quality for the smaller and mid-sized systems. Hydrogen reconstructed at 0.986 fidelity against the reference wavefunction, lithium hydride at 0.964, beryllium hydride at 0.949, molecular oxygen at 0.929 and nitrogen at 0.895 — the latter two on twenty-qubit circuits executed on physical superconducting qubits, not simulated. Carbon monoxide followed at 0.823. These are measured values, obtained from noisy hardware, reconstructed from a compressed observable set rather than from full tomography.
The full record of the campaign is below. Times are wall-clock seconds; the initialisation fidelity loss is the approximation error introduced by the state-preparation step, and the surrogate fidelity is the final reconstruction measured against the reference wavefunction.
| Molecule | Sparse terms | Qubits | CX | U | Submitted depth | Observables | Init. time (s) | Init. fidelity loss | Acquisition time (s) | Fit time (s) | Surrogate fidelity |
|---|---|---|---|---|---|---|---|---|---|---|---|
| H₂ | 2 | 4 | 0 | 2 | 2 | 31 | 0.054 | 0.012666 | 983.9 | 6.7 | 0.985593 |
| LiH | 35 | 12 | 11 | 26 | 22 | 95 | 0.055 | 0.012778 | 466.9 | 14.2 | 0.964424 |
| BeH₂ | 39 | 14 | 14 | 32 | 17 | 111 | 0.052 | 0.020692 | 861.3 | 18.8 | 0.949237 |
| H₄ | 11 | 8 | 0 | 4 | 2 | 63 | 0.051 | 0.040575 | 466.6 | 10.5 | 0.957586 |
| CO | 288 | 20 | 71 | 133 | 184 | 399 | 0.357 | 0.048231 | 1044.4 | 48.4 | 0.823219 |
| N₂ | 160 | 20 | 57 | 104 | 154 | 399 | 0.204 | 0.049670 | 1064.1 | 46.8 | 0.895175 |
| HCN | 525 | 22 | 129 | 229 | 222 | 451 | 0.557 | 0.055947 | 1065.3 | 52.9 | 0.634280 |
| H₂CO | 421 | 24 | 66 | 125 | 194 | 511 | 9.002 | 0.037987 | 1064.4 | 59.1 | 0.761860 |
| C₂H₄ | 521 | 28 | 148 | 266 | 158 | 1791 | 22.128 | 0.046255 | 2842.0 | 234.2 | 0.460694 |
| CO₂ | 653 | 30 | 166 | 297 | 212 | 1807 | 27.559 | 0.053259 | 3203.2 | 228.6 | 0.530111 |
| O₂ | 73 | 20 | 8 | 34 | 16 | 399 | 0.107 | 0.028848 | 1479.7 | 46.9 | 0.929417 |
| Cr₂ | 54,167 | 72 | 1902 | 2693 | 3364 | 14351 | 52.941 | 0.119742 | 8772.0 | 2393.9 | 5.021 × 10⁻¹⁵ |
The device handled everything we submitted. Circuits ranging from two gates to several thousand were transpiled, executed and returned without a single failed job, and the reported standard errors came back consistent with the 8192-shot budget throughout, indicating that the measurement statistics behaved as expected across all twelve systems. Being able to take a seventy-two qubit preparation circuit, submit 14,351 observables against it and receive well-formed data back is itself a meaningful capability, and one that was not routinely available to a company of our size a few years ago.
The three twenty-qubit cases are the most informative part of the campaign, because they share both qubit count and observable count, and so isolate a single variable. Their fidelities span 0.929 to 0.823 in exact correspondence with circuit depth and two-qubit gate count. This is a useful and actionable finding: qubit count alone predicts little about reconstruction quality, whereas circuit complexity predicts a great deal. It tells us directly where to invest — in shallower preparation circuits — rather than in waiting for larger devices.
At the largest scale, dichromium at seventy-two qubits, the reconstruction did not converge to the reference state. We regard this as a well-characterised boundary rather than a setback. The state-preparation stage itself performed well, retaining an initialisation fidelity of 0.88 on a state with more than fifty-four thousand terms, and the analysis of the acquired data separates the remaining contributions cleanly: part is systematic distortion accumulated across a very deep circuit, and part is a structural ceiling in the rank-one surrogate we chose for reconstruction, which we can bound analytically. Both are addressable, and knowing which is which is the point of running the experiment on hardware in the first place.
The rank ceiling is not a matter of opinion. A rank-one Tucker surrogate is separable across the chosen block partition, so its fidelity with the reference state cannot exceed the square of the largest Schmidt coefficient across the most restrictive cut. Evaluating that bound on the reference states gives the following limits, which hold regardless of how clean the measured data are.
| Molecule | Rank-1 fidelity upper bound |
|---|---|
| CO | 0.9411 |
| N₂ | 0.9369 |
| HCN | 0.9357 |
| H₂CO | 0.9387 |
| C₂H₄ | 0.9297 |
| CO₂ | 0.9214 |
| Cr₂ | 0.8656 |
For dichromium the initialisation fidelity of 0.88 already sits above the 0.8656 ceiling, so the prepared state cannot be represented exactly by the reconstruction model we used. That is a clear and fixable model mismatch.
What We Take From This
Reporting the limits matters as much as reporting the successes. The seventy-two qubit case establishes where the current configuration reaches its boundary, and the intermediate cases establish that the boundary is set by circuit complexity rather than by problem size in qubits — which is a considerably more tractable constraint to work against.
The immediate next step is a controlled comparison across four conditions: exact observables of the reference state, exact observables of the prepared circuit, an ideal finite-shot simulation at equivalent precision, and a run performed on the same hardware. That sequence isolates approximation error, statistical error, reconstruction-model error and hardware contribution separately, rather than allowing them to accumulate into a single unattributable number. Raising the surrogate rank is the other obvious lever, since we can prove the rank-one model has a fidelity ceiling below unity for several of the larger molecules regardless of data quality.
These are, then, interim results. A single pass across twelve molecules is enough to establish that the pipeline runs end to end on production hardware, and enough to identify which variables matter most — but it is a calibration exercise, not a characterisation of the method. Repeat runs would let us separate device variability from the effects we have attributed to circuit complexity, and the four-way control comparison would attribute the remainder properly. We will extend the study as further device time becomes available, and we expect the next pass to be considerably better designed for having run this one.
The broader point is that useful quantum experimentation today is largely a systems discipline. The quantum processor contributed a list of numbers with error bars. Everything that determined whether those numbers were meaningful — how the state was compressed, which observables were chosen, how the layout was tracked, what was written to disk — was classical engineering around it.
This work was funded by the German Federal Ministry of Research, Technology and Space (BMFTR) under grant 13N17157 (project QROM). Experiments conducted with quantum credits provided through the IBM Quantum Startup Program.
The cover image was generated with AI. It is a conceptual illustration, not a photograph of IBM hardware or a visualisation of experimental results.
