What is your paper about?
On 4 January 2023, Nature published ‘Papers and patents are becoming less disruptive over time’, receiving worldwide media attention. Our Matters Arising, published after a 32-month delay (see the process), shows that the reported decline can largely be attributed to dataset artefacts.
Detailed answer
The study is based on the CD index, a citation-based measure of how much the successors of a work still rely on its predecessors, running from −1 (consolidating) to +1 (disruptive). Key detail: a paper with zero references automatically gets CD = +1 as soon as it is cited once.
Due to a (then unreported) bug in the seaborn plotting software, the published histograms silently dropped the largest data points. A peak of papers with the maximum disruption score of CD = +1 was hidden from view, while these papers were kept in the analysis.
When removing the hidden outliers with CD = +1, the reported decline in disruption largely disappears across all databases used in the Nature study. For Web of Science, it reduces by 93%.
Why so many papers with a CD index of exactly 1? In the open-source SciSciNet database, 97% of them make zero references. Of 100 randomly sampled source documents, 93 do contain references. They are dataset artefacts.
Why didn't the robustness checks catch this? The linear regression did not control for the discontinuity of the CD index at zero references. Simply adding a dummy for zero references improves the adjusted R² from 0.15 to 0.95, and the adjusted CD index becomes essentially flat.
The Monte Carlo simulations randomly rewired citations while preserving each paper's number of references. Zero-reference papers are thus preserved, and the rewired networks inherit the peak at CD = +1. Their decline mirrors the observed decline in disruption.
Instead of a decline in the disruptiveness of scientific and technological knowledge, the data mainly show an increasing metadata quality over time: relatively fewer papers with zero references due to database errors.