Research Data Reuse

The case for reaching back into archived research data.

The value of the average research dataset does not end when the paper it was collected for is published. Because of what has happened in machine learning, in some respects it begins there.

Before and after

BEFORE .dcm .csv .nii .pdf Scattered · unfindable · ungoverned PhD student graduated. Data sits idle. UNDOSA AFTER Paediatric myopia cohort 1,240 pts Retinal imaging — diabetes 860 pts More cohorts as contributed…

Where the waste comes from

A typical academic dataset is created inside a grant: a PI wins funding, recruits participants, collects imaging, runs the primary analysis, and publishes. The grant ends. The data sits on a university file server. The PhD student who knew the data structure graduates. Within a few years the dataset is functionally dead — technically it exists, but nobody can find it, understand its provenance, get permission to use it, or trust it enough to build on.

Chalmers and Glasziou's 2009 paper in The Lancet estimated that around 85% of biomedical research investment is avoidably wasted, much of it because answerable questions were never asked and data that could have answered them was never made findable. The estimate has been debated. The direction has not.


Why old data has become more valuable

Foundation models increase the value of historical data.

RETFound, the retinal foundation model from Moorfields Eye Hospital and UCL (2023), can extract signals from retinal images that were not extractable when those images were collected. A 2015 OCT sweep is scientifically more useful in 2026 than it was in 2015.

New collection has become harder.

Recruitment is harder post-pandemic. Consent expectations are higher. Imaging is more expensive to run. NHS research capacity is constrained. That makes already-collected data disproportionately more valuable.

Regulatory pressure has moved from carrot to stick.

Under the NIH data-sharing policy (January 2023) and the FDA's enforcement on unreported trial results, non-sharing of publicly funded data is increasingly costly.