Language Selection

Get healthy now with MedBeds!
Click here to book your session

Protect your whole family with Orgo-Life® Quantum MedBed Energy Technology® devices.

Advertising by Adpathway

         

 Advertising by Adpathway

Why biology must prioritize data reanalysis in the era of artificial intelligence

16 hours ago 7

PROTECT YOUR DNA WITH QUANTUM TECHNOLOGY

Orgo-Life the new way to the future

  Advertising by Adpathway

  • Loading metrics

Open Access

Perspective

The Perspective section provides experts with a forum to comment on topical or controversial issues of broad interest.

See all article types »

New technologies have generated huge volumes of biological data, much of which is underutilized. AI-assisted data reanalysis could drive the next era of biological discovery, but realizing its full potential requires standards, validation, and responsible use.

Citation: Ghaderi D, Sahimi A, Esser B, Ruggles KV, Maimon R (2026) Why biology must prioritize data reanalysis in the era of artificial intelligence. PLoS Biol 24(7): e3003888. https://doi.org/10.1371/journal.pbio.3003888

Published: July 22, 2026

Copyright: © 2026 Ghaderi et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Funding: The author(s) received no specific funding for this work.

Competing interests: The authors have declared that no competing interests exist.

Abbreviation:: AI, artificial intelligence

Scientific revolutions do not only emerge from new experiments, but often from new ways of interpreting existing observations. For example, Johannes Kepler transformed astronomy through rigorous reanalysis of observational data collected by Tycho Brahe, revealing the laws of planetary motion before their underlying mechanisms were understood.

Over the past decade, single-cell transcriptomics, spatial omics, and large-scale sequencing technologies have redefined what is technically possible to measure, generating biological datasets of unprecedented scale and complexity [1,2]. Concurrently, public repositories, open-access resources, and FAIR data practices have enabled the wide-scale sharing, integration, and reuse of biological datasets [3,4]. This emerging data ecosystem is creating the opportunity to revisit existing datasets with new hypotheses, analytical tools, and conceptual frameworks. Yet, as biological datasets become increasingly large, complex, and accessible, our capacity to interpret them has not scaled proportionally. Published datasets are frequently underutilized after primary analysis, as investigators shift toward new data synthesis to address specific questions. Additionally, poor quality standards and incomplete metadata often discourage investigators from engaging with shared datasets. This underutilization is exemplified by a methodological survey of more than 454,200 public omics datasets, which found that only 12,162 (2.7%) had undergone at least one documented reanalysis [5]. Consequently, vast amounts of high-quality biological data accumulate in the public domain with substantial unrealized potential for discovery.

A structural shift toward data reanalysis is needed. Reanalysis involves repurposing published datasets to address new questions, reevaluate prior conclusions, and uncover biological insights beyond the scope of the original study. It increases the scientific value of existing data while reducing the need for redundant data generation. By revisiting datasets with improved computational tools [6], refined biological questions, and alternative conceptual frameworks, researchers can uncover patterns and relationships that were invisible at the time of primary analysis. Notably, significant advances in tumor biology have been made possible by robust data sharing and reanalysis paradigms. For example, integration of multi-omic data across the Cancer Genome Atlas and the Broad Institute’s Firehose platform has enabled new tumor subtype classification with implications for patient prognosis and provided molecular phenotyping of histologic tumor subtypes [7,8]. Therefore, just as Kepler extracted hidden celestial order from preexisting astronomical observations, systematic reanalysis of biological datasets may reveal overlooked cellular states, latent genetic programs, and previously unrecognized principles of human biology. In this framework, reanalysis is not secondary science, but an engine for conceptual discovery itself. From a more practical standpoint, reanalysis is less expensive than experimentation and can be conducted from virtually anywhere, lowering financial or geographic barriers to entry while also being environmentally advantageous by reducing the need for redundant, large-scale experiments. Rather than competing with new data generation, reanalysis and experimentation should be viewed as complementary activities, with each informing and strengthening the other.

But biological data today are far richer and more complex than in the era of Kepler, introducing both unprecedented scientific opportunity and substantial analytical challenges. Conventional computational approaches remain limited in their ability to extract subtle biological structure from the high-dimensional, multimodal datasets that now define modern biology. For example, cell type annotation in transcriptomic datasets remains highly sensitive to preprocessing strategies, quality-control thresholds, dimensionality reduction methods, clustering granularity, and marker selection [2]. These technical choices can profoundly influence the identification of rare or quiescent cell populations, including stem cells and immune cell subtypes, limiting reproducibility across studies. Such challenges extend across biological data modalities and increasingly constrain our ability to fully interpret large-scale datasets.

Fortunately, the very scale and complexity of high-throughput data make it inherently compatible with machine learning and artificial intelligence (AI) tools, which draw statistical power from size, variability, and cross-cohort integration. AI can be used to model nonlinear relationships, integrate heterogeneous modalities, and extract structure from high-dimensional biological space, all of which are necessary to disentangle meaning from big data. For instance, AI systems have already demonstrated the ability to reproducibly classify cells based on validated reference data, improving the reliability of annotation pipelines [9]. Beyond improving upon current analytical methods, AI opens new horizons of nuanced data interpretation such as reconstructing differentiation trajectories or modeling intercellular communication dynamics within complex microenvironments [10]. Ultimately, AI has the potential to work synergistically with reanalysis in a restructured data lifecycle that maximizes scientific insight [11]. Models can be used to identify rare, quiescent, or subtle biology that escape primary analyses, allowing for the reannotation of datasets to improve training data for subsequent models. This iterative cycle creates a progressively more accurate representation of biological systems. Importantly, AI does not replace biological reasoning; it amplifies it by enabling deeper interrogation of data already in hand.

But indiscriminate use of AI could substantially distort the research landscape. When trained on poorly powered, homogeneous, or biased datasets, model predictions can diverge from the underlying biological reality, generating inaccurate conclusions that may become amplified through iterative training cycles. In this way, small errors have the potential for pervasive impact. Furthermore, AI can be exploited to rapidly generate poor-quality analyses that further corrupt the scientific literature. A striking example emerged from the National Health and Nutrition Examination Survey, a large public health dataset that was misused to produce hundreds of formulaic analyses based on selectively sampled variables in a process suggestive of large-scale data dredging [12].

To address these risks, rigorous standards for AI-assisted data reanalysis must be established. This begins with careful selection of training datasets containing well-defined ground truth, perturbational structure, longitudinal information, and consistent metadata. Robust training data should be well-annotated, interoperable across modalities, and aligned with FAIR data principles [3]. Importantly, experimentation and primary data generation remain essential, and researchers must remain cognizant of the long-term potential for their datasets to become integrated into future AI systems. Particular attention must be paid to minimizing technical, biological, and demographic biases that may otherwise become amplified through AI-driven analyses. Additionally, data sharing must be practiced with high standards for data and metadata quality to enable reanalysis and improve algorithmic training. Simply increasing the volume of observational data does not automatically improve inference.

Furthermore, validation and reproducibility standards must ensure that machine learning models undergo rigorous testing across diverse datasets and that algorithmic outputs are not mistaken for biological truth. Contrary to concerns about human replacement, critical human reasoning becomes more important than ever in the era of AI. Scientists and clinicians must remain responsible for evaluating model reliability, interpreting biological significance, and experimentally or clinically validating computational predictions. Even in the presence of advanced AI systems, scientific judgment remains indispensable, and algorithmic outputs must ultimately be subject to human scrutiny and validation.

Finally, ethical standards must be upheld to prevent irresponsible or fraudulent use of AI that undermines scientific integrity. Because AI systems can generate misleading, biased, or low-quality findings at scale, human judgment is vital for maintaining rigor, transparency, and trust in science [13]. Accordingly, AI should not be ignored in educational settings. Instead, the next generation of biologists must be trained to use AI critically, responsibly, and with appropriate scientific skepticism.

A new era of biological discovery driven by AI-assisted data reanalysis is on the horizon. Scientific progress has historically advanced through multiple complementary paths, including experimental discovery, conceptual reinterpretation of existing observations, and the communication of transformative ideas. While modern biology continues to emphasize data generation, the coming years may require renewed focus on the equally powerful role of reanalysis. Like Kepler’s reinterpretation of astronomical observations, AI-driven reanalysis may fundamentally reshape biological discovery, but realizing its potential will require rigor, responsibility, and continued human oversight.

Acknowledgments

ChatGPT was used during the preparation of the initial draft to refine the text.

References

  1. 1. Regev A, Teichmann SA, Lander ES, Amit I, Benoist C, Birney E, et al. The Human Cell Atlas. eLife. 2017 Dec 5;6:e27041.
  2. 2. Lähnemann D, Köster J, Szczurek E, McCarthy DJ, Hicks SC, Robinson MD, et al. Eleven grand challenges in single-cell data science. Genome Biol. 2020;21(1):31. pmid:32033589
  3. 3. Wilkinson MD, Dumontier M, Aalbersberg IJJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. pmid:26978244
  4. 4. Yao Z, van Velthoven CTJ, Kunst M, Zhang M, McMillen D, Lee C, et al. A high-resolution transcriptomic and spatial atlas of cell types in the whole mouse brain. Nature. 2023;624(7991):317–32. pmid:38092916
  5. 5. Perez-Riverol Y, Zorin A, Dass G, Vu M-T, Xu P, Glont M, et al. Quantifying the impact of public omics data. Nat Commun. 2019;10(1):3512. pmid:31383865
  6. 6. Stuart T, Butler A, Hoffman P, Hafemeister C, Papalexi E, Mauck WM 3rd, et al. Comprehensive integration of single-cell data. Cell. 2019;177(7):1888-1902.e21. pmid:31178118
  7. 7. Thennavan A, Beca F, Xia Y, Recio SG, Allison K, Collins LC, et al. Molecular analysis of TCGA breast cancer histologic types. Cell Genom. 2021;1(3):100067. pmid:35465400
  8. 8. Coudray N, Ocampo PS, Sakellaropoulos T, Narula N, Snuderl M, Fenyö D, et al. Classification and mutation prediction from non-small cell lung cancer histopathology images using deep learning. Nat Med. 2018;24(10):1559–67. pmid:30224757
  9. 9. Galdos FX, Xu S, Goodyer WR, Duan L, Huang YV, Lee S, et al. devCellPy is a machine learning-enabled pipeline for automated annotation of complex multilayered single-cell transcriptomic data. Nat Commun. 2022;13(1):5271. pmid:36071107
  10. 10. Jin S, Plikus MV, Nie Q. CellChat for systematic analysis of cell-cell communication from single-cell transcriptomics. Nat Protoc. 2025;20(1):180–219. pmid:39289562
  11. 11. Gottweis J, Weng W-H, Daryin A, Tu T, Sirkovic P, Myaskovsky A, et al. Accelerating scientific discovery with Co-Scientist. Nature. 2026;:10.1038/s41586-026-10644-y. pmid:42156544
  12. 12. Suchak T, Aliu AE, Harrison C, Zwiggelaar R, Geifman N, Spick M. Explosion of formulaic research articles, including inappropriate study designs and false discoveries, based on the NHANES US national health database. Munafò M, editor. PLOS Biol. 2025 May 8;23(5):e3003152. https://doi.org/10.1371/journal.pbio.3003152 pmid:40338847
  13. 13. Hanna MG, Pantanowitz L, Jackson B, Palmer O, Visweswaran S, Pantanowitz J. Ethical and bias considerations in artificial intelligence/machine learning. Modern Pathol. 2025;38(3):100686. pmid:39694331
Read Entire Article

         

        

Start the new Vibrations with a Medbed Franchise today!  

Protect your whole family with Quantum Orgo-Life® devices

  Advertising by Adpathway