Back to Home

The case for resilient scientific databases

Bharti Dharapuram
02 Jun 2026
features

Large centralized scientific databases form the backbone of research, but are vulnerable to failure from cyberattacks, funding cuts and geopolitical tensions. A recent Nature Genetics commentary suggests alternative models to build resilient, and shared infrastructure for scientific data. Image: Generated by Dr. Gaurav Sharma via ChatGPT.

Early last March, researchers were jolted when they lost access to PubMed, a free, publicly accessible search engine and citation database, which is a key part of life sciences research. The outage lasted over 24 hours, leaving researchers from different parts of the world unable to access one of the world’s largest catalogues of biomedical literature. This event revived ongoing discussions about the stability of centralized scientific databases, which are not just limited to PubMed.


In a recent commentary published in Nature Genetics, researchers led by Dr. Gaurav Sharma from the Department of Biotechnology, Indian Institute of Technology Hyderabad (IIT Hyderabad) argue for decentralized scientific databases. They propose alternative frameworks of data governance that are more resilient to technical disruptions and institutional funding cuts.


PubMed is maintained by the National Center for Biotechnology Information (NCBI) under the United States government, which also hosts GenBank, an archive of “all publicly available DNA sequences” containing more than 260 million genetic sequences and four billion genomes. Scientists around the world both contribute to and freely access these data, making such repositories an indispensable part of research.

Large scientific databases like these form the backbone of research across disciplines, and are used to address pressing challenges in health and climate.

Centralized governance ensures uniform data standards and makes these repositories easy to maintain and operate.


However, Dr Gaurav and co-authors argue that centralized repositories, given their single points of storage and access, are vulnerable to failure. They give the example of FlyBase, a genetic database of Drosophila, a widely used model organism in biology. Since last year, turbulent government funding has forced the platform to seek community contributions to continue data curation. Maintaining such large scientific databases requires sustained funding, but paywalled access to recover costs disadvantages researchers from economically poorer countries. Apart from budget cuts, cyberattacks are increasingly threatening research institutions, and data breaches are raising serious concerns about the privacy of sensitive health data.


“As scientists, we spend enormous effort generating data, but far less effort ensuring that the infrastructure supporting that data remains resilient for future generations,” says Dr Gaurav.

“Modern biology research runs on databases, yet many of them remain structurally vulnerable. A cyberattack, funding cut, or geopolitical conflict should not disrupt researchers’ collective scientific memory.”

As a solution, the authors describe two alternative models.


In the federated model, multiple institutional or national data nodes governed by their respective organizations work together through shared standards, thereby reducing the risk of failure.


Unlike federated systems, where institutions retain independent control of their data nodes, in a decentralized model, both data and governance are distributed across independent nodes and collectively managed by participating community members. Such systems, enabled by distributed ledgers, aim to improve data security and transparency in data ownership, access, and decision-making.


As a case study, the authors point to ELIXIR, Europe’s intergovernmental network for life science data. It brings together 240 research organisations across over 25 member countries through a federated system of data sharing and governance. The paper argues that such international cooperation can be supported by initiatives like the Committee on Data of the International Science Council, which promotes ethical governance frameworks and common standards for scientific databases. To support sustained funding for such collaborative systems, the authors share the example of the Global Biodata Coalition, a platform that brings together national and philanthropic funding agencies supporting biodata infrastructure.

The paper proposes a hybrid model that combines elements of both frameworks by integrating existing institutional databases with distributed storage, supported through collaborative funding.

They outline a path forward, where scientific organizations across countries, supported by international funding platforms, can collectively build a resilient and shared infrastructure for scientific data.

“For the Global South, open scientific databases are not a convenience; they are critical research infrastructure.”

“Their resilience and accessibility will determine how inclusive the next generation of scientific discoveries can be,” emphasizes Dr Gaurav.


Reference: Sharma, G., Munteanu, V., Ghiasi, N. M., Mahanta, U., Banerjee, J., Varma, S., Foschini, L., Ellrott, K., Mutlu, O., Ciorbă, D., Ophoff, R. A., Bostan, V., Moore, J. H., Sousoni, D., Krishnan, A., Lucaci, A. G., Tull, A., Mason, C. E., Dimian, M., … Mangul, S. (2026). Towards a decentralized future for open-science databases. Nature Genetics, 1–5. https://doi.org/10.1038/s41588-026-02606-x

Biotechnology
#decentralized databases #scientific databases #open science #data sharing #data governance