Research Data Sharing Statistics 2026: Mandates, Compliance & Repository Deposits by Field
Funder mandates for open data are proliferating faster than researcher compliance. The NIH’s landmark Data Management and Sharing (DMS) Policy turned three years old in 2026, Horizon Europe continues to press FAIR data requirements on every grant beneficiary, and UKRI is consolidating nine research-council policies into a single unified framework. Yet a comprehensive 2023 systematic review of more than 2.1 million biomedical articles found that actual verified data sharing remains at approximately 2% — a figure that starkly illustrates the gap between policy ambition and on-the-ground practice. This article brings together the most current research data sharing statistics available in 2026, drawing on peer-reviewed meta-research, funder reports, and repository annual data to map where progress is real and where the policy-practice gap persists.
Major Funder Mandates in 2026
The three most influential public research funders — NIH (United States), Horizon Europe (European Commission), and UKRI (United Kingdom) — have each implemented mandatory data management and sharing requirements that now reach the majority of English-language academic research.
NIH Data Management and Sharing Policy (effective January 2023)
The NIH DMS Policy, which became enforceable on 25 January 2023, is the broadest data sharing mandate in US research history. It applies to all research funded wholly or in part by NIH that generates scientific data, with no lower threshold on grant size. Applicants must submit a Data Management and Sharing Plan (DMS Plan) alongside their grant application; the plan is reviewed by the relevant Institute or Center before award, and investigators must document compliance in their annual Research Performance Progress Report (RPPR) starting in 2024.
NIH Institutes, Centers, and Offices collectively reviewed over 1,100 DMS Plans in the first phase of the policy’s operation and found the majority either acceptable initially or requiring only minor revision. Common deficiencies included excessive length and inclusion of extraneous detail not relevant to the sharing plan itself. Non-compliance can result in special award conditions and may affect future funding decisions.
| Requirement | Detail |
|---|---|
| Applies to | All NIH-funded research generating scientific data (competitive applications from 25 Jan 2023) |
| Submission requirement | DMS Plan submitted with grant application; revised if not approved |
| Compliance reporting | Annual RPPR from 2024 onwards; must confirm adherence to approved plan |
| Enforcement | Special award conditions; may affect future funding decisions |
| Preferred repositories | Domain-specific repositories first; generalist (Zenodo, Dryad, Figshare) acceptable |
Horizon Europe Open Science and FAIR Data Requirements
Horizon Europe, the EU’s flagship research programme with a budget of €95.5 billion for 2021–2027, treats open science as a default obligation rather than an optional commitment. All beneficiaries who generate or reuse digital research data are required to produce a Data Management Plan within six months of project start and to update it throughout the project lifecycle. Data must be made open according to the principle of “as open as possible, as closed as necessary” and must align with the four FAIR principles (Findable, Accessible, Interoperable, Reusable).
In terms of repository infrastructure, a 2023 European Research Council analysis found that only five repositories — including Zenodo, which is co-developed by CERN and funded partly through OpenAIRE — met the “Essential” readiness level for metadata standards required by Horizon Europe. Many repositories were classified as “Close-to-Essential,” indicating that the broader repository ecosystem still has structural gaps relative to funder expectations.
UKRI Policy Framework Development
UKRI announced in December 2024 that it is consolidating the nine separate research-council data policies into a single unified Research Data Policy Framework, aligning with FAIR principles and extending coverage to software and code. UKRI ran a public consultation on its draft policy between April and August 2025. A final policy version, incorporating stakeholder feedback, is expected to be published in summer 2026. Existing UKRI grant conditions already require researchers to make data underlying published findings “findable, accessible and where possible open” via an appropriate repository, with Data Management Plans as a standard award condition.
Data Sharing Compliance Rates
Raw mandate numbers tell only part of the story. The critical question is whether stated policies translate into deposited, accessible datasets. A comprehensive 2023 systematic review and meta-analysis published in BMJ examined 105 meta-research studies covering 2.1 million articles from the health and medical sciences and produced the clearest picture yet of declared versus actual compliance.
| Sharing route | Declared availability | Actual verified sharing |
|---|---|---|
| Public data sharing (all journals) | 8% | 2% |
| Code sharing (all journals) | 0.3% | 0.1% |
| Under mandatory journal policy | 65% | 33% |
| Under “encourage sharing” policy | 17% | 8% |
| No journal policy | 17% | 4% |
Source: Serghiou et al. (2023), BMJ, systematic review and meta-analysis of 2.1 million articles, PMC10334349.
A particularly striking finding is the variation within the “mandatory policy” category: compliance ranged from 0% to 100% across individual journals, indicating that policy wording and enforcement rigour — not just the existence of a mandate — are the determining factors. Journals with active peer-reviewer verification and post-acceptance data validation are at the high end of that range; journals relying solely on author self-certification are at the low end.
Compliance by Data Type
Within the biomedical sciences, verified sharing varies dramatically by the type of data being reported:
- Genomic/sequence data: ~57% actual sharing — the highest rate across any data category, driven by long-established deposition mandates in journals such as Nature and by purpose-built databases like GenBank and the European Nucleotide Archive.
- Systematic review data: ~6% actual sharing.
- Clinical trial individual participant data (IPD): ~1% actual sharing — among the lowest, constrained by patient confidentiality, consent frameworks, and institutional risk aversion.
The Policy–Practice Gap
The 2024 State of Open Data Special Report, produced jointly by Digital Science, Figshare, and Springer Nature, combined survey data with behavioural evidence drawn from Dimensions, Springer Nature Data Availability Statements, and the Wellcome-funded Make Data Count / DataCite citation corpus. Its headline finding was that approximately 2 million datasets are now published annually — a volume roughly matching the total number of journal articles published globally in the year 2000 — but that the policy-practice gap remains structural rather than incidental.
The report identified three interacting dynamics that sustain the gap:
- Substitution rather than addition: As “on request” sharing declines (reductions of 1–9% across most countries), some of that volume is moving not into repositories but into “not applicable” categories or informal venues, suggesting that some researchers are reclassifying data rather than depositing it.
- Incentive misalignment: US researchers show the lowest citation motivation for sharing (4.88%) but the highest funder-requirement motivation (10.23%), suggesting compliance is increasingly administrative rather than intrinsic. By contrast, researchers in Ethiopia and Japan show much higher citation motivation (9.3% and 14.8% respectively) but are in environments where funder enforcement is weaker.
- FAIR adoption lags policy: Research among clinical researchers found that only 11% address all aspects of FAIR in their data management. Although 94.7% of researchers recognise that FAIR data would benefit others, and 89.3% are willing to adopt FAIR practices given adequate support, fewer than 40% add metadata to data elements and just over one third deposit data in a repository.
For a detailed explanation of what FAIR compliance entails, the site’s guide to FAIR data principles for researchers provides a breakdown of all 15 FAIR sub-principles and which repositories implement them.
Repository Deposit Growth: Zenodo, Figshare, Dryad
Three generalist repositories dominate English-language data deposit activity: Zenodo (operated by CERN), Figshare (operated by Digital Science), and Dryad (an independent non-profit). Each serves a somewhat different community but all three accept data from any discipline.
Zenodo
Zenodo, launched in 2013 under a European Commission/CERN partnership, is the preferred generalist repository for Horizon Europe-funded research and the default for many researchers outside the US. As of 2025, Zenodo hosts more than 3 million uploads spanning publications, datasets, software, presentations, and other research outputs, with total data volume exceeding 1 petabyte. The platform receives approximately 25 million site visits per year and is backed by a 5 petabyte CERN EOS storage cluster with comprehensive daily backup infrastructure. Zenodo’s infrastructure is co-developed with over 25 institutional partners through the InvenioRDM open-source project.
Dryad
Dryad is discipline-agnostic but particularly strong in ecology, evolution, and the life sciences. According to Dryad’s 2023–24 Annual Report, the repository released 5,567 new datasets between July 2023 and June 2024. The most prolific depositing journals during that period were Science, Royal Society Open Science, and PLOS ONE. Dryad added 14 new institutional and publishing partners during the fiscal year and launched a dedicated working group to address the growing challenge of very large datasets. Dryad is also an active participant in the NIH’s Generalist Repository Ecosystem Initiative, which coordinates interoperability across generalist repositories.
Figshare
Figshare, which powers institutional repositories for hundreds of universities globally, reported that its 2024 development priorities were centred on streamlining researcher workflows and improving administrative data oversight — reflecting the growing demand from institutions needing to report compliance with NIH, UKRI, and Horizon Europe mandates. Figshare’s cloud infrastructure underpins both the public figshare.com portal and private institutional portals, making it one of the highest-volume data hosting systems in academic publishing.
Data Sharing Rates by Academic Field
Research sharing norms vary significantly across disciplines, shaped by the nature of the data, the existence of community infrastructure, ethical constraints, and the strength of journal mandates within each field.
| Field / Category | Repository deposit rate | Notes |
|---|---|---|
| Genomics / sequence data (biology) | ~57% | Long-standing mandatory deposition; GenBank, ENA |
| Psychology (Open Data Badge journals) | ~23% | Post-badge era at Psychological Science; up from <3% |
| Biomedical/clinical sciences (PLOS) | 19% | PLOS articles; OA comparator cohort 10% |
| Health sciences (PLOS) | 19% | PLOS articles; OA comparator cohort 6% |
| Engineering | 5% | OA comparator cohort; code sharing higher in computing |
| Clinical trials (IPD) | ~1% | Privacy and consent constraints; lowest across all categories |
| Humanities & social sciences | <5% (estimated) | Methodology differences; qualitative data harder to standardise |
Sources: Serghiou et al. (2023) PMC10334349; PLOS open science by discipline analysis (2023); Psychological Science badge programme data.
Why Biology Leads and Humanities Lags
The biological sciences, and genomics in particular, benefit from purpose-built infrastructure (GenBank, the European Nucleotide Archive, UniProt) that reduces the friction of deposition to near zero. Journal mandates in Nature, Science, and other flagship titles have required sequence data deposition for decades, normalising the practice across the entire field. Code sharing in mathematics, information and computing sciences, and physical sciences is similarly high, driven by community norms around reproducibility.
Humanities and qualitative social science face structural obstacles: interview transcripts, ethnographic field notes, and archival materials carry confidentiality obligations, intellectual property complications, and often cannot be made open without damaging research relationships. The Digital Science / Figshare 2024 State of Open Data report explicitly acknowledged that these disciplines require “tailored support and resources” rather than the blanket mandates appropriate for quantitative lab sciences.
Incentive Mechanisms and Their Impact
The most extensively studied incentive mechanism in research data sharing is the Open Data Badge system developed by the Center for Open Science. When Psychological Science introduced voluntary badges for data and materials sharing in January 2014, the share of articles reporting available data rose from under 3% to 23% — roughly an eightfold increase. Follow-up analysis confirmed that among badge-bearing articles, 71.2% stored data in an independent repository (compared with fewer than 8% of pre-badge articles), and 100% of the papers that claimed open data actually fulfilled the claim, with 83% providing complete data, 91% providing correct data, and 76% providing computationally usable data.
The ICMJE Data Sharing Statement policy provides a parallel example at the journal-network level. Since the international committee’s requirements took effect, more than 5,000 journals have adopted Data Sharing Statements as a mandatory submission element. The effect is visible in the trend data: declared data availability in health sciences rose from approximately 4% in 2014 to around 9% by 2020, with the steepest increase occurring in the two years immediately following policy adoption at individual journals. Actual verified sharing, however, did not increase at the same rate, reinforcing that statement requirements alone are insufficient.
Emerging incentive structures identified in the 2024 State of Open Data report include the NIH Data Sharing Index, which creates a citation-equivalent metric for data deposits, and the Make Data Count initiative, which assigns formal usage and download statistics to datasets to enable them to appear in academic impact assessments. These mechanisms address the core structural problem: researchers share data less because they receive no professional credit for it.
Geographic Disparities in Data Sharing
Data sharing practices are not uniform across the globe, and the 2024 State of Open Data report provides the clearest regional breakdown to date. Developed nations — the US, UK, Germany, and France — average approximately 25% repository sharing rates for publications with associated datasets. Brazil, Ethiopia, and India remain significantly below that threshold, which the report attributes to differences in institutional infrastructure, researcher time and technical capacity, and the concentration of mandate enforcement in high-income-country funder ecosystems.
Crucially, researchers in lower-resource environments are not less motivated to share: the report found higher citation motivation in Ethiopia and Japan (9.3% and 14.8% respectively, versus 4.88% in the US), suggesting the barrier is primarily infrastructural rather than attitudinal. The implications for global research equity are significant — if funder mandates and repository infrastructure remain concentrated in the Global North, open data may widen rather than narrow the gap between well-resourced and under-resourced research communities.
Within the UK specifically, UKRI’s current award conditions already require data underlying published findings to be openly available wherever possible. The new unified framework under consultation in 2025 is expected to harmonise requirements across BBSRC, EPSRC, MRC, ESRC, and the other research councils, reducing the compliance burden for multi-funder projects and expanding coverage to include software and code as first-class research outputs.
Implications for Postgraduate Researchers
For master’s and doctoral candidates in 2026, research data sharing is no longer a niche concern — it is an increasingly standard expectation embedded in supervisor guidance, ethics approval processes, and journal submission requirements.
Several practical implications follow from the statistics above:
- Plan your DMP before data collection begins. NIH, UKRI, and Horizon Europe all require plans at the grant or project start stage. For funded PhD research, the supervisory team will need to approve a plan early in the project, not at submission. University research offices increasingly provide DMP templates aligned to funder requirements.
- Choose a repository before you write your data availability statement. The statistic that actual sharing drops to 2% across all journals is driven partly by researchers who write availability statements promising to share “upon reasonable request” but never follow through. Depositing in Zenodo, Dryad, or an institutional repository before or at submission eliminates this ambiguity. Our guide to writing a data availability statement covers templates for every scenario including embargoed and restricted-access data.
- Sequence and quantitative data have the most established norms. If your thesis generates biological sequence data, crystallographic data, or quantitative survey data, there are purpose-built repositories (GenBank, CCDC, the UK Data Service) with well-documented submission workflows. These are typically the most straightforward deposits.
- Qualitative data requires specific ethical planning. If your data includes identifiable human participants, the consent process needs to include explicit language about potential data deposition. Retroactively obtaining consent for sharing is often impossible, which is why ethics planning should address data sharing from the outset.
- FAIR compliance is not the same as open access. A dataset can be FAIR — findable, with a persistent identifier; accessible, with clear conditions; interoperable, with standard metadata; reusable, with a clear licence — while still being under access restrictions. The FAIR data principles guide on this site explains how to implement FAIR in restricted-access scenarios.
- Cite your datasets and code properly. Increasingly, thesis examiners and journal reviewers expect datasets and software to be cited like any other source. The guide to how to cite datasets, software, and code covers APA, MLA, Chicago, and Vancouver formats for repository deposits.
For researchers drafting a thesis that incorporates data analysis, AI-assisted tools such as Tesify can help structure the research methodology and data management sections to align with institutional and funder requirements from the outset.
Frequently Asked Questions
What percentage of researchers actually share their data publicly?
A 2023 systematic review of 2.1 million health and medical science articles found that only around 8% declared public data availability, and actual verified sharing dropped to approximately 2% when independently checked. Under mandatory journal policies, the declared rate rises to 65%, with around 33% actual verified sharing. Rates vary widely by discipline and data type, with genomic sequence data at approximately 57% sharing and clinical trial individual participant data at approximately 1%.
When did the NIH data sharing policy become mandatory?
The NIH Data Management and Sharing (DMS) Policy became effective on 25 January 2023. It applies to all NIH-funded research that generates scientific data — regardless of grant size — and requires a Data Management and Sharing Plan at application stage. Annual compliance reporting began in 2024 via the Research Performance Progress Report. Non-compliance can result in award conditions and affect future funding.
How many datasets does Dryad receive each year?
According to Dryad’s 2023–24 Annual Report, the repository released 5,567 new datasets between July 2023 and June 2024. The most active contributing journals were Science, Royal Society Open Science, and PLOS ONE. Dryad also added 14 new institutional and publishing partners during the same period.
Which academic field has the highest rate of data sharing?
Genomics and experimental biology consistently show the highest verified data sharing rates, with sequence data achieving approximately 57% actual sharing. This is driven by decades-old mandatory deposition requirements in leading journals and purpose-built infrastructure such as GenBank and the European Nucleotide Archive. Psychology has seen dramatic increases through Open Data Badge programmes. Health and biomedical sciences range from 6–19% for repository deposits depending on publication venue. Engineering, humanities, and social sciences generally remain below 10%.
Does Horizon Europe require open data sharing?
Yes. Horizon Europe requires all beneficiaries who generate or reuse digital research data to produce a Data Management Plan within six months of project start, make data available via a repository in line with the “as open as possible, as closed as necessary” principle, and apply FAIR principles throughout the data lifecycle. A 2023 ERC analysis found only five repositories — including Zenodo — meet the “Essential” readiness level for Horizon Europe metadata requirements.
Do open data badges increase data sharing compliance?
Yes, substantially. When Psychological Science introduced voluntary Open Data Badges in January 2014, the share of articles reporting available data rose from under 3% to 23% — roughly an eightfold increase. Follow-up analysis found that 71% of badge-bearing articles stored data in an independent repository (up from under 8% before the programme), and 100% of articles claiming open data fulfilled the claim, with 83% providing complete datasets.
Conclusion
The research data sharing statistics for 2026 reveal a field in transition: funder mandates have reached near-universal coverage across NIH-, UKRI-, and Horizon Europe-funded research; repository infrastructure has scaled to millions of records; and badge and incentive programmes have demonstrated that behavioural change is achievable. At the same time, the policy-practice gap remains real — declared sharing consistently outpaces verified sharing, discipline gaps are wide, and the global distribution of both mandates and infrastructure skews heavily toward high-income countries.
For individual researchers, the practical takeaway is clear: plan your data management before data collection, choose a repository before writing your data availability statement, and align your approach to the specific requirements of your funder and target journal. The norms are tightening, and the cost of retroactive compliance — re-contacting participants, renegotiating consent, reformatting years-old datasets — is far higher than building open data practice into the research design from the outset.
Write your thesis with AI
Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.






Leave a Reply