40+ Free Datasets and Open Data Repositories for Your Thesis in 2026 (By Discipline)
Finding raw data to analyse is one of the earliest — and most frustrating — bottlenecks in thesis research. You have a research question, a methodology, and a supervisor’s blessing, but where to find datasets for research that are free, citable, and appropriate for your discipline is rarely explained in methods courses. This guide fixes that. Below are 42 verified repositories organised by discipline, with honest notes on access requirements, data quality, and how to cite what you download.
One distinction matters before you dive in: this guide covers repositories of raw data — survey microdata, administrative records, sensor readings, trade statistics — not repositories of published papers. If you are building a literature review and need peer-reviewed research, see the companion guide to 35+ open-access thesis and dissertation repositories. The platforms below are what you download and analyse yourself.
Before you commit to any repository, it is also worth clarifying your overall research approach. The research methodology guide on Tesify maps the full range of quantitative, qualitative, and mixed-methods designs — and the data types each requires — helping you select repositories that actually match your study design.
1. General & Cross-Disciplinary Repositories
These platforms host data from multiple fields and are the best starting point when your topic crosses disciplinary boundaries, or when you want to explore what exists before committing to a specialist archive.

1. data.gov (United States)
The US federal government’s central open data portal catalogues datasets from more than 100 agencies — agriculture, climate, education, energy, finance, health, housing, and more. Data is downloadable in CSV, JSON, and XML. Access: Truly open; no registration required. Citation tip: Cite the publishing agency (e.g. US Census Bureau or Bureau of Labor Statistics), not data.gov itself.
2. data.europa.eu (European Union)
The official EU open data portal aggregates datasets from EU institutions, agencies, and member states. Strong coverage of economic, agricultural, social, and environmental statistics across 27 countries. Access: Truly open; no registration required. Citation tip: Cite the producing institution (European Commission, Eurostat, a named agency) rather than the portal.
3. Google Dataset Search
A meta-search engine that indexes more than 25 million datasets from thousands of repositories, government portals, and institutional archives. It does not host data itself — it surfaces datasets with structured metadata, making it ideal for an initial scoping search across disciplines. Access: Truly open; links out to the source repository. Citation tip: Cite the original repository, not Google Dataset Search.
4. Zenodo
Hosted by CERN and supported by the European Commission, Zenodo accepts datasets, software, reports, and publications from all disciplines. Each record receives a DOI, and researchers can deposit up to 50 GB per record for free under Creative Commons or custom licences. Access: Truly open to download public records; free registration to upload. Citation tip: Every Zenodo record has a citable DOI — use it and include the access date.
5. re3data (Registry of Research Data Repositories)
re3data does not hold data; it catalogues more than 3,400 data repositories worldwide, each with metadata on subject area, access conditions, data types accepted, and certification status. Use it to identify the right specialist archive for your discipline. Access: Truly open; no registration required. Tip: Filter by subject, content type (raw data, compiled, databases), and open-access flag.
6. Kaggle Datasets
Kaggle hosts more than 80,000 public datasets contributed by the global data science community. Coverage skews toward machine learning benchmarks, tabular data, and popular topics such as health, sport, and finance. Quality varies considerably — prioritise datasets with high vote counts, recent updates, and a clear data card documenting the source. Access: Free registration required. Citation tip: Cite the contributor, dataset title, Kaggle URL, and download date; cite the original institutional source where one is documented.
7. Harvard Dataverse
Part of the global Dataverse network, Harvard Dataverse hosts research data from Harvard-affiliated researchers and the wider academic community. Strong in social sciences, life sciences, and medicine. All datasets carry DOIs. Access: Truly open for public datasets; some datasets require accepting a terms-of-use agreement before download. Citation tip: Use the DOI provided in the dataset metadata panel.
8. OSF (Open Science Framework)
The OSF is a free research platform that lets researchers share data, materials, code, and preregistrations in a single project space. Widely used in psychology, social science, and medicine for reproducibility-focused projects. Access: Truly open for public projects; free registration to create projects and access some private components. Citation tip: Cite the project DOI and the specific component (data file) URL.
9. Figshare
Figshare enables researchers in any field to upload and share datasets, posters, presentations, and code. It is widely used for releasing data that underlies published articles. All public items receive a DOI and are indexed by Google Dataset Search. Access: Truly open for public items; free account to upload. Citation tip: Cite the Figshare item DOI, contributor name, title, and access year.
2. Social Sciences
10. ICPSR — Inter-university Consortium for Political and Social Research
Based at the University of Michigan, ICPSR holds more than 17,000 datasets spanning sociology, political science, demography, education, law, and criminal justice. It is the world’s largest social science data archive and the natural first stop for any social science thesis requiring secondary data. Access: Free for researchers at member institutions — most UK, US, Canadian, and Australian universities are members; registration required. Citation tip: Include the ICPSR study number and DOI in your reference list entry.
11. UK Data Service
The UK’s national data service provides access to major UK government surveys, longitudinal studies (British Cohort Study, Understanding Society, British Social Attitudes), census microdata, and international datasets. Essential for UK-focused social science. Access: Free registration required; end-user licence under UK GDPR. Open to UK higher education and research institutions; international access is available for some datasets. Citation tip: Include Study Number (SN), title, year, and the persistent URL.
12. Eurostat
Eurostat is the statistical office of the European Union, publishing harmonised data on economy, population, trade, and social conditions for all EU member states plus EFTA countries and EU candidate countries. Bulk downloads are available in CSV and SDMX formats. Data go back to 1960 for some series. Access: Truly open; no registration required. Citation tip: Cite as: Eurostat (year), dataset code, and access URL.
13. World Bank Open Data
More than 2,000 development indicators for over 200 countries — covering poverty, education, health, gender equality, infrastructure, and environment — with some series extending to 1960. Downloadable in CSV, Excel, or via a public API. Access: Truly open; no registration required. Citation tip: Cite World Bank (year) with the specific indicator name and code.
14. European Social Survey (ESS)
A biennial cross-national survey measuring attitudes, beliefs, and behaviour in more than 30 European countries, running since 2001. Data from Rounds 1–11 are available, covering topics from immigration and welfare to trust in institutions and wellbeing. Excellent for comparative attitudes research. Access: Free registration required. Citation tip: Cite the ESS round, production year, and edition (e.g. ESS Round 11, Data file edition 1.0).
15. General Social Survey (GSS)
Conducted by NORC at the University of Chicago since 1972, the GSS is the longest-running survey of American society, tracking social attitudes on religion, politics, family, work, and inequality. The cumulative data file covers more than 50 years of data collection. Access: Free; registration required for some microdata downloads. Citation tip: NORC (year). General Social Survey, [years]. Chicago: NORC at the University of Chicago.
16. Pew Research Center Data Archive
The Pew Research Center releases survey datasets from its polling on media consumption, political polarisation, social trends, religion, and global attitudes. Most datasets become available 12–36 months after the corresponding report is published. Access: Free registration required; datasets released under a research licence for non-commercial use. Citation tip: Cite the dataset name, release year, and Pew Research Center as publisher.
3. Health & Biomedical
17. Global Health Data Exchange (GHDx / IHME)
The GHDx, maintained by the Institute for Health Metrics and Evaluation (IHME) at the University of Washington, catalogues population surveys, censuses, vital registries, and IHME’s own Global Burden of Disease estimates — covering mortality, disease incidence, risk factors, and health system capacity. Access: Free for non-commercial use; registration required for IHME-produced datasets specifically. Citation tip: Cite IHME (year), dataset title, and GHDx record URL.
18. WHO Global Health Observatory (GHO)
The WHO’s primary statistics portal provides health-related data for 194 member states — mortality, disease burden, health system capacity, and the SDG health indicators. Data are accessible by indicator, region, and year, with download options in CSV and JSON. Access: Truly open; no registration required. Citation tip: World Health Organization (year), indicator title, and access URL.
19. CDC WONDER
CDC WONDER (Wide-ranging Online Data for Epidemiologic Research) provides access to US public health databases: mortality, cancer incidence, natality, infectious disease surveillance, and environmental public health data at county, state, and national level. Access: Truly open for most datasets; no registration required. Users accept data-use terms per query session. Citation tip: Cite Centers for Disease Control and Prevention (year) and name the specific database used (e.g. “Multiple Cause of Death, 1999–2022”).
20. NHS England Open Data
NHS England publishes a wide range of secondary care, primary care, urgent care, and waiting-times data through its statistics portal at digital.nhs.uk. Particularly useful for healthcare management, public health, and health policy dissertations in the UK. Access: Truly open for aggregate statistics; no registration required. Patient-level data requires a formal data access application. Citation tip: Cite NHS England (year) and the specific publication title and URL.
21. PhysioNet
Supported by MIT and the US National Institutes of Health, PhysioNet hosts databases of physiological signals — ECG, EEG, blood pressure, gait, polysomnography — used in biomedical engineering, clinical informatics, and critical care research. The MIMIC database (ICU records) is the most widely used. Access: Free registration required; MIMIC and similar databases require additional CITI ethics training and credentialing before access is granted (restricted). Citation tip: Cite the specific PhysioNet database name and DOI.
22. OpenNeuro
OpenNeuro hosts brain imaging datasets — fMRI, EEG, MEG, PET — in the standardised BIDS (Brain Imaging Data Structure) format, enabling reproducible analysis pipelines. Hundreds of publicly contributed datasets from cognitive and clinical neuroscience are available. Access: Truly open; free registration to download datasets. Citation tip: Cite the dataset DOI and BIDS format version alongside the original study reference.
4. Economics & Finance
23. FRED — Federal Reserve Economic Data
Maintained by the Federal Reserve Bank of St. Louis, FRED aggregates more than 800,000 time-series datasets from over 100 national and international sources: the US Bureau of Labor Statistics, the Bureau of Economic Analysis, the IMF, the World Bank, the OECD, and national central banks worldwide. It is the single most useful starting point for macroeconomics theses. Access: Truly open; no registration required. A public API is available. Citation tip: Cite as: Federal Reserve Bank of St. Louis (year), series name, FRED series ID, and access URL.
24. IMF Data
The International Monetary Fund’s data portal provides macroeconomic and financial statistics including the World Economic Outlook (WEO) database, Balance of Payments Statistics, and Financial Soundness Indicators — covering the IMF’s 190 member countries. Access: Truly open; no registration required. Data downloadable in CSV and SDMX. Citation tip: Cite the specific IMF dataset by name and the WEO vintage year.
25. OECD iLibrary
Home to OECD’s statistical databases — OECD.Stat, Education at a Glance, Health Statistics, Revenue Statistics, and dozens more — covering 38 member countries and partner economies. All OECD content became fully Open Access in July 2024. Access: Truly open; no registration required. Citation tip: Cite the database name, variable, and access year (OECD data are updated continuously, so the access date matters).
26. Our World in Data
Published by the Oxford-based Global Change Data Lab, Our World in Data compiles and repackages data from international organisations (WHO, World Bank, FAO, UN) on long-run trends in health, poverty, energy, education, and environment. All charts and underlying data are released under CC BY 4.0 — cite the original source. Access: Truly open; no registration required. Source datasets are linked from each chart. Citation tip: Cite the original source data, not only Our World in Data, in your reference list.
27. UN Comtrade
The United Nations Comtrade database holds detailed import and export trade statistics reported by national statistical offices for more than 200 countries and territories, covering HS commodity codes from 1962 onward. Access: Free registration required; a generous free tier allows up to 100 API calls per hour. Bulk downloads require a premium subscription. Citation tip: Cite as: United Nations Comtrade (year), trade flow, reporter country, partner, commodity code, and access date.
5. Education
28. PISA — Programme for International Student Assessment
Run by the OECD every three years, PISA assesses reading, mathematics, and science literacy in 15-year-olds across 90+ countries. Microdata for each cycle (most recently PISA 2022, released 2023) are available for secondary analysis in SPSS, SAS, and CSV formats — making PISA one of the richest comparative education datasets available to thesis students. Access: Free download; no registration required. Citation tip: OECD (year), PISA [year] Results, and the specific database volume or table used.
29. NCES Data & DataLab (United States)
The National Center for Education Statistics releases datasets from NAEP (National Assessment of Educational Progress), IPEDS (postsecondary institutions), the High School Longitudinal Study, and the Early Childhood Longitudinal Study. The DataLab online tool lets you run custom cross-tabulations without downloading the full microdata file. Access: Truly open for aggregate and restricted-use public files; restricted microdata require a licence application. Citation tip: Cite the specific survey name, year, and NCES publication number.
30. UK Department for Education — Explore Education Statistics
The DfE’s Explore Education Statistics platform (explore-education-statistics.service.gov.uk) publishes school performance tables, pupil census data, GCSE and A-level results, higher education participation statistics, and teacher workforce data for England. Scotland, Wales, and Northern Ireland have equivalent national portals. Access: Truly open; no registration required. Citation tip: Cite Department for Education (year), publication title, and GOV.UK URL.
6. Environment & Earth Sciences
31. NASA Earthdata
NASA Earthdata is the gateway to NASA’s entire Earth observation archive — more than 12,400 datasets spanning land cover and use, ocean colour, atmospheric composition, sea level, ice sheets, soil moisture, and wildfire detection, with records extending back to the 1970s in some instruments. Access: Free Earthdata account required (quick online registration); all datasets are free once registered. Citation tip: Cite the dataset shortname, version, DOI, and the Earthdata platform.
32. NOAA Open Data Dissemination (NODD)
NOAA’s NODD programme makes operational forecast models, historical climate records, ocean data, and severe weather data freely available through cloud storage partners (AWS, Google Cloud, Microsoft Azure). Key datasets include GOES satellite imagery, the Global Historical Climatology Network (GHCN), and Global Surface Summary of the Day (GSOD). Access: Truly open; no registration required for most datasets. Citation tip: Cite the specific NOAA dataset name, product version, and access date.
33. Copernicus Climate Data Store (CDS)
Operated by the European Centre for Medium-Range Weather Forecasts (ECMWF) under the EU Copernicus Earth Observation Programme, the CDS provides ERA5 reanalysis data (global hourly climate from 1940 onward), seasonal forecasts, and climate projections. ERA5 is now one of the most-cited climate datasets in academic literature. Access: Free registration required; all data are then free to download under a Copernicus licence. Citation tip: Cite the dataset DOI and the CDS record identifier alongside the ECMWF author team.
34. Global Forest Watch (GFW)
Operated by the World Resources Institute, GFW provides near-real-time monitoring of global deforestation, fire alerts, and land cover change using Landsat and Sentinel satellite data. Tree cover loss data are updated annually and used widely in environmental economics and conservation biology theses. Access: Truly open; no registration required for data downloads. Citation tip: Cite Hansen et al. for the underlying tree cover loss data and the GFW platform URL and access year.
35. SEDAC — Socioeconomic Data and Applications Center
A NASA Distributed Active Archive Center (DAAC), SEDAC specialises in the intersection of human society and the environment — population density grids, climate vulnerability indices, urban extent maps, and natural disaster risk data. It requires the same free Earthdata account used for NASA Earthdata (#31). Access: Free Earthdata account required. Citation tip: Cite the SEDAC dataset DOI, dataset name, and version.
7. Psychology & Behaviour
36. MIDUS — Midlife in the United States
A longitudinal survey run by the University of Wisconsin-Madison tracking a large national sample of US adults across psychological wellbeing, physical health, biomarkers, and social factors. Now in its third wave of data collection, MIDUS is one of the most widely used datasets in lifespan developmental psychology. Access: Free registration via ICPSR (Study 2760). Citation tip: Cite as Brim et al. (year), MIDUS dataset, ICPSR study number, and DOI.
37. ANES — American National Election Studies
ANES is the longest-running survey of political attitudes and behaviour in the United States, conducted since 1948. It covers voting behaviour, party identification, candidate evaluation, media use, and social trust — making it the primary data source for political psychology, public opinion, and electoral behaviour dissertations. Access: Free registration required. Citation tip: American National Election Studies (year), dataset edition title, and DOI.
38. Databrary
Databrary is a specialised repository for video and audio research data in developmental and learning science, enabling the sharing of observational recordings within appropriate consent frameworks. Operated jointly by NYU and Penn State. Access: Free registration; institutional access requires your university to be a Databrary signatory institution (restricted — confirm before relying on it). Citation tip: Cite the specific volume, session, and contributor as listed in the Databrary record.
8. Humanitarian & Development
39. Humanitarian Data Exchange (HDX)
Run by the UN Office for the Coordination of Humanitarian Affairs (OCHA), HDX is the go-to portal for conflict, displacement, food insecurity, and crisis data. Datasets from UNHCR, the World Food Programme, MSF, and hundreds of NGOs are freely available and regularly updated. Access: Truly open; no registration required for download. Citation tip: Cite the contributing organisation, dataset title, HDX URL, and access date.
40. UN SDG Indicators
The United Nations SDG Global Database tracks progress on all 169 targets of the 2030 Agenda for Sustainable Development, with data submitted by national statistical offices for 193 countries from 2000 onward. Useful for any thesis engaging with development, climate policy, or social equity across countries. Access: Truly open; downloadable in CSV via the UN Stats Division portal. Citation tip: United Nations Statistics Division (year), SDG Indicators, specific goal and indicator code.
41. Dryad
Dryad is a curated repository for research data underlying peer-reviewed publications across the biological, medical, and social sciences. All data are released under CC0 (public domain), making them straightforward to reuse without licence complications. Access: Truly open; no registration required. Citation tip: Cite the Dryad DOI alongside the associated publication — both are needed for full provenance.
42. Registry of Open Data on AWS
Amazon Web Services hosts a catalogue of large-scale publicly available datasets — human genome sequencing data, satellite imagery archives, geospatial datasets, and the Common Crawl web corpus — stored in S3 and accessible without data-transfer fees to AWS users. Best suited to computationally intensive theses in bioinformatics, NLP, or remote sensing. Access: Truly open; a free AWS account is sufficient for most research use. Citation tip: Cite the dataset name, licence, and the AWS Open Data Registry URL.
Quick-Reference: Access Level at a Glance
| Repository | Discipline | Access |
|---|---|---|
| data.gov | General | Truly open |
| data.europa.eu | General | Truly open |
| Google Dataset Search | General (meta) | Truly open |
| Zenodo | General | Truly open (public records) |
| re3data | General (registry) | Truly open |
| Kaggle Datasets | General | Free registration |
| Harvard Dataverse | General | Truly open / terms agreement |
| OSF | General | Truly open (public projects) |
| Figshare | General | Truly open (public items) |
| ICPSR | Social Sciences | Free (member institution) |
| UK Data Service | Social Sciences | Free registration (UK HEI) |
| Eurostat | Social Sciences | Truly open |
| World Bank Open Data | Social / Economics | Truly open |
| European Social Survey | Social Sciences | Free registration |
| GSS | Social Sciences | Free / registration |
| Pew Research Datasets | Social Sciences | Free registration |
| GHDx / IHME | Health | Free (non-commercial) |
| WHO GHO | Health | Truly open |
| CDC Wonder | Health | Truly open |
| NHS England Open Data | Health (UK) | Truly open (aggregate) |
| PhysioNet (MIMIC) | Health / Engineering | Registration + credentialing |
| OpenNeuro | Neuroscience | Free registration |
| FRED | Economics | Truly open |
| IMF Data | Economics | Truly open |
| OECD iLibrary | Economics / Education | Truly open (OA since Jul 2024) |
| Our World in Data | Multi-disciplinary | Truly open (CC BY 4.0) |
| UN Comtrade | Economics / Trade | Free registration (rate-limited) |
| PISA | Education | Truly open |
| NCES / NAEP DataLab | Education (US) | Truly open (aggregate) |
| UK DfE Statistics | Education (UK) | Truly open |
| NASA Earthdata | Environment | Free registration |
| NOAA NODD | Environment | Truly open |
| Copernicus CDS (ERA5) | Environment | Free registration |
| Global Forest Watch | Environment | Truly open |
| SEDAC | Environment / Society | Free registration (Earthdata) |
| MIDUS | Psychology | Free registration (via ICPSR) |
| ANES | Psychology / Political | Free registration |
| Databrary | Psychology / Education | Registration + institutional |
| HDX (OCHA) | Humanitarian | Truly open |
| UN SDG Indicators | Development | Truly open |
| Dryad | Multi-disciplinary | Truly open (CC0) |
| Registry of Open Data (AWS) | Multi-disciplinary | Truly open (free AWS tier) |
How to Cite a Dataset in Your Thesis
Citing data sources is not optional — examiners treat undocumented data the same as missing references. The core elements are: author or originating body, year of release or last update, dataset title, version (if applicable), repository name, and a persistent identifier (DOI preferred over URL where one exists). For a full walkthrough with APA 7 and Vancouver worked examples, see our dedicated guide on how to cite a dataset, software, and code in APA 7 and Vancouver.
Before you download, check the dataset’s licence. CC0 and CC BY datasets can be reused with minimal restriction. Some repositories (MIDUS, UK Data Service, PhysioNet) issue end-user licences that prohibit commercial use or require data destruction after the project ends — record the licence alongside your citation metadata.

If you are using secondary data — data collected by someone else — also check whether your institution requires ethics committee or IRB review before you begin analysis. Many fully de-identified, publicly available datasets qualify for exemption, but this varies by institution and by the sensitivity of the variables involved. The complete guide on whether you need IRB approval for secondary data research walks through the US Common Rule exemptions and UK GDPR considerations in detail.
FAQ
Where is the best place to find datasets for research?
Start with Google Dataset Search for a broad scoping search across 25 million+ indexed datasets. Then move to your discipline’s specialist repository: ICPSR for social sciences, FRED for economics, the Global Health Data Exchange for global health, NASA Earthdata for environmental science, and Zenodo for any field. Kaggle and Harvard Dataverse are reliable cross-disciplinary alternatives when specialist archives do not cover your specific topic.
Are free datasets reliable enough for university research?
Yes — many of the most-cited datasets in academic research come from the repositories in this list. Government portals (Eurostat, ONS, BLS), international organisations (World Bank, WHO, IMF), and institutional archives (ICPSR, UK Data Service) apply rigorous quality-control standards. Community-contributed platforms such as Kaggle vary considerably; check the dataset documentation, collection methodology, and update frequency before committing to a dataset from these sources.
Do I need ethics approval to use publicly available datasets?
In most cases, no. Using fully de-identified, publicly available datasets qualifies for exemption under the 2018 Common Rule (US) and equivalent UK frameworks. However, some datasets with quasi-identifiers or sensitive health variables require an institutional review regardless of their public availability. Check with your supervisor before starting analysis — requirements vary by institution and by the nature of the variables you plan to use.
What is the difference between a data repository and a literature database?
A data repository stores raw or processed data — survey responses, sensor readings, administrative records, genomic sequences — that researchers download and analyse themselves. A literature database such as PubMed, Scopus, or Web of Science indexes published articles and abstracts. Some platforms overlap (Zenodo stores both datasets and articles), but the distinction matters for your methodology chapter: a data repository is where you find primary empirical material, while a literature database is where you find published evidence to synthesise in your review.
How do I know if a dataset is appropriate for quantitative analysis?
A dataset is suitable for quantitative analysis if it contains numeric or ordinal variables with a documented sampling strategy, a clear codebook or data dictionary, and a known (or estimable) target population. Check the documentation for sample size, collection period, and known limitations before committing. For guidance on structuring your analysis approach, see the step-by-step guide to quantitative research methods.
Can I use Kaggle datasets in a university thesis?
Yes, with caveats. Many Kaggle datasets carry CC0 or CC BY licences and are fully citable. The key risks are provenance (not all contributors document how data were collected) and persistence (datasets can be updated or removed). Prefer datasets that link back to an original institutional or government source, document your download date and version, and — where a primary source exists — cite that source directly rather than only the Kaggle page.
Plan your data chapter before you start downloading
Having the data is only half the challenge. Before you commit to a specific repository, you need a clear methodology: which variables you will analyse, which statistical tests you will apply, and how the results map to your research questions. Tesify helps you build that framework — defining your research design, structuring your methodology chapter, and connecting your data sources to your thesis argument — so you arrive at the analysis stage with a coherent plan rather than a folder of CSV files and no clear next step.
For the chapter itself, the guide on how to write a dissertation methodology chapter walks through exactly what to include when documenting secondary data sources for an examiner.
If you have not yet settled on your overall research approach, the research methodology guide on Tesify maps the full spectrum of quantitative, qualitative, and mixed-methods designs — helping you match your study design to the right data sources before you start downloading.
Write your thesis with AI
Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.






Leave a Reply