Where to Get Data for a Political Science or International Relations Thesis: V-Dem, Polity5 and 8 More Sources (2026)

·

Key finding: Political science and international relations theses draw on a small set of well-established cross-national databases — V-Dem for democracy measurement, Polity5 and Freedom House for regime type, UCDP/PRIO and Correlates of War for conflict, the World Bank’s Worldwide Governance Indicators for institutional quality, and survey programmes such as the World Values Survey for public opinion — most of which are free to download and citable by a standard dataset DOI or working paper.

This guide sorts ten sources by what kind of question they actually answer, not just by name, so you can match a data source to your research question rather than reverse-engineering a question from whatever dataset you found first. It complements the discipline-by-discipline 40+ free datasets and open data repositories directory, which lists ICPSR as one general social-science entry point; the sources below go one level deeper into what a political science or IR thesis specifically needs.

1. V-Dem (Varieties of Democracy)

Based at the University of Gothenburg, V-Dem is the most widely used democracy-measurement dataset in current political science, covering more than 200 countries and territories across five high-level democracy principles: electoral, liberal, participatory, deliberative and egalitarian. Version 16, released in 2026, holds more than 600 indicators built from expert coder ratings aggregated through a measurement model rather than a single coder’s judgement. Access: free download from v-dem.net; both the raw indicators and pre-built composite indices are available. Best for: any thesis comparing democratic quality across countries or tracking backsliding over time, since the disaggregated principles let you test a specific claim (e.g., about judicial constraints) rather than relying on a single blended democracy score.

2. Polity5

Maintained by the Center for Systemic Peace, Polity5 codes each country-year on a −10 (full autocracy) to +10 (full democracy) scale based on the competitiveness and openness of executive recruitment, constraints on executive authority, and political competition. It has one of the longest time series in the field, running from 1800 for many states, but its coverage currently ends in 2018, so pair it with V-Dem if your period runs past that year. Access: free download, no registration required. Best for: theses needing the longest possible historical time series or wanting to replicate the large existing body of quantitative IR literature that uses the Polity score as a control variable.

3. Freedom House “Freedom in the World”

Illustration of matching research questions to political science datasets
Match the dataset to the question, not the other way round.

An annual expert-scored index rating political rights and civil liberties on a 1–7 scale for every country and a set of territories, published since the 1970s. Its methodology is more transparent about scoring criteria than some alternatives, which makes it easier to justify to a committee, though it has also drawn methodological criticism for potential Western-liberal bias — a limitation worth naming explicitly if you use it. Access: free, downloadable country and aggregate data tables from freedomhouse.org.

4. World Bank Worldwide Governance Indicators (WGI)

Six composite indicators — voice and accountability, political stability, government effectiveness, regulatory quality, rule of law, and control of corruption — aggregated from dozens of underlying expert and survey sources, updated annually for over 200 economies. Access: free via the World Bank’s Data Catalog and API. Best for: theses on governance quality, corruption, or institutional determinants of economic or development outcomes, since the six sub-indicators let you isolate the specific governance dimension your theory predicts matters.

5. UCDP/PRIO Armed Conflict Dataset

A joint product of the Uppsala Conflict Data Program and the Peace Research Institute Oslo, this is the standard reference dataset for state-based armed conflict, coding conflict onset, actors, conflict type (interstate, intrastate, extrasystemic) and battle-related deaths from the mid-twentieth century onward. UCDP also maintains a separate Georeferenced Event Dataset (GED) for sub-national, event-level conflict data. Access: free from ucdp.uu.se. Best for: conflict onset, duration, or civil-war-specific theses; use GED rather than the country-year file if your question is about location or timing at finer resolution than the calendar year.

6. Correlates of War (COW) Project

One of the oldest continuously maintained IR datasets, COW provides state system membership lists, militarized interstate disputes (MID), national material capabilities (the CINC composite index of military and economic power), and alliance data. Access: free from correlatesofwar.org. Best for: theses working in the realist or power-transition traditions that need a standard operationalisation of state capability or historical alliance structure; note some COW sub-datasets update on a slower cycle than UCDP, so check the most recent coding date before committing to it as your primary source.

7. World Values Survey (WVS) and regional barometers

The WVS has run representative national surveys on values, trust, political engagement and social attitudes across roughly 100 countries since the early 1980s, in coordinated waves. For sub-national or regional public-opinion questions, the regional barometer surveys — Afrobarometer, Latinobarómetro, Asian Barometer, and the Eurobarometer for EU member states — often ask more region-specific questions than the WVS core module. Access: free registration and download for all of these. Best for: any thesis on political trust, populism, democratic satisfaction or civic attitudes, where cross-national survey microdata (not just aggregate country scores) is what your hypotheses actually need.

8. Comparative Manifestos Project / MARPOR

Codes the policy content of political party election manifestos across democracies going back to 1945, producing left-right and issue-salience scores derived from sentence-level content coding rather than expert judgement about the party as a whole. Access: free with registration via manifesto-project.wzb.eu. Best for: theses on party competition, issue ownership, or how manifesto rhetoric tracks (or fails to track) governing behaviour once in office.

9. ParlGov

A single, continuously updated database of election results, cabinet composition and party positions for parliamentary democracies, useful specifically because it standardises party and election identifiers across countries, which is otherwise a persistent headache when merging national election data. Access: free from parlgov.org. Best for: comparative theses on coalition formation, cabinet duration or electoral system effects that need a clean, cross-nationally consistent party-level dataset.

10. GDELT (Global Database of Events, Language and Tone)

An automated, machine-coded event dataset built from global news monitoring, updated in near real time and covering a far larger volume of events than any hand-coded alternative, at the cost of more noise and less validated coding accuracy per event. Access: free, queryable via Google BigQuery or bulk download. Best for: theses that specifically need event-level, high-frequency or very recent data (crisis escalation, protest waves, media framing) where hand-coded datasets like UCDP or COW have not yet caught up, provided you address coding-reliability limitations directly in your methods chapter. If your design tries to draw a causal claim from any of these observational data sources rather than a purely descriptive comparison, work through directed acyclic graphs for causal inference before finalising your model — cross-national panel data is exactly the setting where unaddressed confounding is easiest to miss.

Matching a source to your research question

Your question is about… Start with
Democratic quality or backsliding over time V-Dem (disaggregated) or Polity5 (long time series, to 2018)
Institutional quality / corruption / governance World Bank WGI
Civil war, interstate conflict or conflict duration UCDP/PRIO (country-year or GED)
Power, capability or alliance structure Correlates of War
Public trust, populism, or civic attitudes World Values Survey / regional barometer
Party competition or manifesto content Comparative Manifestos Project
Very recent or high-frequency events GDELT (with caveats on coding noise)

Citing these sources correctly

Illustration of a researcher merging cross-national datasets for a thesis
Merge on one country-identifier system and log every version you download.

Most of these datasets publish a specific, versioned citation (author list, version number, year) on their own website — cite that exact version, not just the project name, since coding and coverage can change between releases (V-Dem alone has moved through more than a dozen major versions). Keep the exact download date and version number in your methods chapter or an appendix; a reviewer re-checking your numbers against a newer version is one of the most common sources of “your figures don’t match” queries at defence. Once your merged data set is built, free statistical software for students covers the tools (R, JASP, jamovi) most political science and IR theses use to analyse cross-national panel data at no licensing cost.

Licensing and data ethics for secondary cross-national data

Because none of the ten sources above involve you collecting data from human participants directly, most institutional review boards treat secondary analysis of these datasets as exempt or minimal-risk, but that is a determination your IRB or ethics committee makes, not an assumption you should write into your methods chapter unchecked. Check two things specifically: whether the dataset’s terms of use restrict redistribution of the raw microdata (most do, which affects what you can put in a public thesis repository appendix), and whether any individual-level survey microdata (as opposed to aggregated country scores) carries additional confidentiality conditions. The World Values Survey and the regional barometers, for instance, release de-identified individual respondent records, and their terms of use typically require citing the specific wave and prohibit attempting re-identification.

How many countries or years is “enough” for a cross-national design

There is no universal minimum, but committees generally expect the sample to be justified by your research question and estimation strategy rather than by whatever the dataset happens to include. A study using country-fixed-effects panel regression needs enough within-country variation over time to identify an effect, which argues for a longer time series (Polity5 or COW) over a source with only a handful of recent waves; a study comparing regime types at one point in time can use a cross-sectional slice of V-Dem or the WGI with a much larger country count instead. State the logic behind your case selection explicitly, including which countries or country-years you excluded and why.

Where Tesify fits

Once you have identified and downloaded the right dataset, Tesify’s thesis workspace, used by 9,000+ students and 15,000+ chapters, helps you structure the data-and-methods section. Every word stays 100% written by you, and Tesify cannot select a dataset for you or verify a citation you have not opened yourself, so treat the sourcing work above as the non-negotiable first step.

Frequently asked questions

Do I need to register to access these datasets?

Some do: the World Values Survey and the Comparative Manifestos Project ask for a simple registration (name, institution, intended use), while V-Dem, Polity5, Correlates of War and GDELT offer direct downloads. Access terms change, so check each project’s download page before you plan your timeline.

Which dataset should I use if my thesis compares democracies and autocracies?

V-Dem’s disaggregated principles let you test specific claims (for example, about media freedom versus judicial independence) rather than relying on a single blended score; Polity5 is a reasonable alternative if you specifically need the longest possible historical time series and your period ends by 2018.

Can I combine several of these datasets in one thesis?

Yes, and it is common — for example, merging V-Dem democracy scores with UCDP conflict data to test whether democratic backsliding predicts conflict onset — but you must merge on a consistent country-identifier system (the Correlates of War numeric codes or ISO-3 codes are the usual choices) and document any country-year mismatches.

Is GDELT reliable enough for a thesis?

It can be, for research questions that specifically need event-level or very recent data, but its automated coding is noisier than hand-coded datasets like UCDP or COW, so pair it with a validation step (spot-checking a sample of coded events against news sources) and say so explicitly in your limitations section.

What if the dataset I need only covers a specific region?

Use the matching regional barometer (Afrobarometer, Latinobarómetro, Asian Barometer) rather than forcing the World Values Survey’s global module to answer a question it was not designed to ask regionally.

How do I handle missing country-years in these datasets?

Report the extent of missingness for your specific sample and country-years, state whether you used listwise deletion or an imputation method, and check whether the missingness is itself substantively meaningful (states that collapse or are excluded from coding are rarely missing at random). If you need to report an estimate’s precision rather than just a point value, see how to interpret a confidence interval for the reporting conventions reviewers expect.

Are these datasets appropriate for an undergraduate thesis, or only graduate-level work?

Most are entirely appropriate at undergraduate level for a focused, well-scoped question (a handful of countries or a specific region over a defined period); the skill being assessed is precise use of an existing dataset, not building one from scratch.

Do I need IRB or ethics approval to use these secondary datasets?

Usually a lighter-touch exempt or minimal-risk review, since you are not collecting new data from participants yourself, but confirm this with your own institution’s ethics committee rather than assuming it — the determination is theirs to make, and some committees still require a short secondary-data-use application.

Can I request a custom extract instead of downloading the full dataset?

Several of these projects (the World Bank’s Data Catalog, GDELT via BigQuery, and some WVS access points) support querying a subset directly rather than downloading the entire file, which is worth doing if your country or time-period selection is narrow and the full file is large.

Write your thesis with AI

Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.

Leave a Reply

Your email address will not be published. Required fields are marked *