Identification Strategy for an Economics Dissertation: Fixed Effects, DiD, IV or RD — and the Data Each One Needs (2026)
Having data is not having an answer. The examiner’s question is always the same: why should we believe this correlation is causal? An economics dissertation is graded on how convincingly you answer it — on which variation in your data identifies the effect, and on whether you have defended the assumption that makes that variation credible. The dataset is a means to that end, never the point.
This article is the identification-strategy specialisation of our broader guide on how to write an economics thesis, which covers the discipline end to end from research question to final chapter. Here we go deep on one decision: the five strategies that carry dissertation-scale empirical work, what each one assumes, how examiners test that assumption — and, in the second half, twenty sources that actually supply the variation each design requires.
Choosing an identification strategy
Five strategies cover the vast majority of dissertation-scale empirical work. Pick yours before you pick your dataset, not after: the design tells you what variation you need, and only then does the search for data become tractable.
Panel fixed effects
With repeated observations on the same unit, entity and time fixed effects absorb all time-invariant unobserved heterogeneity and all common shocks. This is the accessible workhorse. It does not solve reverse causality or time-varying confounders, so be precise in your claims: fixed effects rules out a class of confounders, not all of them. Cluster your standard errors at the level of treatment assignment, and say that you have done so.
Difference-in-differences
A policy applied to some units and not others, with data before and after, gives you a treatment × post interaction. State the parallel-trends assumption explicitly and support it with a pre-trend plot — an event-study specification with leads and lags is now the expected presentation, not a bonus. Be aware of the recent literature on staggered adoption: when units are treated at different times, the standard two-way fixed effects estimator can be biased, and modern estimators exist to address it. Citing that debate demonstrates you are reading current work.
Instrumental variables
An instrument must be relevant (strongly correlated with the endogenous regressor) and valid (affecting the outcome only through it). Report the first-stage F statistic; below roughly 10 you have a weak instrument and your two-stage estimates are unreliable. The exclusion restriction cannot be tested, only argued, so devote a full paragraph to defending it. Weak, borrowed instruments are among the most heavily criticised features of student empirical work.
Regression discontinuity
Where treatment is assigned by a threshold — an income cut-off for a benefit, a test score for admission, a firm-size threshold for a regulation — comparing units just either side gives credible local identification. Show the density of the running variable to rule out manipulation, and demonstrate that your result survives changes in bandwidth and polynomial order.
Honest descriptive work
A carefully constructed, well-documented descriptive analysis of an underused dataset is a legitimate dissertation and often scores better than a badly identified causal claim. If your design cannot support causality, say “associated with” throughout and mean it.
Whichever route you take, the mechanics of your outcome variable still determine your estimator: continuous outcomes to linear models, binary participation outcomes to probit or logit — our comparison of linear versus logistic regression covers when the linear probability model is defensible and when it is not. Where individuals are grouped within regions or firms within industries, our guide to multilevel models for nested data explains the alternative to simply clustering standard errors.
Where the variation comes from: twenty sources you can actually obtain
A design is only as good as the data that feeds it. Cross-country panels support fixed effects but rarely a discontinuity; regional and administrative data support DiD and RD because policies change at borders and thresholds; household panels support within-person identification. Below are twenty sources genuinely available to a student in 2026, grouped by what they let you study, with the limitation that most often catches people out. For a discipline-by-discipline directory beyond economics, see our reference list of free datasets and open data repositories for a thesis.
Macroeconomic and cross-country panels
1. World Bank World Development Indicators. The default starting point for cross-country work: roughly 1,400 indicators covering growth, poverty, health, education, trade and infrastructure for over 200 economies from 1960 onward. Free, with a bulk download and an API. Limitation: coverage is deeply unbalanced — low-income countries have large gaps precisely where the interesting variation is, so check your panel’s balance before committing to a fixed-effects specification.
2. IMF data portal (World Economic Outlook and International Financial Statistics). Best for macro aggregates, fiscal and monetary series, balance of payments and exchange rates, plus IMF forecasts you can use as an expectation measure. Limitation: WEO vintages are revised, so cite the release month you downloaded.
3. Penn World Table. The standard source for cross-country comparisons of real income, output, inputs and productivity using purchasing power parities. Essential for anything on growth accounting or convergence. Limitation: versions are not interchangeable; results can shift meaningfully between releases, so state the exact version.
4. OECD.Stat. Detailed and comparatively harmonised data for member economies: labour markets, taxation, social expenditure, productivity, regional accounts. Limitation: membership is a selected sample, which restricts external validity.
5. Eurostat. The richest harmonised source for European work, including regional NUTS-level data that supports within-country identification and much finer geography than most cross-country panels allow.
6. FRED (Federal Reserve Bank of St Louis). Over 800,000 US and international time series with an excellent interface, an API and Excel add-in. Best-in-class for anything monetary, financial or high-frequency US macro. Practical tip: FRED’s vintage database, ALFRED, lets you use data as it was originally published rather than as later revised, which matters for anything about policy in real time.
7. Our World in Data. Not a primary source, but a well-documented harmonisation layer over many of the above, with clear provenance for each series. Use it for exploration and to find the primary source; cite the primary source in your dissertation.
Household, labour and micro survey data
8. UK Data Service. The main route for UK microdata: the Labour Force Survey, Family Resources Survey, Living Costs and Food Survey, Wealth and Assets Survey and much else. Registration is free for students and most datasets are available under an End User Licence. Our step-by-step walkthrough of registering with the UK Data Service and downloading data covers the account, the licence tiers and the file formats.
9. Understanding Society (UK Household Longitudinal Study). A large annual household panel following the same individuals over time, with income, employment, health, wellbeing and household composition. The single best UK resource for anything requiring within-person variation. Limitation: the file structure is complex and merging waves takes real time — budget a fortnight.
10. IPUMS. Harmonised census and survey microdata: IPUMS-CPS for the US Current Population Survey, IPUMS International for census extracts from more than 100 countries, IPUMS Health Surveys. Free, with an extract system that lets you request only the variables you need.
11. Panel Study of Income Dynamics. The longest-running household panel anywhere, following US families since 1968 and now into multiple generations. Unmatched for intergenerational mobility questions.
12. Living Standards Measurement Study and the Demographic and Health Surveys. The workhorses of development economics microdata, covering consumption, agriculture, health and education across low- and middle-income countries. Both are free after registration.
13. European Social Survey, World Values Survey and Afrobarometer. Repeated cross-national attitude surveys with careful sampling documentation. Useful for anything at the boundary of economics and political economy — trust, preferences for redistribution, institutional confidence.
14. SHARE (Survey of Health, Ageing and Retirement in Europe). Longitudinal, cross-country data on older adults covering health, retirement decisions, pensions and family transfers.
Several of these are longitudinal, which changes what your dissertation can claim; our guide to UK longitudinal cohort studies compares the major panels on cohort, coverage and access route.
Firm, financial and trade data
15. Kenneth French’s Data Library. Free, canonical factor returns and portfolio sorts for asset pricing work — the market, size, value, profitability, investment and momentum factors, in monthly, daily and annual frequency. If your dissertation runs a factor model, this is where the right-hand side comes from.
16. WRDS (CRSP, Compustat, IBES). The standard for empirical finance and industrial organisation, but licensed — check whether your business school subscribes before designing around it. Practical tip: access is often granted to taught master’s students on request even where it is not advertised.
17. Orbis / Bureau van Dijk and the World Bank Enterprise Surveys. Firm-level accounts and ownership for Orbis (subscription); nationally representative firm surveys covering access to finance, informality and constraints for Enterprise Surveys (free).
18. UN Comtrade and the CEPII gravity database. Bilateral trade flows by product and partner, plus CEPII’s ready-made distance, contiguity, language and colonial-tie variables. Together they let you estimate a gravity model in an afternoon rather than a month of variable construction.
19. Bank for International Settlements. Credit-to-GDP series, property prices, effective exchange rates and cross-border banking statistics — the standard source for anything on financial cycles.
20. Institutional and policy indices. V-Dem for democracy measures, the Chinn-Ito index for capital account openness, the Economic Policy Uncertainty indices, the KOF Globalisation Index, and Barro-Lee for educational attainment. All free, all constructed indices — read the methodology before using one as a right-hand-side variable, and never treat a composite index as an objective measurement.
Two further routes worth remembering for UK-focused topics. Administrative data you cannot find published can sometimes be obtained directly: our guide to using Freedom of Information requests for a dissertation covers how to write a request that is actually answered. And for secure access to detailed government microdata, our explainer on whether a student can use the ONS Secure Research Service sets out the accreditation reality.
Practical rules that save weeks
- Download the data before you write the proposal. Not the documentation — the actual file. Confirm your key variable exists, for your countries, for your years.
- Check the panel balance immediately. A cross-country study of 190 economies frequently becomes 47 once every variable in your specification is required simultaneously.
- Record the exact version and download date of every source. Revisions are routine and irreproducible results are penalised.
- Write your cleaning as a script, never by hand in Excel. You will need to re-run it, and a documented pipeline plus a data dictionary and cleaning log is what makes your results reproducible.
- Deflate and convert consistently. Decide once on your price base year and PPP versus market exchange rates, document it, and apply it everywhere.
- Plan for a null result. Well-identified nulls are publishable and markable; our guide on what to do when results are not significant shows how to frame one as a contribution rather than a failure.
Get the empirical chapters written while the data work is fresh
Your data section, identification strategy and robustness discussion all follow a predictable structure — and all of them are easier to write the week you build the dataset than three months later. Tesify turns your variable definitions, sample construction and estimation choices into properly structured chapters, keeps every dataset and paper citation formatted, and flags the assumptions you have not yet defended.
Frequently asked questions
Can I write an economics dissertation using only secondary data?
Yes — the overwhelming majority of undergraduate and taught master’s economics dissertations use existing datasets, and departments expect this. Original data collection is rare because sample sizes achievable by one student rarely support econometric inference. Your contribution comes from the question, the identification strategy and the interpretation, not from having gathered the numbers yourself.
How many observations do I need for an econometrics dissertation?
It depends on the design rather than a fixed threshold. A country panel of 40 economies over 25 years gives 1,000 observations but only 40 clusters, and it is the number of clusters that governs your standard errors — fewer than about 30 to 40 clusters makes conventional clustered inference unreliable and calls for wild bootstrap methods. Household microdata typically offers tens of thousands of observations, in which case precision is rarely the binding constraint; identification is.
Should I use Stata, R or Python?
Use whatever your department teaches and supports, since supervisor help matters more than software elegance. Stata remains the standard in applied microeconomics and has the most convenient panel and instrumental-variables commands; R is free, excellent for graphics and increasingly common; Python is strongest when your project involves scraping or large unstructured data. Whichever you choose, submit a do-file or script rather than describing clicks.
What makes an instrumental variable acceptable to an examiner?
Two things demonstrated separately. Relevance is empirical: report the first-stage coefficient and F statistic, with values below roughly 10 signalling a weak instrument. Validity is argued, not tested: you must explain why the instrument could not plausibly affect your outcome through any channel other than the endogenous regressor. Borrowing an instrument from a published paper without re-arguing exclusion in your own setting is the most common weakness in student IV chapters.
Is a descriptive dissertation without causal identification acceptable?
Yes, provided you are explicit about it. A careful descriptive analysis with clear measurement, honest language and thoughtful interpretation typically scores better than a weakly identified causal claim dressed up with an unconvincing instrument. What loses marks is describing correlations in causal language. Use “associated with”, state clearly which confounders remain unaddressed, and explain what design would be required to answer the causal question.
Write your thesis with AI
Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.






Leave a Reply