How to Choose a Validated Psychology Scale for Your Dissertation in 2026: Reliability, Permission and Reporting
You have a research question, an ethics deadline, and a Measures section that is still blank. The single decision that will do most to determine whether your psychology dissertation survives the viva is not your sample size or your statistical test — it is which instrument you use to operationalise your construct. Choose a well-validated scale with a documented factor structure and your analysis chapter almost writes itself. Choose a questionnaire you invented over a weekend, and your marker will spend the whole discussion asking whether you measured anything at all.
This guide walks through the decision in the order you actually have to make it: define the construct, shortlist candidate instruments, check their psychometric credentials, confirm you are allowed to use them, plan the scoring before you collect data, and write the Measures paragraph that examiners look for. Every step is something you can do this week.
Step 1: Define the construct narrowly enough to measure it
Most instrument problems are really definition problems. “Wellbeing” is not a construct you can measure; it is a family of constructs. Hedonic wellbeing (positive affect, life satisfaction) and eudaimonic wellbeing (purpose, autonomy, personal growth) are measured by different instruments and behave differently in regression models. If your research question says “wellbeing,” your supervisor’s first question will be “which kind?”
Write a one-sentence conceptual definition before you look at any scale. Then write the operational definition underneath it: “Perceived stress, defined as the degree to which situations in the past month were appraised as unpredictable, uncontrollable and overloading, measured by the 10-item Perceived Stress Scale.” Notice that the operational definition contains a time frame. Recall window matters enormously: the PHQ-9 asks about the last two weeks, the PSS-10 about the last month, and the Satisfaction With Life Scale about life in general. Two instruments that appear to measure the same thing can be non-interchangeable purely because of their reference period.
A useful discipline: state what your construct is not. Depression is not sadness, resilience is not optimism, burnout is not fatigue. Each of those distinctions maps onto a different instrument.
Step 2: Shortlist instruments that are actually free to use
Undergraduate and taught-master’s dissertations have no measurement budget. Before falling in love with a scale, check whether it is in the public domain, free for non-commercial student research, or commercially licensed at a per-response cost you cannot pay. This single check saves more dissertations than any other step in this guide.
Commonly free or student-accessible instruments
- PHQ-9 (depressive symptoms, 9 items) and GAD-7 (generalised anxiety, 7 items) — no permission required for use, reproduction or distribution; widely used, well-normed, with established cut-offs at 5/10/15/20.
- DASS-21 — three 7-item subscales for depression, anxiety and stress; free for research use; note that its stress subscale often correlates strongly with the anxiety subscale, which matters if you plan to enter both as predictors.
- Perceived Stress Scale (PSS-10) — free for academic research; contains four positively worded items that must be reverse-scored.
- Rosenberg Self-Esteem Scale (RSES) — 10 items, free; famously produces a method factor from its five reverse-worded items, which is why some studies fit a two-factor solution.
- Satisfaction With Life Scale (SWLS) — 5 items, free, unidimensional and extremely robust for a scale that short.
- PANAS / PANAS-X — positive and negative affect, free for research.
- WEMWBS / SWEMWBS — mental wellbeing; free but requires a short online registration and an acknowledgement statement.
- UCLA Loneliness Scale (Version 3) and the De Jong Gierveld 6-item scale — both available for research use.
- MSPSS (perceived social support, 12 items, three sources: family, friends, significant other) — free and cleanly three-factor.
- Big Five Inventory-2 and the IPIP item pools — free personality measures; the IPIP is explicitly public domain.
- FFMQ (mindfulness), ERQ (emotion regulation), AAQ-II (psychological inflexibility), IRI (empathy) — all generally available for non-commercial research.
Instruments that usually cost money
The Beck Depression Inventory-II, the State-Trait Anxiety Inventory, the Maslach Burnout Inventory, the NEO-PI-R and most clinical diagnostic batteries are commercially published. Licences are typically priced per administration and may require a qualified purchaser. If your department already holds a licence, ask — many psychology schools do, and the answer is often yes for supervised student projects. If not, substitute early rather than negotiating late. The Oldenburg Burnout Inventory and the Copenhagen Burnout Inventory are free alternatives to the MBI, and the GAD-7 or DASS-21 anxiety subscale will usually stand in for the STAI.
Step 3: Interrogate the psychometrics before you commit
Do not accept “it is a validated scale” as an answer — including from yourself. Find the original validation paper and at least one recent study in a population resembling yours, and check four things.
Reliability
Cronbach’s alpha above .70 is the conventional floor for group-level research, .80 is comfortable, and above .95 suggests redundant items rather than excellence. Alpha is inflated by scale length, so a 40-item scale with alpha of .88 is less impressive than a 5-item scale with the same figure. Increasingly, reviewers prefer McDonald’s omega, which does not assume every item loads equally on the factor. Reporting both costs you one extra line and looks careful. For any repeated-measures design, you also want test-retest reliability, usually as an intraclass correlation.
Critically: reliability is a property of scores in a sample, not of an instrument. You must compute and report alpha in your data, not merely quote the developer’s figure.
Factor structure and subscale use
If you intend to analyse subscales separately, confirm the subscale structure has been replicated in populations like yours. Many scales developed on North American undergraduates behave differently in clinical, older or non-Western samples. If the literature is contested, say so and justify your choice — that single sentence often earns more credit than a clean but uncritical decision. If you have a large enough sample and the structure is genuinely uncertain, you can test it yourself; our walkthrough of running an exploratory factor analysis in SPSS covers KMO, Bartlett’s test, extraction and rotation in the order SPSS asks for them.
Validity evidence
Look for convergent validity (correlates with related measures), discriminant validity (does not correlate too highly with things it should be distinct from) and, ideally, criterion or predictive validity. Where a diagnostic cut-off exists, note its sensitivity and specificity in the population you are sampling — a screening threshold validated in primary care does not automatically transfer to university students.
Population fit and reading level
Check the age range, language and clinical status of the validation sample. If your participants are adolescents, an adult-normed instrument may need the youth version (PHQ-A rather than PHQ-9, for example). If you are collecting data in another language, do not translate the scale yourself in an afternoon — use an existing validated translation, and if none exists, follow a formal procedure. We have a full walkthrough of back-translation and cross-cultural scale adaptation, including what to report when you have adapted rather than adopted.
Step 4: Get permission in writing before ethics submission
Even free scales frequently carry conditions: registration, a citation requirement, a prohibition on modification, or a restriction to non-commercial use. Ethics committees increasingly ask for evidence of permission, and a screenshot of a licence page or an email from the author is far easier to obtain in week two than in week ten. The practical mechanics — who to email, what to say, what counts as evidence, and what to do when the author never replies — are set out in our guide to getting permission to use a validated questionnaire.
One rule with no exceptions: do not silently alter items. Dropping items, changing the response format from 7 points to 5, or rewording for your population all break comparability with published norms and invalidate the developer’s reliability evidence. If you must adapt, document every change, report it as an adaptation, and re-examine the internal structure in your own data.
Step 5: Plan scoring and reverse-keying before data collection
Write your scoring syntax the day you finalise the questionnaire, not the day after data collection closes. For each scale, record in a single document: the response anchors and their numeric values, which items are reverse-scored, whether the total is a sum or a mean, how many missing items are tolerated before a case is dropped, and the direction of interpretation.
Reverse-keyed items cause more analysis errors in psychology dissertations than any other single mechanic. On a 1-to-5 scale the transformation is 6 − x; on a 0-to-4 scale it is 4 − x. Get this wrong and your alpha collapses to near zero or even goes negative — which, usefully, is the standard diagnostic. If alpha is implausibly low, check reverse-scoring first. Keeping a versioned record of these decisions is exactly what a data dictionary and cleaning log is for, and examiners notice when you have one.
Decide your missing-data rule in advance too. A common convention is prorating a subscale score when no more than 10-20% of its items are missing and excluding the case otherwise. State the rule; do not improvise it.
Step 6: Think about how the scale will behave in your analysis
Your choice of instrument constrains your statistics. Three consequences worth anticipating:
- Distribution. Clinical symptom scales in non-clinical samples are usually floor-heavy and positively skewed. A PHQ-9 total in a healthy undergraduate sample will pile up near zero, which pushes you toward robust or non-parametric approaches. Our overview of Mann-Whitney, Wilcoxon and Kruskal-Wallis covers when that switch is justified and how to report it.
- Sum scores versus latent variables. If your model has several latent constructs and mediation paths, structural equation modelling uses the items rather than a single total and handles measurement error explicitly — but it demands a bigger sample and a defensible factor structure.
- Common method variance. If every variable in your model comes from one self-report questionnaire completed at one sitting, some of your correlation is attributable to the method rather than the constructs. Designing around this — temporal separation, different response formats, a marker variable — is much stronger than apologising in your limitations. See our treatment of common method bias in survey research for design remedies that a marker will accept.
If you are still deciding which analysis your design implies, work backwards from your variables using our decision guide to choosing a statistical test before you finalise the questionnaire. It is far cheaper to add a demographic item now than to discover in April that you cannot run the model you promised in your proposal.
Step 7: Write the Measures section examiners want
Every instrument in your Measures subsection should be described in a compact paragraph containing eight elements, in roughly this order:
- Full name of the scale, abbreviation and citation of the original development paper.
- What it measures, in one clause.
- Number of items and any subscales, with item counts.
- Response format and anchors, verbatim.
- A sample item in quotation marks.
- Scoring: sum or mean, score range, reverse-scored items, direction of higher scores, and any cut-offs you apply.
- Previously reported reliability, with the population it came from.
- Reliability in the present sample.
A worked example: “Perceived stress was measured with the 10-item Perceived Stress Scale (PSS-10; Cohen & Williamson, 1988), which assesses the extent to which respondents appraised situations in the past month as unpredictable, uncontrollable and overloading. Items (e.g. ‘In the last month, how often have you felt that you were unable to control the important things in your life?’) are rated from 0 (never) to 4 (very often). Four positively worded items are reverse-scored and all items are summed, giving a range of 0-40, with higher scores indicating greater perceived stress. Internal consistency has been reported between .78 and .91 in community samples; in the present sample, Cronbach’s alpha was .87 and McDonald’s omega was .88.”
That paragraph is replicable, citable and complete. Write one for every instrument, including single-item demographic measures where the wording could matter. If your wider chapter architecture still needs shaping, our complete guide to writing a psychology dissertation covers how the Measures subsection sits inside the Method chapter alongside design, participants and procedure.
Common mistakes that cost marks
- Quoting the developer’s alpha and never computing your own. This is the most frequently flagged omission in psychology dissertation marking.
- Using subscales the validation evidence does not support. Splitting a unidimensional scale into “factors” you found in a small sample is not a contribution, it is noise.
- Stacking six long questionnaires into one survey. Attrition rises sharply past about 10 minutes, and careless responding rises with it. Three well-chosen instruments beat six.
- Treating a screening cut-off as a diagnosis. Say “scores above the clinical threshold,” never “participants with depression.”
- Forgetting that a scale needs a citation in the reference list as well as the Measures paragraph.
Draft your Measures section while the details are still in front of you
The Measures subsection is one of the few parts of a psychology dissertation that can be written completely before you collect a single response — and it is far easier to write while the scale documentation is open on your screen. Tesify turns your instrument details, response anchors and reliability figures into properly structured Method prose in your own academic voice, keeps every citation formatted in APA 7, and flags the elements you have left out.
A one-week action plan
- Day 1. Write conceptual and operational definitions for every construct in your model, including recall windows.
- Day 2. Shortlist two or three candidate instruments per construct and record their licence status.
- Day 3. Read each original validation paper plus one recent study in your target population. Note alpha, factor structure and sample characteristics.
- Day 4. Request permission in writing where required. Save the responses in your ethics folder.
- Day 5. Build the questionnaire, estimate completion time, and pilot it with five people.
- Day 6. Write the scoring key and analysis syntax, including reverse-keying and missing-data rules.
- Day 7. Draft the Measures section in full, leaving only the present-sample reliability figures blank.
Frequently asked questions
Can I create my own questionnaire instead of using a validated scale?
You can, but you should only do so when no suitable instrument exists, and you must then present evidence of reliability and validity yourself — typically factor analysis and internal consistency in a pilot sample. For a single dissertation cycle this is a large extra workload, and markers will scrutinise it heavily. For measuring an established psychological construct, an existing validated scale is almost always the stronger choice. Bespoke items are appropriate for context-specific facts such as course of study, hours of placement, or frequency of a particular behaviour.
What Cronbach’s alpha is acceptable for a dissertation?
Above .70 is the conventional threshold for group-level research, .80 or higher is comfortable, and values above .95 may indicate redundant items. Short subscales of three or four items often produce alphas in the .60s despite being acceptable, because alpha rises with scale length; in that case report the mean inter-item correlation as well, where .15 to .50 is generally considered adequate. Always report reliability computed in your own sample, not just the figure published by the scale’s developer.
Do I need permission to use the PHQ-9 or GAD-7 in my dissertation?
No. Both instruments were placed in the public domain by their developers and may be reproduced, translated and used without permission or fee, provided they are cited correctly. Many other widely used scales — including the Satisfaction With Life Scale and the IPIP personality item pools — are similarly free. Others, such as the Beck Depression Inventory-II, the State-Trait Anxiety Inventory and the Maslach Burnout Inventory, are commercially licensed and require a paid purchase. Always verify the current licence status yourself and keep a copy of the terms in your ethics file.
Can I shorten a validated scale to save my participants time?
Not arbitrarily. Removing items breaks the correspondence with published norms, reliability estimates and cut-offs, and many licences explicitly prohibit modification. If you need a shorter instrument, look for an officially validated short form — such as the SWEMWBS instead of the WEMWBS, the PSS-4 instead of the PSS-10, or the BFI-2-S instead of the full BFI-2 — and cite the short-form validation paper directly.
How many scales can I reasonably include in one dissertation survey?
Keep total completion time under roughly 10 to 15 minutes, which for most undergraduate samples means three to five instruments alongside your demographic items. Longer surveys produce higher dropout rates and more careless responding, both of which damage your data more than a missing variable would. Include only measures that appear in your hypotheses or serve as pre-registered covariates.
What should I do if my scale’s reliability is poor in my sample?
First check reverse-scoring, which explains the majority of unexpectedly low alphas. Then inspect the corrected item-total correlations for items below about .30, and consider whether the item is ambiguous in your population. Do not simply delete items until alpha looks acceptable and then report the result as though nothing happened. If reliability remains poor, report it honestly, interpret the associated findings cautiously, and discuss it as a measurement limitation — examiners reward transparency far more than a suspiciously tidy table.
Write your thesis with AI
Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.






Leave a Reply