How to Design a Likert Scale Questionnaire for Your Thesis (2026)
Designing a Likert scale questionnaire is one of the most consequential decisions in quantitative thesis research. Get it right and you collect clean, interpretable data that holds up under examiner scrutiny. Get it wrong — with ambiguous wording, the wrong number of scale points, or skipped piloting — and you spend your viva defending measurement artefacts rather than real findings. This guide walks you through every step of how to design a Likert scale questionnaire that is valid, reliable, and appropriate for academic research in 2026.
The technique dates to Rensis Likert’s 1932 publication A Technique for the Measurement of Attitudes in Archives of Psychology, where he introduced the summated rating approach as a more efficient alternative to the Thurstone scale. Nearly a century later, Likert-type instruments remain the most widely used self-report measurement tool in the social sciences, education, business, and health research — which means examiners have high expectations for how you construct, validate, and report them.
What Is a Likert Scale?
A Likert scale is a psychometric instrument in which respondents indicate their level of agreement, frequency, importance, or satisfaction with a series of statements using an ordered response format. The full questionnaire — consisting of multiple such items measuring the same underlying construct — is correctly called a Likert scale or summated rating scale. A single statement with its response options is, strictly speaking, a Likert-type item.
The critical feature of a Likert scale is that individual item scores are combined (summed or averaged) into a composite score that represents the respondent’s overall position on the construct. This summation is what gives the instrument its statistical power and is why writing items that all genuinely tap the same construct matters so much.
For your thesis, Likert scales are appropriate when you want to measure:
- Attitudes (e.g. towards a policy, intervention, or technology)
- Perceptions and beliefs (e.g. perceived ease of use, perceived competence)
- Behavioural intentions or motivations
- Self-reported experiences or satisfaction
They are not suitable for measuring factual knowledge (use a scored knowledge test), objective behaviour (use observation or log data), or demographic variables (use categorical questions).
Once you have collected and scored your Likert data, you will typically enter the composite scores as predictor or outcome variables in further analysis. For a full walkthrough of that next step, see our guide on how to run multiple regression in SPSS and report it in APA.
Step 1: Define Your Construct
Before writing a single item, write a precise operational definition of the construct you intend to measure. A construct is a latent variable — something you cannot observe directly, such as motivation, job satisfaction, or attitude towards climate change — that you infer from observable indicators (your items).
Your operational definition should answer three questions:
- What does this construct mean? Write two to three sentences grounded in your theoretical framework, citing the scholars whose definition you are adopting. For example: “Academic self-efficacy is defined, following Bandura (1997), as a student’s belief in their capability to complete specific academic tasks at a desired level of performance.”
- What are its dimensions? Many constructs are multidimensional. Self-efficacy, for instance, may have dimensions across different task types. Decide whether you are measuring a unidimensional construct (one composite score) or a multidimensional one (separate subscale scores).
- What does it not include? Boundary conditions prevent construct contamination. Academic self-efficacy is not the same as academic achievement, general self-esteem, or intrinsic motivation.
Document this definition in your methodology chapter before presenting the questionnaire. Examiners will assess whether your items actually reflect this definition — a process called content validity. For a deeper treatment of how construct definition relates to validity threats in research design, see our guide on construct, internal, and external validity.
Step 2: Review Existing Validated Scales
One of the most common mistakes in thesis questionnaire design is starting from scratch when well-validated instruments already exist. Adapting an established scale gives you:
- Prior reliability and validity evidence you can cite
- Comparability with published studies
- Examiner confidence that your measurement approach has scholarly precedent
Search databases such as PsycINFO, ERIC, or PubMed using the construct name plus “scale”, “questionnaire”, or “instrument”. Also check APA PsycTests and the Measurement Instrument Database for the Social Sciences (MIDSS) — both catalogue validated instruments with psychometric properties.
When you find a suitable scale, check:
- Whether its reliability (Cronbach’s alpha) is reported at ≥ 0.70 in studies similar to yours
- Whether it has been validated with a population comparable to your participants
- Whether its items are freely available or require permission to use
If you are adapting rather than adopting a scale wholesale — changing wording, removing subscales, or translating it — you must re-validate it for your context. Minor adaptations (e.g. changing “at work” to “at university”) require at minimum a reliability pilot. Major adaptations require full psychometric testing.
Step 3: Write Your Item Pool
Whether you are creating a new scale or supplementing an existing one, item writing is the most craft-intensive part of questionnaire design. Start by drafting 1.5 to 2 times more items than your target scale length, knowing that some will be removed after piloting.
Rules for Well-Written Likert Items
| Rule | Poor Example | Improved Version |
|---|---|---|
| One idea per item | I feel confident and motivated when writing my thesis. | I feel confident when writing my thesis. (Separate item for motivation) |
| No leading language | Most students find research difficult. Do you agree? | I find carrying out primary research challenging. |
| Avoid universals | I always struggle to find academic sources. | I find it difficult to locate relevant academic sources. |
| Avoid negatives | I do not feel unprepared for my viva. | I feel prepared for my viva examination. |
| Plain language | The epistemological grounding of my methodology is clear to me. | I understand the reasoning behind my chosen research approach. |
| Add a timeframe where relevant | I feel anxious about my research. | Over the past two weeks, I have felt anxious about my research progress. |
The double-barrelled item is the most frequent error in thesis questionnaires. It occurs when a single statement contains two distinct ideas — such as “I find my supervisor supportive and knowledgeable.” A respondent who has a supportive but less knowledgeable supervisor cannot answer honestly. Split every such item before proceeding.

Step 4: Choose Your Scale Format (5 vs 7 Points, Odd vs Even)
This is the decision most thesis students agonise over. Here is a practical framework.
5-Point vs 7-Point
The evidence base, summarised in peer-reviewed survey methodology research including work published in Global Business and Organizational Excellence (Wiley, 2025), supports the following guidance:
- 5-point scales are appropriate for general student populations, online surveys where cognitive load matters, and studies where you do not need to distinguish fine-grained attitudinal differences. They also produce cleaner distributions with less mid-point bunching.
- 7-point scales are better suited to populations with high verbal ability and survey experience, or where your research question requires detecting subtle variations in attitude. They also tend to produce slightly higher test-retest reliability in some contexts.
For most undergraduate and master’s thesis research, a 5-point scale is sufficient and preferable. If your construct is multidimensional and you are conducting doctoral-level psychometric research, a 7-point format may be justified.
Odd vs Even Number of Points
An odd number of points (5, 7) includes a neutral midpoint (“Neither agree nor disagree”). An even number (4, 6) forces respondents to lean one way or the other.
- Include a midpoint (odd scale) when genuine neutrality is a meaningful response — for example, when asking about a topic respondents may have no strong opinion about.
- Remove the midpoint (even scale) when you suspect acquiescence bias, when neutrality is not a theoretically meaningful position, or when you want to increase scale variance.
Most academic surveys use an odd scale. If you force a choice, justify this decision explicitly in your methodology.
Labelling Your Response Options
Always use fully labelled anchors — a word or phrase attached to every point on the scale, not just the endpoints. Research consistently shows that partial labelling introduces ambiguity about what the intermediate points mean, leading to inconsistent interpretations across respondents.
Standard labels for a 5-point agreement scale:
- Strongly disagree
- Disagree
- Neither agree nor disagree
- Agree
- Strongly agree
Use the same label set consistently throughout a subscale or section. Switching between “Strongly agree–Strongly disagree” and “Very satisfied–Very dissatisfied” within the same construct introduces measurement noise.
Step 5: Include Reverse-Coded Items
Reverse-coded items are statements worded in the opposite direction to the construct being measured. If you are building a 10-item scale measuring research confidence, three or four items might be worded negatively: “I doubt my ability to collect valid data” sits alongside “I am confident I can collect reliable data.”
Why this matters:
- It detects acquiescence bias — respondents who agree with everything regardless of content, producing inflated scores.
- It forces respondents to read each item rather than responding on autopilot.
- It is standard psychometric practice; an examiner who sees no reverse-coded items in a 15-item scale will question your instrument’s rigour.
As a guideline, include approximately 20–30% reverse-coded items. Avoid clustering them together — distribute them throughout the questionnaire.
How to Reverse Score During Analysis
Before calculating composite scores, recode reverse-coded item responses using this formula for a 5-point scale:
New Score = (Maximum scale value + 1) − Original Score
For a 5-point scale: New Score = 6 − Original Score
For a 7-point scale: New Score = 8 − Original Score
So a response of 1 (Strongly disagree) on a reverse-coded item becomes 5 (Strongly agree) in the composite, correctly indicating high confidence rather than low. Document every reverse-coded item in your appendix and note them in your data analysis chapter. For a detailed walkthrough in SPSS, see the guide on how to calculate Cronbach’s alpha in SPSS step by step.
Step 6: Arrange Items and Write Instructions
The order and framing of your questionnaire affects response quality more than most researchers expect.
Item Arrangement
- Group items by construct (subscale) rather than mixing them randomly — this reduces cognitive load and produces cleaner factor structures.
- Distribute reverse-coded items randomly within each subscale’s block, not at the end of the section.
- Place your most important or sensitive subscales in the first half of the questionnaire, before response fatigue sets in.
- Begin with neutral demographic items or easy, non-threatening questions to ease respondents into the survey.
Writing the Instructions Block
Your instruction block should appear above the first Likert item and include:
- What the scale measures (without priming a socially desirable response)
- The response format and what each anchor means
- A reminder that there are no right or wrong answers and that responses are confidential
- A prompt to answer every item (to reduce missing data)
Example instruction block: “The following statements are about your experience of writing your thesis. Please indicate how much you agree with each statement by selecting one response. There are no right or wrong answers — please give your honest first reaction. All responses are anonymous.”
Step 7: Run a Cognitive Pilot
A cognitive pilot (also called a think-aloud protocol) involves asking 5 to 10 participants from your target population to complete your draft questionnaire while verbalising their thoughts. You are looking for:
- Items interpreted differently from your intent
- Ambiguous wording or academic jargon that confuses respondents
- Response options that seem inadequate (“I wanted to say something between ‘agree’ and ‘strongly agree’”)
- Items that produce the same response from nearly everyone (floor or ceiling effects that will reduce variance)
- Survey length issues — cognitive pilot data often reveals that a 30-item questionnaire feels like 60
Run the cognitive pilot before collecting any formal data. Record participants’ comments systematically and revise items that generate confusion from two or more participants.
This step is often skipped under time pressure, but it is among the most cost-effective ways to improve data quality. A poorly worded item across 200 responses cannot be fixed retroactively.
Step 8: Collect Pilot Reliability Data
After revising the instrument based on cognitive pilot feedback, administer the questionnaire to a pilot sample of at least 30 participants, with 50 being preferable for stable reliability estimates. The pilot sample should come from the same population as your main study participants — convenience samples of different populations can produce misleading reliability statistics.
Cronbach’s Alpha: The Standard Reliability Measure
Cronbach’s alpha (α) is the most widely reported measure of internal consistency for Likert scales. It ranges from 0 to 1, where higher values indicate that items are more consistently measuring the same construct.
| Alpha Value | Interpretation | Action |
|---|---|---|
| < 0.60 | Unacceptable | Revise or replace items; reconsider construct definition |
| 0.60–0.69 | Questionable | Acceptable for exploratory research; justify in methodology |
| 0.70–0.79 | Acceptable | Meets standard threshold for social science research |
| 0.80–0.89 | Good | Strong reliability; proceed with confidence |
| 0.90–0.95 | Excellent | May indicate item redundancy; consider trimming |
| > 0.95 | Redundancy likely | Remove near-duplicate items to improve parsimony |
The Ordinal Alpha Debate
Cronbach’s alpha technically assumes continuous (interval-level) data. Because Likert responses are ordinal, some methodologists argue that ordinal alpha — which uses a polychoric correlation matrix rather than Pearson correlations — is the more technically correct reliability coefficient for Likert scales. Research published in PLOS ONE and BMC Medical Research Methodology has demonstrated that Cronbach’s alpha can underestimate true reliability for ordinal data, particularly when item distributions are skewed.
In practice, most thesis examiners will accept Cronbach’s alpha as your primary reliability measure. However, if you are conducting doctoral-level psychometric research or using strongly skewed items, calculate ordinal alpha (available in R via the psych package) as a supplementary estimate. State both values and note that they converge (if they do).
Step 9: Refine and Finalise Your Instrument
Analyse your pilot reliability data to identify weak items for removal. The key diagnostic is the corrected item-total correlation — the correlation between each item and the total composite score after removing that item. Items with corrected item-total correlations below 0.30 are not contributing meaningfully to the scale and should be revised or removed.
Also examine the “Alpha if item deleted” column in your SPSS or R output. If removing an item would increase alpha by more than 0.02, consider removing it. However, do not blindly maximise alpha by removing items — always ask whether removing an item would leave a dimension of your construct unrepresented.
After removing or revising items, recalculate alpha on the refined scale. If the revised alpha is acceptable, your instrument is ready for main data collection. Document every change made between the pilot and final versions in your methodology appendix.
Step 10: Analyse and Report Likert Data Correctly
How you analyse and report your data must align with your measurement level — and this is a methodological debate your examiner will likely raise. Here is the current guidance for thesis writers in 2026.
Individual Likert Items (Ordinal Data)
A single Likert item is an ordinal variable. The intervals between response categories are not guaranteed to be equal — the distance between “Disagree” and “Neither agree nor disagree” may not be the same perceived distance as between “Agree” and “Strongly agree.” Treating single items as interval data is a technical error.
For single items:
- Report median and interquartile range (IQR), not mean and SD
- Use frequency tables or bar charts to display distributions
- Use non-parametric tests (Mann-Whitney U, Kruskal-Wallis) for group comparisons
Composite Likert Scale Scores (Interval-Treated Data)
When you sum or average responses across multiple items into a composite score, the resulting variable has more interval-like properties — the law of large numbers and central limit theorem effects mean that composite scores approximate a continuous distribution more closely. This is the mainstream statistical practice in social science research, and treating composite scores as interval data is widely accepted provided your scale has at least four to five items and reasonable distributional properties.
For composite scale scores:
- Report mean and standard deviation
- Use parametric tests (t-test, ANOVA, Pearson correlation, regression) for inferential analysis
- Check normality of composite scores before applying parametric tests
- Explicitly acknowledge in your methodology that you are treating composite scores as continuous, cite supporting methodological literature, and note it as a limitation
For reporting composite scores in your results chapter, see the guide on how to write a thesis results chapter step by step.
Reporting Likert Reliability in Your Methodology
In your methodology chapter, include:
- The name of each scale or subscale and number of items
- The response format (e.g. “5-point Likert scale from 1 = Strongly disagree to 5 = Strongly agree”)
- Which items are reverse-coded (by item number or label)
- Pilot sample size and Cronbach’s alpha from the pilot
- Main sample Cronbach’s alpha (reported again in your results)
- Your position on the ordinal vs interval debate and which statistics you use as a result
Worked Example: Student Academic Self-Efficacy Scale
The following illustrates the full design process for a 10-item scale measuring academic self-efficacy in postgraduate thesis students. This example is for illustration purposes only and draws on established self-efficacy theory (Bandura, 1997) without claiming to be a validated published instrument.
Construct Definition
Academic self-efficacy: a postgraduate student’s belief in their ability to complete the specific academic tasks required to produce an acceptable thesis, including writing, data analysis, and oral examination. Measured as a unidimensional construct yielding a single composite score.
Scale Format
5-point agreement scale (1 = Strongly disagree, 2 = Disagree, 3 = Neither agree nor disagree, 4 = Agree, 5 = Strongly agree).
Items (R = Reverse-coded)
- I am confident I can produce a well-structured thesis.
- I find the writing demands of my thesis manageable.
- I believe I can analyse my data competently.
- I struggle to know whether my arguments are academically sound. (R)
- I feel prepared to defend my research decisions in my viva.
- I am able to identify and use appropriate academic sources.
- I doubt my ability to meet my supervisor’s expectations. (R)
- I can produce clear, accurate academic writing.
- I find it difficult to structure my ideas coherently in writing. (R)
- I believe my research will make a valid academic contribution.
Three of ten items (30%) are reverse-coded (items 4, 7, 9). During analysis, these are recoded using the formula New Score = 6 − Original Score before summing or averaging across all ten items.
Pilot Results
Administered to 42 postgraduate students. Cronbach’s alpha for the full 10-item scale: α = 0.82. Item-total correlations ranged from 0.34 to 0.61, all above the 0.30 threshold. No items raised alpha if deleted. Scale retained in full form for main data collection.
Reporting Example (Results Chapter Sentence)
“The Academic Self-Efficacy Scale demonstrated good internal consistency in the main sample (α = 0.83, n = 187). Composite scores (M = 34.2, SD = 5.6, possible range 10–50) were treated as continuous for parametric analysis, consistent with Norman (2010) and accepted practice for multi-item Likert composites.”
For detailed step-by-step guidance on running this analysis in SPSS, refer to the article on how to calculate Cronbach’s alpha in SPSS. If you collected complementary qualitative data alongside your questionnaire, the methodology for that component is covered in the guide on how to conduct semi-structured interviews for your thesis.
Frequently Asked Questions
Should I use a 5-point or 7-point Likert scale for my thesis?
Both are widely accepted. A 5-point scale is simpler and works well with general populations, especially for online surveys where cognitive load matters. A 7-point scale provides more granularity and is better suited to populations with higher verbal ability or when your research question requires detecting subtle attitudinal differences. Choose one format and use it consistently throughout the questionnaire. For most undergraduate and master’s theses, the 5-point format is sufficient.
How many items do I need for a reliable Likert scale?
A minimum of four to five items per construct is generally recommended for adequate internal consistency. Fewer than four items makes it difficult to achieve a Cronbach’s alpha of 0.70 or above. For exploratory research, six to eight items per construct is a safer target. If your construct is multidimensional, each subscale needs at least four items independently.
What is a good Cronbach’s alpha for a thesis questionnaire?
A Cronbach’s alpha of 0.70 or above is the widely accepted minimum threshold for acceptable internal consistency in social science research. Values between 0.80 and 0.89 indicate good reliability. Values above 0.90 indicate excellent reliability but may also signal item redundancy. Values above 0.95 strongly suggest some items are near-duplicates and the scale could be shortened without losing construct coverage. Report your alpha value alongside the sample size it was calculated from.
Can I treat Likert scale data as interval data in my analysis?
This is debated in the methodology literature. Technically, Likert responses are ordinal — the distances between scale points are not guaranteed to be equal. However, treating composite Likert scale scores (sums or averages of multiple items) as approximately interval data is widely practised and broadly accepted when the scale has at least four to five items and the composite score distribution is approximately normal. Individual Likert items should be analysed as ordinal data using medians and non-parametric tests. Address this distinction explicitly in your methodology chapter to demonstrate methodological awareness.
What is reverse coding and why is it important?
Reverse coding means re-scoring negatively worded items so that high scores consistently indicate high levels of the construct across all items. For a 5-point scale, the reverse score is calculated as: New Score = 6 − Original Score. So a response of 1 becomes 5, and 2 becomes 4. Reverse-coded items help detect acquiescence bias — the tendency of respondents to agree with statements regardless of content. Including approximately 20–30% reverse-coded items in your scale is standard practice and will be expected by your examiner.
How many participants do I need to pilot a Likert scale?
A cognitive pilot (think-aloud interviews to check item clarity and wording) requires 5 to 10 participants from your target population. A quantitative reliability pilot to calculate Cronbach’s alpha requires a minimum of 30 participants, with 50 being a safer target to obtain stable estimates. The pilot sample must come from the same population as your main study — do not use convenience samples from a different group, as this can produce misleading reliability statistics.
Write Your Thesis Methodology Faster with Tesify
Designing a questionnaire is only one part of a methodology chapter. Tesify helps you draft, structure, and refine every section of your thesis — from your research design rationale through to your discussion of findings — while keeping your academic voice and referencing requirements intact. Start your methodology chapter today.
Write your thesis with AI
Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.






Leave a Reply