Which Scales Should You Use in a Marketing Dissertation? Brand, Loyalty and Purchase Intention Measures for 2026
Use established multi-item scales from published marketing journals rather than writing your own items. Nearly every construct a marketing dissertation needs — brand equity, brand trust, perceived value, service quality, satisfaction, loyalty, purchase intention, electronic word of mouth, technology adoption — already has a validated instrument with reported reliability, and adapting one takes an afternoon instead of a semester.
What follows is the working shortlist, organised by the question your dissertation is asking, plus the adaptation and reporting rules that decide whether a marker treats your measurement section as competent.
Why can’t I just write my own questionnaire items?
You can write single-item factual questions — how often you buy the category, which platform you use, your age bracket. You should not invent multi-item measures of psychological constructs. An established scale arrives with published evidence of internal consistency, a known factor structure, discriminant validity against neighbouring constructs, and a body of prior studies your findings can be compared against. Inventing items throws all of that away and hands your examiner an obvious line of attack: how do you know these seven statements measure brand love rather than general positivity?
There is also a practical argument. If you use a scale that appears in 200 published papers, you can benchmark your Cronbach’s alpha, anticipate the factor structure, and cite precedent for every methodological decision you make. That is a large amount of defensive armour for very little effort.
Which scales measure brand constructs?
Brand equity
The most-used operationalisation in student research is Yoo and Donthu’s multidimensional consumer-based brand equity scale, which separates brand awareness/associations, perceived quality and brand loyalty, alongside a four-item overall brand equity measure. Aaker’s conceptual framework underpins it, but Aaker’s model is a theory rather than a ready-made questionnaire, so cite Aaker for the concept and Yoo and Donthu for the items.
Brand trust and brand attachment
Chaudhuri and Holbrook’s brand trust and brand affect scales are short, well-replicated and pair naturally with loyalty outcomes. For deeper emotional bonds, Batra, Ahuvia and Bagozzi’s brand love scale and Thomson, MacInnis and Park’s emotional attachment scale are both established, though brand love in particular has several competing versions of differing length — pick one, cite the specific paper, and note the version in your methods.
Brand image and personality
Aaker’s Brand Personality Scale (sincerity, excitement, competence, sophistication, ruggedness) is the standard, but be aware that its dimensional structure has been contested across cultures. If you are running it outside a US context, either cite the replication evidence for your setting or treat the dimensional analysis cautiously.
Which scales measure service quality and satisfaction?
SERVQUAL, with its five dimensions of tangibles, reliability, responsiveness, assurance and empathy, remains the reference point in services research. Its complication is the gap-score design: SERVQUAL asks for expectations and perceptions separately and analyses the difference. Difference scores are psychometrically awkward and reduce reliability, which is precisely why SERVPERF — the performance-only version — is common in student work and often performs better. Choose one, and explain the choice in a sentence; that sentence alone demonstrates you understand the debate.
For satisfaction, use a short established multi-item measure (Oliver’s expectancy-disconfirmation items or the American Customer Satisfaction Index items) rather than a single “how satisfied are you” question, since single items cannot have their reliability assessed. Net Promoter Score can appear as a practitioner benchmark but should not be your only dependent variable — it is a single item recoded into categories, and academic reviewers treat it as commercial metric rather than a validated scale.
Which scales measure purchase intention and value?
Purchase intention is most often measured with the three-item Dodds, Monroe and Grewal scale (“the likelihood of purchasing this product is high”, and variants), typically on a seven-point semantic differential or Likert format. Willingness to pay premium prices has established items from Netemeyer and colleagues’ brand equity work.
For perceived value, the PERVAL scale by Sweeney and Soutar breaks value into emotional, social, quality/performance and price/value-for-money dimensions, which is far more informative in a dissertation than a single global value item. Zaichkowsky’s Personal Involvement Inventory measures product-category involvement, a common moderator.
Which scales measure digital, social and influencer constructs?
- Electronic word of mouth — eWOM intention and eWOM credibility items adapted from Goyette and colleagues’ eWOM scale or from information-adoption research.
- Source credibility of an influencer — Ohanian’s three-dimensional scale covering attractiveness, trustworthiness and expertise is the standard, and still the most defensible choice for influencer marketing dissertations.
- Parasocial relationship — established parasocial interaction items adapted from media research, increasingly common in creator-economy projects.
- Consumer engagement — Hollebeek’s cognitive, emotional and behavioural engagement dimensions, or Vivek’s customer engagement scale.
- Privacy concern — Malhotra’s Internet Users’ Information Privacy Concerns instrument, useful for anything involving personalisation or data use.
If your dissertation is about adoption of a new app, payment method or retail technology, you are in technology acceptance territory rather than classic brand research, and the choice between the parsimonious TAM and the richer UTAUT determines both your model and your sample requirements. Our comparison of TAM versus UTAUT for a dissertation sets out which one your research question actually needs.
Which scales measure sustainability and ethical consumption?
Green marketing dissertations are among the most common topics in 2026, and the measurement here is more fragmented. Green purchase intention items are usually adapted from Chan’s green purchase behaviour work; environmental concern from the New Ecological Paradigm scale; greenwashing perception from Chen and Chang’s green scepticism items; and the attitude-behaviour gap is typically operationalised by measuring stated intention and reported behaviour separately rather than by a dedicated instrument. Consumer ethnocentrism has the long-established CETSCALE, with a validated 10-item short form that is usually the better choice for a student survey.
If your model is behavioural rather than attitudinal — you want to explain why intention does not convert into action — a behaviour-change framework may serve you better than a stack of attitude scales. Our guide to COM-B and the Behaviour Change Wheel in a dissertation shows how to use one as an organising model.
How do I adapt a scale to my specific brand or category?
Substituting the brand or category name into the item stem is standard, expected practice: “I trust [brand]” is a legitimate adaptation of “I trust this brand.” Note the substitution in your methods and cite the original.
What is not acceptable is quietly rewriting item content, cutting items to shorten the survey, or changing the response format without saying so. Each of those breaks comparability with the published reliability and factor structure. If you must adapt substantively — say, translating for a non-English sample — follow a documented procedure and report it; our walkthrough of back-translation and cross-cultural adaptation covers the standard steps and what to state in the methods.
Keep the response format consistent across your questionnaire where possible. Mixing five-point and seven-point Likert scales within a single instrument confuses respondents and complicates comparison of standardised coefficients. Seven-point Likert is conventional in marketing research; five-point is acceptable if applied consistently.
How many items and constructs should my survey contain?
Three items is the practical minimum per latent construct for structural equation modelling; four gives you room to drop a poorly loading item without dropping below identification. Beyond that, extra items buy diminishing returns.
Constrain the total. Marketing dissertation surveys routinely balloon to 80 items across nine constructs, and the resulting dropout and careless responding cost more data than the extra construct gains. Six constructs at four items each, plus demographics, is a serious study that respondents will finish. Screening the responses you do collect is a separate skill — our guide to screening junk survey responses, bots and careless answering covers attention checks, straight-lining detection and completion-time thresholds, all of which you should build in before you launch.
What sample size do the analyses need?
If you are running regression with five or six predictors, 150 to 200 usable responses is a workable target. If you are running covariance-based structural equation modelling with six latent constructs, you want 250 to 400, and the common heuristic is at least ten responses per estimated parameter. Partial least squares SEM tolerates smaller samples, which is why it is so widespread in marketing dissertations — but “PLS works with small samples” is not itself a justification for choosing it. Our comparison of CB-SEM versus PLS-SEM sets out the theory-testing versus prediction distinction that should drive the decision, and if you have already chosen PLS, our walkthrough of running PLS-SEM in SmartPLS covers the measurement model, bootstrapping and the exact statistics to report.
What do I have to report about my measures?
For each construct, your measurement section should state the source paper, the number of items, the response anchors, one sample item, and any adaptation. Then, in your results, report the psychometric evidence:
- Internal consistency — Cronbach’s alpha and, increasingly expected, composite reliability, both above .70.
- Convergent validity — average variance extracted above .50, with standardised item loadings above .70 (loadings between .40 and .70 may be retained if AVE and reliability thresholds are still met).
- Discriminant validity — the Fornell-Larcker criterion, and in current practice the HTMT ratio, which should fall below .85 or .90 depending on how similar the constructs are conceptually.
- Common method bias — if all your constructs come from one self-report questionnaire at one time point, address this. Harman’s single-factor test is widely used and widely criticised as insufficient; our explainer on common method bias in survey research covers procedural remedies that carry more weight than a post-hoc test.
If you plan to examine the factor structure in your own data before modelling, our step-by-step guide to exploratory factor analysis in SPSS covers KMO, Bartlett’s test, extraction and rotation choices.
Write the measurement section while your scale sources are open
The measurement subsection of a marketing dissertation is dense, formulaic and entirely writable before your survey closes. Tesify turns your construct list, item sources and response formats into a properly structured methodology section, keeps every scale citation formatted correctly, and flags the reliability and validity statistics you still need to report.
Frequently asked questions
Do I need permission to use a marketing scale published in a journal article?
Scales published in the body of a journal article are conventionally used in academic research with citation rather than formal permission, and most marketing instruments fall into this category. Commercially owned instruments are the exception and require licensing. Where the items are not printed in full in the paper, email the corresponding author — they usually reply, and a short email is worth keeping in your ethics folder as evidence of good practice.
Should I use SERVQUAL or SERVPERF?
SERVPERF is generally the better choice for a student dissertation. It measures perceptions of performance only, is half the length, avoids the reliability problems associated with difference scores, and has repeatedly matched or outperformed SERVQUAL in predicting satisfaction. Use SERVQUAL when your research question is specifically about the gap between expectations and experience, since that gap is the whole point of the instrument.
Can I use Net Promoter Score as a loyalty measure?
Include it as a practitioner benchmark if it is useful to your context, but do not make it your only loyalty measure. NPS is a single item collapsed into categories, so its reliability cannot be assessed and its treatment of the 0-to-10 scale discards information. Pair it with an established multi-item loyalty scale — Yoo and Donthu’s loyalty items or Zeithaml’s behavioural intentions battery — and use that for your inferential analysis.
How many responses do I need for a marketing dissertation survey?
For multiple regression with a handful of predictors, aim for 150 to 200 usable responses. For covariance-based structural equation modelling with several latent constructs, aim for 250 to 400. Partial least squares SEM can run on smaller samples, often 100 to 150 depending on model complexity, but sample convenience should not be your stated reason for choosing it. Whatever the target, plan for 20 to 30% of raw responses to be removed during data screening.
Is a five-point or seven-point Likert scale better?
Seven points is the marketing convention and gives slightly more variance and better discrimination, which helps in structural models. Five points is easier on mobile screens and is perfectly defensible. What matters most is consistency: use the same format across constructs where the original scales allow it, reproduce the original anchors when you can, and state the format explicitly in your measurement section.
What is HTMT and why is my supervisor asking for it?
The heterotrait-monotrait ratio of correlations is a test of discriminant validity — evidence that two constructs in your model are genuinely distinct rather than two labels for the same thing. It has become the expected standard because the older Fornell-Larcker criterion is now known to miss discriminant validity problems in many realistic conditions. Values below .85 indicate clear distinction; below .90 is acceptable for conceptually similar constructs such as satisfaction and loyalty.
Write your thesis with AI
Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.






Leave a Reply