How Do You Establish Validity and Reliability of a Survey in a Nursing Thesis? Content Validity Index, Pilot Test and Alpha (2026)

Tesify Team Avatar

·

Direct answer: To establish validity and reliability of a survey in a nursing thesis, you show three things in your methods chapter. Content validity: a panel of experts rates each item for relevance and you report the content validity index. Pilot evidence: you test the draft on a small group from your population and revise unclear items. Reliability: you report an internal consistency coefficient such as Cronbach’s alpha, or KR-20 for yes/no items, computed in your own sample, and test-retest reliability if the construct should be stable over time.

If you adopt a published instrument rather than build one, you still report the original validation and your own reliability coefficient. The general concepts are explained in the guides to construct, internal and external validity and Cronbach’s alpha; this article shows how a nursing thesis applies them, with a worked illustration.

What is the difference between validity and reliability in a nursing survey?

Validity asks whether the survey measures what you say it measures, for example nurses’ attitudes toward a patient safety practice rather than their general job mood. Reliability asks whether it measures consistently, so that the same nurse in the same state would give similar answers on repeated administration, and so that the items meant to measure one concept hang together. A survey can be reliable without being valid, but it cannot be valid if it is unreliable, which is why committees expect both and expect them reported separately.

How do you establish content validity with an expert panel?

Content validity means the items adequately cover the concept. In nursing research the usual route follows the two-stage process described by Lynn (1986, Nursing Research, 35(6), 382-386): a development stage, in which you define the concept and write items from the literature and your framework, and a judgment-quantification stage, in which experts rate the items and you compute an index.

  1. Define the construct and its domains. Write a one-paragraph definition and list the domains the items should cover, such as knowledge, attitude and practice.
  2. Recruit the panel. Choose clinicians, educators or researchers with direct expertise in the topic. Document who they are, their credentials and years of experience, because examiners ask.
  3. Collect ratings. Ask each expert to rate each item for relevance on a four-point scale, for example 1 = not relevant, 2 = somewhat relevant, 3 = quite relevant, 4 = highly relevant, and to comment on clarity and wording.
  4. Compute the index. The item-level content validity index (I-CVI) is the number of experts rating the item 3 or 4 divided by the number of experts. The scale-level index averaging the I-CVIs (S-CVI/Ave) summarizes the whole instrument.
  5. Decide what to revise. Polit, Beck and Owen (2007, Research in Nursing and Health, 30(4), 459-467) report that, after adjusting for chance agreement, items with an I-CVI of .78 or higher with three or more experts could be considered evidence of good content validity. Revise or drop items below your stated threshold, and report both what you changed and why.

What does a worked CVI look like?

The numbers below are illustrative, not from a real study. Imagine a ten-item draft scale on handover communication rated by seven experts.

Item Experts rating 3 or 4 (of 7) I-CVI Decision
1 7 1.00 Keep
2 7 1.00 Keep
3 6 .86 Keep
4 6 .86 Keep
5 7 1.00 Keep
6 5 .71 Revise wording and re-rate
7 7 1.00 Keep
8 6 .86 Keep
9 7 1.00 Keep
10 6 .86 Keep

The I-CVI values add up to about 9.14, so S-CVI/Ave = 9.14 / 10 = .91. Five of the ten items were rated relevant by every expert, so the universal agreement version (S-CVI/UA) is 5 / 10 = .50. Report both and explain the difference: the average version describes typical item quality, while universal agreement is much stricter and falls quickly as the panel grows. Item 6 sits below .78, so you would rewrite it using the experts’ comments and either re-rate it or drop it. Your methods chapter should state the scale, the panel size, the threshold, the indexes and the changes made.

Panel of experts rating survey items on a grid with agreement marks
Document the panel, the rating scale, the threshold and every change you made.

What should a pilot test show?

A pilot is a small trial of the survey on respondents like your target population, whose data you do not include in the final analysis. Its job is practical: time to complete, confusing wording, items people skip, technical problems with the online form, and a first look at reliability. Report the number of pilot participants, how they were recruited, what problems appeared and what you changed. There is no single required pilot size; many nursing theses use a few dozen participants, but your committee’s expectation and your population size decide. The step-by-step guide to conducting a pilot study covers the procedure, and a pilot on a small unit may also reveal that your intended sample size is unrealistic.

Where do face validity and construct validity fit?

Content validity is the evidence most nursing committees ask for first, but two neighbours often come up in the viva. Face validity is the informal impression that the survey looks like it measures the concept; it is worth checking with a few members of your target group during the pilot, by asking whether any item seems off-topic or confusing, but it is weak evidence on its own and should never replace the expert panel. Construct validity asks whether scores behave the way theory predicts. With a large enough sample you can support it with exploratory or confirmatory factor analysis, reporting the sampling adequacy statistics and factor loadings you used, or with correlations between your scale and a related, already validated measure. If your sample is too small for factor analysis, say so as a limitation rather than skipping the question; examiners respect an honest account of what the design can and cannot show.

How do you report reliability?

Choose the coefficient that matches your items and say why.

  • Cronbach’s alpha for Likert-type scales: report it for each subscale and for the total, with the number of items. Values of .70 or above are commonly cited as acceptable for research use, but interpret alpha together with the number of items and the breadth of the construct.
  • KR-20 for items scored right or wrong or yes or no, such as a knowledge test.
  • Test-retest reliability when the construct should be stable over a short interval: administer the survey twice to the same people, typically one to two weeks apart, and report an intraclass correlation or a correlation coefficient with the interval stated.
  • Inter-rater agreement when observers score practice, for example with Cohen’s kappa.

An illustrative reporting sentence: “Internal consistency was acceptable for the 12-item attitude subscale (Cronbach’s alpha = .86, n = 30 in the pilot and .84, n = 212 in the main sample).” Always recompute reliability in your final sample; the coefficient from a published validation does not transfer automatically, especially if you translated or adapted the wording. For adapted or translated tools, the guide to back-translation and adapting a validated scale explains the extra steps.

What if I use an existing instrument?

You still do four things. Cite the original validation study and the reliability it reported. Obtain and document permission, as described in the guide to getting permission to use a validated questionnaire. State the version, number of items and scoring. And report your own reliability coefficient. If you shorten the instrument or change wording, say so and treat it as a new version that needs its own content validity evidence and reliability.

What mistakes do nursing committees flag most?

  • Reporting alpha from a published study instead of your own sample.
  • Calling a draft “validated” because a supervisor read it, without an expert panel or an index.
  • Listing the expert panel without credentials or the rating scale.
  • Using Cronbach’s alpha on yes/no items, where KR-20 is the matching coefficient.
  • Including pilot responses in the final dataset.
  • Reporting a high alpha as proof of validity.
  • Forgetting to reverse-score negatively worded items before computing reliability.

How do you write the methods paragraph?

A compact structure that satisfies most committees: one paragraph on the instrument and its source, one on content validity (panel, rating scale, index, threshold, revisions), one on the pilot (size, findings, changes), and one on reliability (coefficient, values, interpretation, sample). Put the expert-rating table and the final item list in an appendix, and refer to them from the text. For how this section sits inside the wider chapter, see the guide to the methodology chapter of a nursing dissertation or DNP project.

An illustrative version of the content-validity paragraph reads: “Seven nurse educators and clinicians with at least five years of experience rated each of the ten draft items for relevance on a four-point scale. Item-level content validity indexes ranged from .71 to 1.00, and the scale-level index (average) was .91. One item with an index below .78 was reworded following the experts’ comments and re-rated.” Notice that it names the panel, the scale, the range, the scale-level result and the action taken, in four sentences.

How long should you plan for validation?

Validation takes longer than most students expect because it depends on other people. Allow time for expert invitations and reminders, for a second rating round if several items fall below your threshold, for the review board to approve the pilot, and for the pilot itself. Build these steps into your proposal timeline, and start recruiting experts while you are still writing the literature review so the draft instrument is ready for their ratings as soon as it is written.

Small pilot group completing a survey beside a reliability gauge on a laptop
Report reliability from your own sample, not only from the original validation.

Where Tesify fits

After you have your expert ratings and reliability results, Tesify’s thesis workspace helps you structure the instrumentation and validity sections of your methods chapter; 9,000+ students have written 15,000+ chapters with Tesify. Every word stays 100% written by you, and Tesify cannot recruit your expert panel or run your reliability analysis.

Frequently asked questions

How many experts do I need for content validity?

Guidance varies. The Polit, Beck and Owen guidance quoted above applies to panels of three or more experts, and many nursing theses use between three and ten; state your number and justify it.

Is a content validity index enough to call my survey valid?

No. It is evidence of content validity only. Many theses add pilot testing, reliability and, where the sample allows, factor analysis to support construct validity.

What alpha value is good enough?

A value of .70 or above is commonly cited as acceptable for research, but interpret it together with the number of items and the purpose; very high values can indicate redundant items.

Do I need to pilot a published instrument?

If you use it unchanged in the population it was designed for, a pilot is optional but useful. If you adapt wording, translate it or use it with a new population, a pilot is strongly advisable.

Can I include pilot respondents in my final sample?

Generally no, because the pilot may lead you to change the survey, and mixing versions would blur your data. State clearly that pilot responses were excluded.

What is the difference between S-CVI/Ave and S-CVI/UA?

S-CVI/Ave is the average of the item-level indexes; S-CVI/UA is the proportion of items that every expert rated relevant. The universal agreement version is much stricter, so report both and explain the difference.

Which reliability coefficient fits a yes/no knowledge test?

KR-20 is designed for dichotomously scored items, while Cronbach’s alpha is its generalization for scales with more than two response options.

Do I need ethics approval for the pilot and the expert panel?

Often the pilot needs approval because it involves participants, and expert review may need a brief notification or exemption. Your institution’s review board makes that decision, so ask before you begin.

Write your thesis with AI

Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.

Tesify Team Avatar

Leave a Reply

Your email address will not be published. Required fields are marked *