Construct, Internal, and External Validity in Research Design Explained (2026)

·

Construct, Internal, and External Validity in Research Design Explained (2026)

Validity is the cornerstone of rigorous research. Without it, even the most elegant experimental design, the most sophisticated statistical model, and the most carefully collected dataset amount to little more than noise dressed as signal. Yet for many researchers — from early-stage PhD candidates to experienced faculty — the four principal types of validity in the Cook and Campbell tradition remain muddled: conflated with one another, confused with reliability, or acknowledged only in a perfunctory paragraph of the limitations section. This guide offers a systematic account of construct, internal, external, and statistical conclusion validity, the major threats to each, and the design and analytic strategies that make your causal inferences credible and your findings worth citing.

The canonical framework derives from a lineage of foundational texts: Campbell and Stanley (1963), Cook and Campbell (1979), and — most comprehensively — Shadish, Cook, and Campbell (2002), Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Understanding this typology is not an abstract methodological exercise. Grant reviewers, journal editors, and thesis examiners routinely use it to evaluate the epistemic quality of a study. Researchers who can name the relevant threats to their specific design — and articulate how they have addressed them — produce stronger proposals, stronger papers, and stronger theses.

Quick Answer

The Shadish–Cook–Campbell framework defines four validity types for causal research: statistical conclusion validity (is there a real covariation?), internal validity (does that covariation reflect a causal relationship?), construct validity (do your operationalisations accurately represent the intended constructs?), and external validity (does the causal relationship generalise beyond the study?). Each type has distinct threats — and each threat demands a distinct remedy.

Why Four Validity Types?

The plurality of validity types reflects the plurality of ways a causal inference can fail. Campbell and Stanley (1963) originally distinguished only two forms — internal and external — asking, respectively: “Is there a causal relationship here?” and “Does that relationship generalise?” Cook and Campbell (1979) added two further forms: construct validity (whether the constructs involved in the causal claim are accurately operationalised) and statistical conclusion validity (whether there is defensible statistical evidence for covariation at all). Shadish, Cook, and Campbell (2002) refined the full taxonomy, cataloguing dozens of specific threats across the four types.

The key insight is that the four types address different levels of a causal question. Working roughly from the data upwards:

  • Statistical conclusion validity concerns the numbers: Is there a statistically reliable relationship between the variables as measured?
  • Internal validity concerns the mechanism: If that relationship is real, is it genuinely causal — or confounded?
  • Construct validity concerns the interpretation: Do the variables actually represent the theoretical constructs the researcher intended?
  • External validity concerns the scope: Under what conditions, in what populations, and at what times does the causal relationship hold?

Addressing these questions in order makes logical sense: there is little point assessing generalisation (external validity) if the underlying causal inference (internal validity) is not secured, or if the measures used do not correspond to the constructs of interest (construct validity).

Statistical Conclusion Validity

Statistical conclusion validity is the validity of inferences about whether a covariation between the independent variable (IV) and dependent variable (DV) exists, based on statistical evidence. It is the most basic form of validity, and it is threatened whenever the analysis yields a misleading picture of whether an effect is present at all.

Threats to Statistical Conclusion Validity

Shadish, Cook, and Campbell (2002) enumerate nine threats. The most consequential for most research contexts are:

Threat Description Direction of Error
Low statistical power Sample too small to detect a real effect of the expected magnitude Type II error (false negative)
Violated test assumptions Using a parametric test when data are severely non-normal or heteroscedastic Either direction
Multiple comparisons (fishing) Running many tests without error-rate correction inflates the family-wise Type I error rate Type I error (false positive)
Unreliable measures Measurement error attenuates observed effect sizes Type II error
Restriction of range Sampling only a narrow band of the IV or DV artificially reduces correlations Type II error
Heterogeneity of units High within-group variance reduces power; averaging over subgroups may mask effects in both directions Either direction

Strengthening Statistical Conclusion Validity

  • A priori power analysis: Use G*Power, R’s pwr package, or an equivalent tool to calculate minimum sample size before data collection. Specify your expected effect size, acceptable Type I error (α = .05), and desired power (1 − β ≥ .80).
  • Pre-registration: Register your hypotheses, outcome variables, and analysis plan on the Open Science Framework or AsPredicted before data collection. This distinguishes confirmatory from exploratory analyses and prevents post-hoc hypothesis adjustment.
  • Correction for multiple comparisons: Apply Bonferroni, Holm, or Benjamini–Hochberg corrections when running multiple significance tests on correlated outcomes.
  • Report effect sizes and confidence intervals: p-values alone are insufficient. Cohen’s d, η², or r with 95% CIs communicate the magnitude and precision of effects.
  • Check assumptions: Test normality (Shapiro–Wilk for n < 50), homogeneity of variance (Levene’s test), and independence before applying parametric tests.

Internal Validity

Internal validity is the degree to which a study supports the inference that variation in the independent variable caused variation in the dependent variable, within the specific conditions of the study. A study has high internal validity when all plausible alternative explanations for the observed relationship have been ruled out. This is the primary concern of experimental and quasi-experimental research.

The Major Threats to Internal Validity

Campbell and Stanley (1963) initially catalogued eight threats; Shadish, Cook, and Campbell (2002) extended and refined the list. The core threats that researchers encounter most often are:

History

Events occurring between pre-test and post-test measurements, beyond the experimental treatment, that could account for observed change. In a longitudinal study of student anxiety across an academic year, a widely publicised mental health crisis occurring during the data collection window would constitute a history threat. Longer studies are more vulnerable.

Maturation

Biological or psychological change in participants as a natural function of time. In a 12-month intervention study with adolescent participants, cognitive development attributable to age alone could mimic a treatment effect.

Selection Bias

Pre-existing differences between groups that are not attributable to the treatment. In non-randomised designs — the majority of educational and social research — participants self-select into conditions, meaning the groups may differ systematically on variables correlated with the outcome. Selection bias is arguably the most pervasive threat in applied social research.

Testing Effects

Exposure to a pre-test instrument sensitising participants or causing practice effects that alter their post-test performance independently of any intervention. Repeated administration of cognitive assessments, for example, can inflate post-test scores through familiarity.

Instrumentation

Changes in the calibration of measurement instruments, or in the criteria applied by human raters, over the course of a study. Inter-rater drift — where coders gradually shift their scoring conventions over time — is a well-documented instrumentation threat in qualitative research that relies on human coding. Researchers who design studies involving human raters should consult established procedures for measuring coding consistency, including inter-rater reliability and Cohen’s kappa.

Regression to the Mean

When participants are selected on the basis of extreme scores on a variable, subsequent measurement tends to produce scores closer to the group mean, regardless of any treatment. Studies targeting high-risk individuals based on screening scores are particularly susceptible.

Attrition (Differential Dropout)

When dropout from the study is not random — when those who leave one condition differ systematically from those who leave another, or from those who remain — the retained sample is no longer comparable across groups. Attrition is especially threatening in randomised controlled trials where control-group participants, receiving no benefit, are more likely to withdraw.

Confounding and Spuriousness

A third variable causally related to both the IV and DV can create or mask an apparent relationship between them. Confounding is the arch-enemy of causal inference in observational research, and its control is the central rationale for randomisation.

Strategies to Strengthen Internal Validity

Design hierarchy for internal validity (strongest to weakest):

  1. Randomised controlled experiment (RCT) with double blinding
  2. Randomised experiment without blinding
  3. Quasi-experimental design (e.g., regression discontinuity, difference-in-differences)
  4. Propensity score matched observational study
  5. Unmatched observational study with statistical covariate control
  6. Purely descriptive or correlational study

Specific mitigations include:

  • Random assignment: The only technique that controls for both known and unknown confounds simultaneously.
  • Control groups: A no-treatment or active-control group provides the counterfactual comparison necessary for causal inference.
  • Blinding: Participants blind to condition assignment cannot exhibit demand characteristics; assessors blind to assignment cannot exhibit experimenter bias.
  • Standardised procedures: Detailed protocols, training for research assistants, and fidelity checks minimise instrumentation drift.
  • Intent-to-treat analysis: Including all randomised participants regardless of dropout preserves the protection against selection bias that randomisation originally provided.
  • Covariate adjustment: ANCOVA, hierarchical regression, or mixed-effects models can statistically partial out pre-existing group differences when full randomisation is not possible.

Construct Validity

Construct validity is perhaps the most intellectually demanding of the four validity types, because it requires moving between the theoretical and empirical levels simultaneously. The term was introduced by Cronbach and Meehl in their landmark 1955 paper “Construct Validity in Psychological Tests,” published in Psychological Bulletin. Their central contribution was the concept of the nomological network: a system of theoretical propositions and empirical relationships that define what a construct is, how it relates to other constructs, and what observable indicators it should predict.

Construct validity is threatened whenever the operationalisation of a construct — the specific measure, treatment, or outcome used in the study — does not adequately capture the intended theoretical entity. As Shadish, Cook, and Campbell (2002) put it, construct validity is the degree to which inferences are warranted about the higher-order constructs that the instances of the study represent.

Subtypes of Construct Validity

Contemporary psychometrics distinguishes several interlocking subtypes:

Content Validity

The degree to which the items or components of a measure adequately sample the full domain of the construct. A scale purporting to measure “research self-efficacy” would lack content validity if it only assessed confidence in quantitative methods while omitting qualitative, mixed-methods, and writing-related self-efficacy. Content validity is typically evaluated through systematic expert review and comparison against a theoretically derived content map.

Convergent Validity

High convergent validity is demonstrated when a new measure correlates strongly with established measures of the same construct (or closely related constructs). If a novel scale measuring academic burnout correlates highly with established burnout inventories (e.g., the Maslach Burnout Inventory — Student Survey), convergent validity is supported. Multitrait-multimethod (MTMM) matrix analysis, proposed by Campbell and Fiske (1959), remains the classical technique for jointly assessing convergent and discriminant validity.

Discriminant Validity

The complement of convergent validity: a measure should not correlate strongly with measures of theoretically unrelated constructs. A scale measuring academic burnout should discriminate from a scale measuring, say, introversion — even if some empirical overlap exists. Confirmatory factor analysis (CFA) with the average variance extracted (AVE) criterion is the modern standard for demonstrating discriminant validity, following Fornell and Larcker (1981).

Criterion-Related Validity

The degree to which a measure predicts a specified criterion. Concurrent validity refers to prediction of a contemporaneous criterion; predictive validity refers to prediction of a future criterion. A measure of PhD completion intention has predictive validity if it significantly predicts actual completion rates two years later.

Threats to Construct Validity

Shadish, Cook, and Campbell (2002) enumerate several construct validity threats specific to experimental and quasi-experimental research:

  • Inadequate explication of constructs: If the theoretical construct is not defined precisely before operationalisation, any measure will be partially valid at best. Vague constructs (“wellbeing,” “engagement”) require systematic conceptual analysis before measurement. A detailed treatment of how theoretical constructs relate to your overall research framework appears in our companion guide on theoretical vs conceptual frameworks.
  • Construct confounding: The operationalisation contains multiple distinct constructs, making it impossible to attribute effects to the intended one alone. A social skills training programme that simultaneously delivers peer mentoring and cognitive restructuring confounds the two components at the construct level.
  • Mono-operation bias: Using only a single operationalisation of a construct. If “academic stress” is measured only via self-report, the findings may reflect response style as much as true stress. Multiple operationalisations (self-report, cortisol, heart rate variability, observer ratings) triangulate the construct more convincingly.
  • Mono-method bias: Assessing all constructs via the same method (e.g., all via self-report surveys) inflates correlations through shared method variance, a problem that MTMM designs are specifically built to detect.
  • Treatment-sensitive factorial structure: The factor structure of an outcome measure may differ between treated and control groups, meaning the measure is not measuring the same construct in both conditions — a threat to the comparability of comparisons.
  • Demand characteristics and experimenter expectancy: Participants infer the study’s hypotheses and adjust their behaviour accordingly; researchers unknowingly communicate expected results. Both introduce systematic construct-irrelevant variance.

Strategies to Strengthen Construct Validity

  • Define constructs via systematic review of the theoretical literature before selecting or developing measures. For qualitative research methods, this may involve working definitions derived from phenomenological or interpretive traditions.
  • Pilot-test instruments; conduct exploratory factor analysis (EFA) followed by confirmatory factor analysis (CFA) in independent samples.
  • Use multiple operationalisations and methods to converge on constructs (triangulation). When your measure is a Likert scale, the guide on how to design a Likert scale questionnaire covers psychometric best practices for ensuring adequate construct coverage.
  • Implement blinding to reduce demand characteristics and experimenter expectancy.
  • Report construct reliability statistics (Cronbach’s α, McDonald’s ω) alongside validity evidence; construct validity requires adequate reliability as a prerequisite.
Validity types at a glance

Validity type Core question Primary threat Key remedy
Statistical conclusion Is there a real covariation? Low power A priori power analysis
Internal Is the covariation causal? Selection bias / confounding Random assignment
Construct Do measures match constructs? Mono-operation / mono-method bias Multiple operationalisations
External Do results generalise? WEIRD samples / lab settings Diverse sampling / field studies

Based on Shadish, Cook, & Campbell (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference.

External Validity

External validity is the degree to which a causal relationship, once established within a study, holds across different people, settings, treatments, and outcomes — the “PSTO” framework advanced by Shadish, Cook, and Campbell (2002). External validity is often colloquially glossed as “generalisability,” though the Shadish–Cook–Campbell framework is considerably more precise: external validity concerns generalisation of a causal relationship, not merely descriptive statistics.

A study with high external validity supports claims such as: “This intervention to improve thesis completion rates works not only for full-time PhD students in STEM at Russell Group universities in 2026, but also for part-time students in social sciences at other institution types and in different calendar years.”

Threats to External Validity

Interaction of the Causal Relationship with Units (Persons)

The causal effect may be moderated by person-level characteristics. An effect found in undergraduate student samples may not replicate in community or clinical samples. The “WEIRD” problem — the predominance of Western, Educated, Industrialised, Rich, and Democratic samples in psychological and educational research — represents an overarching threat to cross-population external validity that remains actively debated in 2026.

Interaction of the Causal Relationship with Settings

Laboratory settings impose artificial constraints that may not exist in naturalistic contexts. The controlled environment that bolsters internal validity (removing history threats, standardising treatment) is precisely the feature that undermines ecological validity — a related but distinct concept referring to fidelity to real-world conditions.

Interaction of the Causal Relationship with Treatments

The specific operationalisation of the treatment may differ from how the treatment would be implemented in practice. A mindfulness intervention delivered by trained researchers over 8 weeks in a research context may differ substantially from a brief digital mindfulness app used independently — even if both are labelled “mindfulness.”

Interaction of the Causal Relationship with Outcomes

The specific outcome measures used may not represent the full range of possible outcomes. An intervention may improve researcher-designed outcome measures while leaving clinician-assessed or patient-reported outcomes unchanged.

Context-Dependent Mediation

The mechanism through which a cause produces its effect may be context-specific. The effect of peer feedback on academic writing quality may operate through social motivation in collectivist educational cultures but through cognitive elaboration in individualist ones, meaning the same surface-level intervention operates through different causal pathways in different settings.

Strategies to Strengthen External Validity

  • Diverse, representative sampling: Use probability sampling from a well-defined target population rather than convenience samples where possible. Where convenience samples are unavoidable, explicitly characterise the sample and qualify the scope of conclusions accordingly.
  • Multi-site replication: Conducting the same study across multiple sites, institutions, or countries increases confidence that effects are not site-specific. For a systematic literature review, meta-analysis can formally test for heterogeneity of effects across studies differing in populations and settings.
  • Field experiments and natural experiments: Moving from laboratory to naturalistic settings (field experiments) or exploiting naturally occurring assignment variation (natural experiments, regression discontinuity designs) improves ecological validity while preserving some causal rigour.
  • Conceptual replication: Replicating a finding with different operationalisations of the treatment and outcome — conceptual rather than direct replication — tests whether the causal claim generalises across alternative constructions of the phenomenon.
  • Explicit population definition: Specify the target population at the design stage and document how the sample was drawn. The research proposal should include this specification as a matter of routine.

Validity vs. Reliability: The Critical Distinction

Reliability and validity are related but logically distinct properties of measurement. Reliability refers to the consistency or reproducibility of scores — the same instrument should yield the same results when administered to the same respondents under the same conditions (test-retest reliability), across different items purportedly measuring the same construct (internal consistency), or across different raters applying the same coding scheme (inter-rater reliability).

The logical relationship between the two is asymmetric:

  • Reliability is necessary but not sufficient for validity. A highly consistent measure that consistently measures the wrong construct is reliable but not valid. A bathroom scale that consistently reads 5 kg too heavy is reliable; it is not a valid measure of true body weight.
  • Validity implies a minimum level of reliability. An instrument that produces random, inconsistent scores cannot systematically capture any construct, and therefore cannot be valid. Unreliability introduces random measurement error, which attenuates correlations and reduces construct and statistical conclusion validity simultaneously.

This distinction has direct implications for research design. Researchers should establish reliability evidence (Cronbach’s α for internal consistency, ICC or Cohen’s κ for inter-rater agreement) as a prerequisite step before making validity claims. A coefficient alpha of .70 is a widely cited minimum threshold for adequate internal consistency (Nunnally, 1978), though McDonald’s ω is now preferred as a less assumption-laden alternative.

Managing Validity Trade-Offs by Design

One of the most enduring observations in research methodology is that validity types trade off against one another. Cook and Campbell (1979) were explicit that researchers cannot simultaneously maximise all four types in a single study. Understanding the trade-offs allows researchers to make principled design choices rather than simply defaulting to whichever design is most convenient.

The Internal–External Validity Trade-Off

The most widely discussed trade-off is between internal and external validity. Laboratory experiments maximise control over confounds (internal validity) at the cost of artificial conditions that may not generalise (external validity). Conversely, large-scale observational studies sampling from diverse, representative populations are highly externally valid but face severe internal validity threats from confounding. Neither extreme is universally preferable; the appropriate balance depends on the research question.

A study asking “Can intervention X produce any effect on outcome Y under ideal conditions?” prioritises internal validity. A study asking “Does intervention X produce a meaningful effect on Y in routine practice?” prioritises external validity. Both are legitimate scientific questions requiring different designs.

The Construct–Internal Validity Tension

Tight experimental control may narrow the operationalisation of constructs to the point of under-representing them (poor construct validity). A laboratory analogue of “workplace stress” created by imposing a cognitive load task bears limited resemblance to the multi-faceted stressors of real occupational contexts. Researchers must weigh whether a tractable, controllable operationalisation sufficiently represents the target construct.

Statistical Power and Other Validity Types

Increasing sample size improves statistical conclusion validity but does not, by itself, improve internal, construct, or external validity. Conversely, a highly internally valid RCT with a small sample may fail statistical conclusion validity if it is underpowered. This is the logic behind power analysis as a design tool rather than a post-hoc rationalisation.

Worked Examples Across Research Designs

Example 1: RCT of a Writing Intervention for PhD Students

A researcher randomly assigns 120 PhD students to either a structured daily writing intervention or a control condition (free writing). The primary outcome is weekly word count after eight weeks.

  • Statistical conclusion validity: With n = 60 per group and an expected medium effect size (d = 0.50), a priori power analysis confirms adequate power at .82 (α = .05). Internal consistency of the word-count diary (ICC = .91) is strong.
  • Internal validity: Random assignment controls selection bias. Standardised intervention delivery and assessor blinding reduce instrumentation threats. Attrition is monitored; an intent-to-treat analysis is pre-specified.
  • Construct validity: “Daily writing productivity” is operationalised via verified word count and a validated writing self-efficacy scale, not self-report alone — reducing mono-operation and mono-method bias.
  • External validity: The sample is drawn from a single UK research-intensive university. Findings may not generalise to part-time students, students in non-Anglophone countries, or professional doctorate candidates. These limits are explicitly acknowledged.

Example 2: Observational Study of Assessment Format and Academic Achievement

A large-N observational study uses administrative data from 8,000 undergraduates across 12 universities to examine whether open-book examinations are associated with higher final-year grades than closed-book formats.

  • Statistical conclusion validity: Large sample size provides high power; multilevel modelling accounts for within-university clustering, avoiding inflated standard errors.
  • Internal validity: Selection bias is the dominant threat: universities may systematically differ in student ability, course difficulty, or instructional quality, all of which correlate with both examination format adoption and achievement outcomes. Propensity score matching is used to create pseudo-comparable groups, but residual confounding cannot be ruled out.
  • Construct validity: “Assessment format” is operationalised as a binary (open vs. closed book), which ignores variation within each category. “Academic achievement” is operationalised as final-year grade — a valid but narrow operationalisation that excludes deeper learning or transferable skills.
  • External validity: The multi-university design significantly enhances external validity relative to a single-site study. Explicit characterisation of the university types in the sample allows readers to assess applicability to their own contexts.

Example 3: Qualitative Case Study

A researcher conducts semi-structured interviews with 18 doctoral supervisors across three institutions to explore how they conceptualise their role during candidates’ writing-up phase. The study applies thematic analysis following Braun and Clarke’s reflexive framework.

In qualitative research, the classical Cook–Campbell validity vocabulary requires translation. Internal validity maps onto credibility (Lincoln and Guba, 1985): are the findings an accurate representation of participants’ experiences? Construct validity maps onto the adequacy of conceptual definitions guiding the analysis. External validity maps onto transferability: rather than asserting that findings generalise statistically, the researcher provides “thick description” enabling readers to judge the applicability of the findings to their own contexts. Statistical conclusion validity, as a category, does not apply to non-inferential qualitative work; some qualitative methodologists apply parallel criteria such as confirmability and dependability instead.

Reporting Validity in Your Thesis or Paper

Validity threats and their mitigations belong in several sections of a research paper or thesis chapter:

In the Methods Section

Describe how each major threat was addressed prospectively. For internal validity, justify your design choice and describe randomisation or matching procedures. For construct validity, cite the psychometric evidence for your measures (α, ω, CFI, RMSEA). For statistical conclusion validity, report your power analysis with parameters.

In the Results Section

Report assumption checks alongside inferential statistics. Include effect sizes with 95% confidence intervals for every key comparison. Transparency here supports replication and meta-analytic synthesis.

In the Discussion / Limitations

Discuss residual threats — those that the design could not fully eliminate — and their likely direction of bias. Use the four-validity typology explicitly: “A threat to the internal validity of this study is…”, “The external validity of these findings is bounded by…”. Reviewers and examiners recognise this vocabulary and expect it in doctoral-level work.

For practical guidance on structuring the discussion and limitations chapter, researchers frequently consult the research methodology and citation standards guide as a reference for academic writing norms across disciplines.

Frequently Asked Questions

What is the difference between internal validity and external validity?

Internal validity concerns whether a causal relationship exists within your study — specifically, whether the independent variable genuinely caused changes in the dependent variable, ruling out confounds. External validity concerns whether that causal relationship generalises beyond the specific sample, setting, and time period of your study. High internal validity tells you your study is causally sound; high external validity tells you those causal conclusions apply elsewhere. Most research designs involve an inherent trade-off between the two.

What is construct validity and why does it matter?

Construct validity is the degree to which your operationalisation — your measure, manipulation, or outcome — accurately captures the theoretical construct it is supposed to represent. Introduced by Cronbach and Meehl (1955), it matters because if your measure does not actually tap the construct of interest, all downstream conclusions about that construct are flawed, regardless of how internally or externally valid the study is otherwise. Subtypes include content, convergent, discriminant, and criterion-related validity.

Can a study have high internal validity but low external validity?

Yes, and this is a classic tension in experimental research. A tightly controlled laboratory experiment typically maximises internal validity by eliminating confounds, but its highly artificial setting may mean the causal relationship found does not generalise to real-world populations or contexts — reducing external validity. Researchers must make deliberate trade-off decisions depending on whether their primary aim is causal proof-of-concept (prioritise internal validity) or applied generalisation (prioritise external validity).

What are the main threats to internal validity?

Cook and Campbell (1979) and Shadish, Cook, and Campbell (2002) identified core threats including: history (external events occurring between measurements), maturation (natural change in participants over time), selection bias (pre-existing differences between groups), instrumentation (changes in measurement tools or procedures), attrition (differential dropout from study conditions), testing effects (pre-test exposure altering responses), and regression to the mean (extreme initial scorers moving toward the average on retest). Randomisation is the most powerful single mitigation, as it controls simultaneously for all known and unknown confounds.

What is statistical conclusion validity?

Statistical conclusion validity is the validity of inferences about whether a statistical covariation exists between the independent and dependent variables. Threats include low statistical power (increasing Type II error risk), violation of statistical test assumptions, fishing for significance across multiple tests without correction (inflating Type I error), unreliable measures, and restriction of range. It is the foundational validity type and must be established before addressing internal, construct, or external validity.

How do validity and reliability differ in research methodology?

Reliability refers to the consistency or reproducibility of a measurement — whether the same instrument yields the same results under the same conditions. Validity refers to whether the measure actually captures what it claims to measure. A measure can be reliable without being valid (consistently measuring the wrong thing), but a valid measure must have some degree of reliability, since a highly inconsistent measure cannot systematically capture any true construct. Researchers should establish reliability evidence before making validity claims; the two properties are complementary, not interchangeable.


Primary sources: Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin. | Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302. | Scribbr (2026). Internal validity in research.

Write your thesis with AI

Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.

Leave a Reply

Your email address will not be published. Required fields are marked *