Which Statistical Test Should an Education Dissertation Use? The 2026 Guide to Classroom, School and Large-Scale Assessment Data

·

Which Statistical Test Should an Education Dissertation Use? The 2026 Guide to Classroom, School and Large-Scale Assessment Data

Education data break the assumptions of the standard test-selection chart in three predictable ways: students sit inside classrooms inside schools, so observations are not independent; interventions are assigned to intact classes rather than to individuals, so the pretest matters more than the design chart admits; and the outcomes are test scores, rubric ratings and survey scales that were built by someone else. This guide starts from those three facts and works through the analyses that education dissertations in the United States actually run, from the classroom study with two intact sections to the secondary analysis of a national dataset.

Key points first

  • If students are clustered in classrooms or schools and the treatment varies at the cluster level, the default analysis is a multilevel model, not an ordinary regression or ANOVA; the intraclass correlation from an empty model tells you how much the clustering matters.
  • For the two-classroom pretest-posttest design, ANCOVA with the pretest as covariate is the standard analysis when groups were not randomized; gain-score t-tests are acceptable only when baseline equivalence has been shown.
  • Report effect sizes as Hedges’ g with a confidence interval; the What Works Clearinghouse, whose Procedures and Standards Handbook Version 5.0 (2022) is the reference most US committees use for evidence quality, expects g and describes baseline differences by the same metric.
  • Secondary analysis of NCES datasets (ECLS-K, HSLS:09, NAEP, PISA and TIMSS in the international sphere) requires the sampling weights and, for assessment scores, the plausible values; running an unweighted regression on these files is a methods-chapter failure.
  • Survey scales need a reliability estimate in your sample and, if you built or adapted the instrument, an exploratory or confirmatory factor analysis before any score is used as a variable.

Start with the unit of analysis, not with the test

Three questions settle most of the choice. Who received the treatment: individual students, whole classrooms or whole schools? What is the outcome: a continuous score, a category such as graduated or not, an ordinal rating, or a count? And how many times was each student measured: once, twice around an intervention, or repeatedly across a year? The general decision matrix in our guide to which statistical test to use answers the second and third questions for any field; the first question is the one that makes education different, and it is the one students most often skip.

Matching the education design to the analysis

Common education dissertation designs and the analysis each one calls for
Design Typical example Analysis What to report
Two intact classes, pretest and posttest, no randomization Flipped versus lecture section of the same course ANCOVA with pretest as covariate; check homogeneity of regression slopes Adjusted means, F, partial eta squared, Hedges’ g on adjusted means
Students randomized within one class Two feedback formats assigned by lottery Independent-samples t-test or ANCOVA if a pretest exists Means, t, p, g with confidence interval
Several classrooms or schools per condition District curriculum pilot in 12 schools Two-level multilevel model, students in schools, treatment at level 2 Intraclass correlation, fixed effects, variance components, cluster-adjusted g
Same students measured three or more times Reading fluency each month across a year Growth-curve or repeated-measures multilevel model Intercept and slope, slope difference by group
Binary outcome On-time graduation, course pass, program completion Logistic regression, multilevel if clustered Odds ratios with confidence intervals, classification accuracy
Ordinal outcome Rubric levels, proficiency bands Ordinal logistic regression, or non-parametric comparison for two groups Proportional odds check, cumulative odds ratios
Survey scales as variables Teacher self-efficacy predicting retention intention Reliability, factor analysis, then regression or SEM Alpha or omega per subscale, factor loadings, model fit
Policy change at a known date New attendance policy, state test change Interrupted time series or difference-in-differences Level and slope change, parallel-trends evidence
Single student or very small group, repeated measures Behavioral intervention in special education Single-case design with visual analysis and Tau-U or a between-case standardized mean difference Phase data, effect estimate, replication across cases
Secondary analysis of a national dataset HSLS:09 predictors of STEM major choice Survey-weighted regression with replicate weights; plausible values for assessments Weighted estimates, design-based standard errors

Nested data: the education problem the standard charts ignore

When a teacher delivers an intervention to a whole class, every student in that class shares the teacher, the room and the time of day. Their scores are correlated, and a test that treats 60 students as 60 independent observations will report standard errors that are too small and p-values that are too optimistic. The fix is a multilevel model with students at level 1 and classrooms or schools at level 2; our pillar on multilevel models for nested data explains the intraclass correlation and the design effect in general terms. For an education dissertation the practical rules are these: report the intraclass correlation from an unconditional model (in US school data it commonly falls between .10 and .25 for achievement outcomes), put the treatment variable at the level where it was assigned, and count degrees of freedom by clusters, not by students. If you have only two classrooms, one per condition, no model can separate the treatment effect from the classroom effect; say so in the limitations and analyze at the student level with the caveat, rather than pretending the problem is not there.

Pretest-posttest with intact groups: ANCOVA or gain scores?

The most common master’s and EdD design is two existing sections, one taught the new way, both tested before and after. The choice is between ANCOVA on the posttest with the pretest as covariate and a t-test on gain scores. ANCOVA is usually preferred because it adjusts for baseline differences that non-random assignment produces and has more power; it assumes the relationship between pretest and posttest is the same in both groups, which you test by adding the interaction term and hoping it is not significant. Gain scores are defensible when the groups are equivalent at baseline. The What Works Clearinghouse treats a baseline difference of up to 0.05 standard deviations as equivalent, requires statistical adjustment for differences between 0.05 and 0.25, and does not accept designs with larger baseline gaps for non-randomized comparisons. The comparison of ANOVA variants in which ANOVA do you need covers the mechanics of ANCOVA and its repeated-measures cousin.

Concentric layers showing student figures inside a classroom inside a school outline, with a clustered-dots chart beside them, illustrating nested education data
Students nested in classrooms nested in schools: the structure that decides whether an ordinary test is even admissible.

Effect sizes the education literature expects

Education committees increasingly read effect sizes before p-values, partly because the What Works Clearinghouse reports every finding as Hedges’ g and partly because policy audiences think in months of learning. Report g (the small-sample-corrected standardized mean difference) with its confidence interval for any two-group comparison, partial eta squared for ANOVA-family analyses, and odds ratios for logistic models. Do not translate g into months of learning without citing the conversion you used, and do not describe a g of 0.10 as small just because Cohen’s benchmarks say so; in K-12 intervention research the median effect in rigorous evaluations is well below 0.20, and a 0.10 effect on a state test at scale is a policy-relevant result. Our guide to effect sizes and confidence intervals covers the computation; what matters here is the benchmark you compare against, which should come from education meta-analyses, not from psychology.

Test scores, scales and the instrument question

Three kinds of outcome dominate education dissertations, and each carries its own analytic obligation.

  • Standardized test scores. Use scale scores, not percentiles or proficiency labels, as the dependent variable; percentiles are ordinal and proficiency bands throw away most of the variance. If the test changed form between pretest and posttest, you need vertically scaled scores or a within-year z-score, and you need to say which.
  • Researcher-made tests and rubrics. Report inter-rater reliability for rubric scores (Cohen’s kappa or intraclass correlation, depending on the scale), and treat a rubric level as ordinal unless you can defend interval spacing.
  • Survey scales. Whether you use the Motivated Strategies for Learning Questionnaire, the Teacher Sense of Efficacy Scale or an instrument you adapted (our guide to choosing a theoretical framework for an education dissertation lists which instrument belongs to which theory), report Cronbach’s alpha or McDonald’s omega for each subscale in your sample. If you changed items or translated the scale, run an exploratory factor analysis on a development sample or a confirmatory factor analysis against the published structure before using the scores. Sum scores from an instrument whose structure you have not checked are the most common reason a Chapter 4 is sent back.

Secondary analysis of NCES and international datasets

Dissertations using the Early Childhood Longitudinal Study, the High School Longitudinal Study of 2009 (which followed more than 23,000 ninth graders from 944 schools through a 2012 follow-up, a 2013 update, high school transcripts, a 2016 second follow-up and postsecondary transcripts), the National Assessment of Educational Progress, or PISA and TIMSS face two requirements that classroom studies do not. The samples are complex, so every estimate needs the study’s sampling weight and its replicate weights or design variables for standard errors; running the same regression in SPSS without the complex-samples module gives standard errors that are wrong by a factor you cannot predict. And assessment scores in NAEP, PISA and TIMSS are released as plausible values, usually five or ten per student, which must be analyzed separately and combined by Rubin’s rules rather than averaged. Public-use files are enough for most dissertations; geocoded and some transcript files require a restricted-use license through your institution. Our directory of free datasets and open data repositories lists the education section.

Three traps specific to education dissertations

  1. Analyzing classroom-level treatments at the student level with no acknowledgment. Reviewers trained on What Works Clearinghouse standards will flag this immediately. If clustering is unavoidable and clusters are few, say so, adjust what you can, and interpret conservatively.
  2. Testing the same students on a pretest that teaches the posttest. A pretest identical to the posttest given two weeks apart inflates gains in both groups; use parallel forms or a longer interval, and report which.
  3. Running a dozen subscale comparisons and reporting the one that reached .05. Preregister the primary outcome in your proposal, treat the rest as exploratory, and apply a correction or a hierarchical testing plan.

Write the analysis plan for your education dissertation

Give Tesify your design, your unit of assignment and your outcome type, and it drafts the data analysis section: the model matched to the nesting, the assumption checks, the effect size you will report and the reporting template the committee expects, with every methods source formatted in APA 7 by Auto Bibliography.

Start your education dissertation analysis plan with Tesify, free to begin

Frequently asked questions

Do I need a multilevel model if I only have two classrooms?

No, and it would not run reliably with two clusters. Analyze at the student level, report that the treatment was assigned at the classroom level, and state in the limitations that the classroom effect cannot be separated from the treatment effect. Multilevel models need enough level-2 units, commonly cited as at least 20 to 30, to estimate variance components with any stability.

Should I use ANCOVA or repeated-measures ANOVA for a pretest-posttest study?

Use ANCOVA when the question is whether groups differ at posttest after adjusting for pretest; it is the standard for non-randomized intact groups. Use repeated-measures ANOVA when the question is about change over three or more time points within the same students.

What effect size should I report in an education dissertation?

Hedges’ g with a 95 percent confidence interval for two-group comparisons, partial eta squared for ANOVA designs, and odds ratios for binary outcomes. Compare the result with effect sizes from education meta-analyses rather than with Cohen’s general benchmarks.

Can I run a t-test on Likert-scale survey items?

On a single item, treat the response as ordinal and use a Mann-Whitney U or ordinal regression. On a validated multi-item scale score with acceptable reliability, a t-test or regression on the scale score is standard practice, provided you have reported the reliability in your sample.

Do I have to use sampling weights with NCES data?

Yes. The samples are stratified and clustered by design, and the weights are what make estimates representative of the population. Use the weight and replicate weights specified in the study documentation, and use software that supports complex-samples analysis.

What test do I use for teacher-level survey data collected in several schools?

If teachers are your unit and schools are the cluster, the same nesting logic applies: check the intraclass correlation for the outcome, and use a multilevel model if it is non-trivial and you have enough schools. With few schools, use school as a fixed effect or cluster-robust standard errors and say why.

Write your thesis with AI

Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.

Leave a Reply

Your email address will not be published. Required fields are marked *