Writing the Results Chapter of a Sports Science Dissertation in 2026: Effect Sizes, Reliability and Worked Examples

·

Writing the Results Chapter of a Sports Science Dissertation in 2026: Effect Sizes, Reliability and Worked Examples

Sports science results chapters go wrong in a predictable way. You tested fourteen participants, your intervention group improved countermovement jump height by 1.8 cm, the p value came out at 0.09, and you now have to write eight pages about a result you have been told is “not significant”. So you write a chapter that sounds apologetic, fill it with tables nobody can read, and hand your examiner a document that undersells a perfectly respectable piece of work.

The discipline has moved on from that framing. Modern sports and exercise science reporting is built on effect sizes with confidence intervals, on measurement reliability, and on whether a change is large enough to matter practically — not on whether p crossed 0.05 in a sample of fourteen. This guide shows how to write the chapter that way, with sentences you can adapt directly.

What belongs in the results chapter — and what does not

Results present what you found. They do not explain it, compare it to the literature, or speculate about mechanisms — all of that belongs in the discussion. The single most common structural error in sports science dissertations is a results chapter that starts interpreting on page two.

A defensible order:

  1. Participant characteristics and flow (who was recruited, who completed, who was excluded and why)
  2. Reliability and measurement quality for your key outcomes
  3. Assumption checks, briefly
  4. Descriptive statistics for each outcome
  5. Inferential results, hypothesis by hypothesis, with effect sizes
  6. Individual responses, where relevant
  7. Secondary and exploratory outcomes, clearly labelled as such

Step 1: Report participants and flow honestly

Open with a table of participant characteristics: n, sex, age, stature, body mass, training history and any performance benchmark relevant to your outcome (VO2max, one-repetition maximum, competitive standard). Report mean and standard deviation, and state the sample descriptor precisely — “recreationally active university students” is not “trained athletes”, and the difference constrains everything you can claim.

Then account for everyone. If nineteen volunteered and fourteen completed, state what happened to the other five: injury unrelated to the study, failure to attend the post-test, non-compliance with the training programme. A CONSORT-style flow diagram is standard for any intervention study and takes twenty minutes to draw.

Worked sentence: “Nineteen participants were recruited, of whom 14 completed all testing sessions (10 male, 4 female; age 21.3 ± 2.1 y; stature 1.76 ± 0.08 m; body mass 74.2 ± 9.6 kg). Three withdrew due to injuries sustained outside the study and two missed more than 20% of the prescribed sessions and were excluded per the a priori compliance criterion.”

Step 2: Report reliability before you report change

This is the step that separates competent sports science dissertations from the rest, and it is the one most often skipped. If you are claiming a 1.8 cm improvement in jump height, the reader needs to know how much your measurement wobbles when nothing changes. Otherwise the improvement is uninterpretable.

Report, for each key measure:

  • Intraclass correlation coefficient (ICC) with its 95% confidence interval, specifying the model and form used (for example, two-way mixed effects, absolute agreement, single measurement).
  • Typical error of measurement, in the original units, and/or the coefficient of variation as a percentage.
  • Smallest worthwhile change, commonly estimated as 0.2 × the between-participant standard deviation, which gives you a practical benchmark for interpreting your effect.

These come from your familiarisation or pilot sessions, and if you did not run repeat trials, you can often cite established reliability values for the test from the literature — while noting that you did not establish them in your own laboratory. That caveat is worth writing; the absence of any reliability information is worth marks lost.

Worked sentence: “Test-retest reliability from the two familiarisation sessions was high for countermovement jump height (ICC = 0.96, 95% CI 0.89 to 0.99), with a typical error of 1.1 cm and a coefficient of variation of 3.2%. The smallest worthwhile change, calculated as 0.2 × between-participant SD, was 1.0 cm.”

Where you are comparing two measurement devices rather than repeat trials — a validation study of a new jump mat against a force platform, for instance — reliability is the wrong framework and agreement is the right one. Use a Bland-Altman plot with the systematic bias and 95% limits of agreement, not a correlation coefficient. Two devices can correlate at 0.99 while one reads consistently 3 cm high.

Step 3: State assumption checks briefly, then move on

One short paragraph. Which distributional assumptions you checked, how, and what you found. If a repeated-measures ANOVA violated sphericity, name the correction you applied — this is expected and its absence is noticed. Our walkthrough of repeated-measures ANOVA, Mauchly’s test and Greenhouse-Geisser covers the sequence and exactly what to report.

Worked sentence: “Normality of residuals was assessed by inspection of Q-Q plots and confirmed by Shapiro-Wilk tests (all p > 0.05). Mauchly’s test indicated a violation of sphericity for the time × group interaction, χ²(2) = 8.94, p = 0.011, so Greenhouse-Geisser corrected values are reported (ε = 0.78).”

Resist the temptation to delete awkward data points to make assumptions hold. In sports science an extreme value is often a genuinely exceptional athlete rather than an error, and our discussion of whether to remove outliers covers where that line falls and how to document the decision.

Step 4: Descriptives that a reader can actually use

Report mean ± SD for normally distributed variables, median and interquartile range for skewed ones — RPE, session counts and injury data are usually skewed. Give units and decimal places consistently, and match precision to the measurement: jump height to one decimal place in centimetres, VO2max to one decimal in ml·kg⁻¹·min⁻¹, sprint times to two decimals in seconds.

For pre-post data, the informative number is the change score with its 95% confidence interval, not just the two means. Give all three.

Step 5: Inferential results, with effect sizes as the headline

Every inferential statement in a modern sports science chapter should contain four elements: the test statistic with degrees of freedom, the exact p value, the effect size, and a confidence interval on the effect or the mean difference.

Choosing the effect size

  • Cohen’s d for the difference between two independent means. With samples below roughly 20 per group — that is, most dissertations — use Hedges’ g, which applies a small-sample correction. Stating that you did so signals genuine methodological awareness.
  • Partial eta squared (η²ₚ) for ANOVA effects, though note it is not comparable across studies with different designs.
  • Standardised mean difference using the between-participant SD for pre-post designs, which is the convention in sport and exercise science and is more interpretable than using the SD of the change scores.
  • Percentage change with confidence interval for performance variables, which practitioners find far more usable than a standardised value.

Common magnitude thresholds in sports science are 0.2 small, 0.6 moderate, 1.2 large and 2.0 very large — deliberately more conservative than Cohen’s original psychology-derived benchmarks, because meaningful performance changes in trained athletes are small. Cite the scale you use rather than asserting the labels.

Worked sentence: “Countermovement jump height increased in the plyometric group from 34.2 ± 4.8 cm to 36.9 ± 5.1 cm, a mean change of 2.7 cm (95% CI 1.4 to 4.0 cm), while the control group changed from 33.8 ± 5.2 cm to 34.1 ± 5.4 cm (0.3 cm, 95% CI −0.9 to 1.5 cm). The time × group interaction was significant, F(1, 12) = 9.84, p = 0.009, η²ₚ = 0.45, with a between-group effect size for the change of g = 1.06 (95% CI 0.31 to 1.78), representing a moderate to large effect. The mean change in the plyometric group exceeded the smallest worthwhile change of 1.0 cm.”

That last clause is what turns a statistic into a sports science finding.

Choosing the test itself

If you are still deciding, work backwards from your design. Two independent groups on one occasion calls for an independent-samples t test — our walkthrough of running one in SPSS covers Levene’s test and the write-up. One group tested twice calls for a paired-samples t test, covered in our paired t test guide. Two groups across two or more time points — the classic training-intervention design — calls for a mixed ANOVA, and our comparison of ANOVA designs sets out which member of the family your design needs. For relationships between physiological and performance variables, our guide to correlation in SPSS covers Pearson, Spearman and reporting both. Where sample sizes are very small or data are ordinal, such as RPE or wellness scores, our overview of non-parametric tests covers the alternatives.

Step 6: Report individual responses

Group means hide the thing coaches most want to know: did everyone improve, or did four people improve substantially and ten not at all? Individual response variability is a live topic in sport science and reporting it well is a straightforward way to add value.

Present it as a scatter or line plot of individual pre-post changes, or a bar chart of each participant’s change with the smallest worthwhile change marked as a shaded band. Then state how many participants exceeded that threshold.

Worked sentence: “Individual responses varied considerably (Figure 4). Nine of the 14 participants in the plyometric group exceeded the smallest worthwhile change of 1.0 cm, three showed changes within the typical error of measurement, and two decreased. Baseline jump height was not associated with the magnitude of change (r = −0.21, p = 0.47).”

Step 7: Tables and figures that earn their space

  • Every table and figure must be referenced in the text and interpreted in the following sentence.
  • Never present the same data in both a table and a figure.
  • Use error bars that you have defined in the caption — SD, SEM and 95% CI look similar and mean different things. Confidence intervals are increasingly preferred.
  • Keep decimal places consistent within a column.
  • Report exact p values to three decimals (p = 0.009), reserving p < 0.001 for values below that threshold.
  • Follow APA 7 conventions for statistical notation unless your department specifies otherwise: italicised test statistics, no leading zero on values that cannot exceed one.

Small formatting inconsistencies of this kind are among the most frequently flagged issues in marking, and they are entirely avoidable — as our analysis of statistical reporting errors in published research shows, even peer-reviewed work gets them wrong at a surprising rate.

What to do when nothing reached significance

With 12 to 20 participants, a non-significant result is the statistically expected outcome for anything but a large effect, and examiners know this. What loses marks is treating it as a failure.

Write it as a finding. Report the point estimate and its confidence interval, and interpret the interval’s range: “the mean difference was 0.9 cm (95% CI −0.7 to 2.5 cm); the study was therefore unable to rule out an effect as large as 2.5 cm, which would exceed the smallest worthwhile change.” Then state the effect size the design could have detected — a post-hoc sensitivity analysis in G*Power gives you this in five minutes and is far more informative than observed power, which is a function of your p value and tells the reader nothing new. Our guide on what to do when your results are not significant covers how to frame the whole chapter around this.

For the wider structure — how this chapter connects back to your methods and forward into the discussion — see our discipline guide to writing a sports science dissertation.

Turn your SPSS output into a finished results chapter

You have the numbers. What takes the days is writing them up in correct statistical prose, with effect sizes, confidence intervals and consistent APA notation across every paragraph. Tesify takes your design, your variables and your output and drafts a results chapter that reports each hypothesis properly, names the assumptions you checked, and keeps every table and citation formatted to your department’s requirements.

Start your thesis with Tesify — free to begin

Frequently asked questions

Should I use Cohen’s d or Hedges’ g in a sports science dissertation?

Use Hedges’ g when your groups have fewer than about 20 participants each, which describes most dissertation samples. Cohen’s d is upwardly biased in small samples, and Hedges’ g applies a correction factor that removes most of that bias. The two converge as sample size grows, so with large samples the choice is immaterial. Whichever you report, name it explicitly, give a confidence interval, and cite the magnitude thresholds you are using to label it.

What is the smallest worthwhile change and do I have to report it?

The smallest worthwhile change is the smallest alteration in a performance measure considered practically meaningful, most commonly estimated as 0.2 multiplied by the between-participant standard deviation. It is not universally required, but reporting it substantially strengthens a sports science results chapter because it lets you say whether an observed change matters in practice rather than only whether it was statistically detectable. Compare it against your typical error of measurement: a change smaller than your measurement error cannot be confidently attributed to the intervention.

Do I need to report post-hoc power for a non-significant result?

Do not report observed post-hoc power calculated from your obtained effect — it is a direct transformation of the p value and adds no information. What is genuinely useful is a sensitivity analysis: given your sample size, alpha and desired power, what is the smallest effect you could have detected? Report that alongside the confidence interval on your effect, and interpret both. This tells the reader what your study could and could not have found.

How long should a sports science results chapter be?

In a 10,000 to 12,000 word undergraduate dissertation, the results chapter is typically 1,500 to 2,500 words plus tables and figures — usually the shortest substantive chapter. Concision is a virtue here: the chapter should be dense with numbers and free of interpretation. If yours is running long, check whether you are explaining findings rather than reporting them, since that material belongs in the discussion.

Should error bars show standard deviation or confidence intervals?

Use standard deviation when your purpose is to show the spread of individual values in the sample, and 95% confidence intervals when your purpose is to show the precision of an estimated mean or mean difference. Confidence intervals are increasingly preferred in sport and exercise science because they connect directly to inference. Whichever you choose, define it explicitly in every figure caption — undefined error bars are ambiguous and are routinely queried by markers.

Write your thesis with AI

Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.

Leave a Reply

Your email address will not be published. Required fields are marked *