You run your analysis, SPSS spits out p = .003, and for a moment everything feels worth it. The result is statistically significant. Then your supervisor reads the draft and writes one question in the margin: "Yes, but how big is the effect?" That question separates statistical significance from practical significance, and it trips up more thesis students than any assumption check or missing-data problem. A result can be real in the statistical sense and still be too small to matter to anyone. This article explains the difference, shows why large samples make it worse, and walks through what you actually need to report.
What statistical significance actually tells you
A p-value answers a narrow question: if there were truly no effect in the population, how often would random sampling alone produce data at least as extreme as yours? When p = .003, the answer is about 3 times in 1,000. That is unlikely enough that most researchers reject the null hypothesis and conclude something other than chance is going on.
Notice how little that statement contains. It says nothing about the size of the effect. It says nothing about whether the effect matters clinically, educationally, or financially. It does not even tell you the probability that your hypothesis is true, which is the single most common misreading students bring to their defense. If you want to check your own interpretation, paste your result into our P-Value Interpreter and compare its plain-English output against what you wrote in your results chapter. We also cover the common misreadings in detail in our guide to interpreting p-values.
The threshold itself is a convention, not a law of nature. Fisher suggested .05 as a convenient benchmark in the 1920s, and the field kept it. Nothing changes in reality between p = .049 and p = .051, even though one gets published more easily than the other.
A p-value measures how surprising your data would be if the null hypothesis were true. It measures compatibility with "no effect," never the size or importance of the effect itself.
What practical significance means
Practical significance asks a different question: is the effect large enough to matter in the real world? The answer depends on context, cost, and consequences rather than on any statistical calculation.
Consider a study finding that a new teaching method raises exam scores by 0.4 points on a 100-point scale, with p < .001 thanks to a sample of 8,000 students. The effect is almost certainly real. It is also almost certainly worthless. No school will retrain its staff, rewrite its materials, and restructure its timetable for less than half a point. A statistician would say the result is statistically significant but practically negligible.
Clinical fields have formalized this idea. Depression researchers work with the concept of a minimal clinically important difference: on the Hamilton Depression Rating Scale, the UK's NICE guidelines historically treated a 3-point drug–placebo difference as the smallest change worth caring about. Kirsch and colleagues (2008), analyzing FDA trial data for four antidepressants, found a mean difference of about 1.8 points. The trials were statistically significant. Whether they were practically significant became one of the loudest debates in psychiatry that decade, precisely because significance testing could not settle it.
The reverse situation exists too. A pilot study with 15 participants per group might show a 12-point improvement that fails to reach p < .05. The effect could be practically enormous and statistically invisible, simply because the study lacked power. Dismissing it as "no effect" would be a mistake of the same kind.
The sample size problem
Here is the mechanical reason the two concepts come apart: p-values depend on both the size of the effect and the size of the sample. Hold the effect constant and grow the sample, and the p-value shrinks toward zero. With enough participants, any nonzero difference, no matter how trivial, becomes statistically significant.
The arithmetic is blunt. For a correlation to reach significance at the .05 level, r needs to exceed roughly 1.96 divided by the square root of N. With 100 participants, you need r = .20. With 10,000, r = .02 will do. With a million, r = .002 clears the bar. Nobody believes a correlation of .002 tells us anything useful about human behavior, yet it earns the same asterisk in a results table as r = .60.
The Facebook emotional contagion experiment is the textbook case. Kramer, Guillory, and Hancock (2014) manipulated the news feeds of 689,003 users and found statistically significant effects on the emotional tone of subsequent posts. The effect sizes ran from d = 0.001 to d = 0.02. Translated into raw behavior, that meant something on the order of one altered word per thousand. The headlines said Facebook could manipulate emotions. The effect sizes said it could barely nudge them.
This cuts the other way for underpowered studies, which is why power analysis belongs at the design stage rather than the write-up stage. Our article on sample size and power analysis covers how to pick an N that can detect the smallest effect you would actually care about, which is the cleanest way to keep statistical and practical significance aligned from the start.
With a large enough sample, everything becomes statistically significant. A tiny p-value in a study with thousands of participants may reflect precision, not importance.
Effect sizes: the bridge between the two
Effect sizes quantify how big an effect is, on a scale that does not change when the sample grows. They are the standard answer to the supervisor's marginal question, and every major journal in psychology, education, and medicine now expects them. You can compute the common ones from your existing output with our Effect Size Calculator, and our full effect size guide covers the less common ones.
Cohen's d
For comparing two group means, Cohen's d expresses the difference in standard deviation units: the mean difference divided by the pooled standard deviation. A d of 0.5 means the groups differ by half a standard deviation. Cohen (1988) offered benchmarks of 0.2 for small, 0.5 for medium, and 0.8 for large, while warning that they were rough defaults for when no field-specific standard existed. In some areas a d of 0.2 is impressive; a d of 0.2 for a cheap public-health intervention applied to millions of people can save lives, while the same d for an expensive individual therapy may not justify a single session.
Eta-squared and partial eta-squared
For ANOVA designs, eta-squared reports the proportion of total variance in the outcome explained by the factor. A partial eta-squared of .06 means the factor accounts for 6% of the relevant variance. Conventional benchmarks are .01 (small), .06 (medium), and .14 (large). SPSS prints partial eta-squared on request, which makes the frequent absence of effect sizes from student theses hard to excuse. Be aware that partial eta-squared values from designs with different numbers of factors are not directly comparable, a detail reviewers sometimes check.
Odds ratios
For binary outcomes, the odds ratio tells you how much the odds of an event change between conditions. An OR of 2.0 means the odds double. Odds ratios come with a trap of their own: when the baseline event is rare, a dramatic-sounding OR can describe a minuscule absolute change. Doubling a risk of 1 in 100,000 still leaves a risk of 2 in 100,000. Reporting the absolute risk difference alongside the OR keeps readers honest about scale.
| Measure | Typical use | Small | Medium | Large |
|---|---|---|---|---|
| Cohen's d | Two-group mean comparisons (t-tests) | 0.2 | 0.5 | 0.8 |
| Pearson's r | Correlations | .10 | .30 | .50 |
| Eta-squared | ANOVA | .01 | .06 | .14 |
| Odds ratio | Binary outcomes | 1.5 | 2.5 | 4.3 |
Treat the benchmarks in this table as a starting point for interpretation, then argue from your own field's literature. A sentence like "the observed d of 0.35 is smaller than the d of 0.51 reported in Miller (2019) for a comparable intervention" carries far more weight in a discussion chapter than "the effect was small to medium."
Famous results that were real but tiny
The aspirin arm of the Physicians' Health Study (1989) is a classic illustration. Among 22,071 male physicians randomized to aspirin or placebo, aspirin reduced heart attacks with p < .00001, and the trial's monitoring board stopped it early on ethical grounds. Rosnow and Rosenthal later pointed out that the effect corresponded to r² of about .0011, meaning aspirin explained roughly a tenth of one percent of the variance in heart attack occurrence. The finding changed medical practice anyway, because heart attacks are lethal, aspirin costs pennies, and even a 0.9 percentage-point absolute reduction across millions of men adds up. Practical significance depends on stakes, not just size.
Contrast that with the drug-trial pattern that regulators see constantly: a pain medication demonstrating p < .001 for a 0.3-point improvement on a 100-point visual analogue scale. Patients cannot feel a 0.3-point change; most studies put the minimal detectable improvement on such scales at 8 to 12 points. Statistical significance here certifies only that the trial was large enough to measure a trivial difference precisely.
Education research produces the same pattern. Large-scale studies of school interventions routinely report significant effects with d values around 0.05, which corresponds to moving the average student from the 50th to the 52nd percentile. Whether that justifies the cost of implementation is a policy question, and no p-value can answer it.
What supervisors and reviewers actually look for
Thesis committees and journal reviewers read results sections with a fairly predictable checklist. They want the test statistic with its degrees of freedom, the exact p-value, an effect size, and a confidence interval. Then they read the discussion chapter to see whether you interpreted the effect size or merely reported it.
The single most damaging pattern in student writing is what one might call significance-only reasoning: the results chapter reports p-values, and the discussion chapter treats every p < .05 as a confirmed, important finding and every p > .05 as proof of no effect. Reviewers flag both moves. A significant result with d = 0.11 deserves a sentence acknowledging that the effect, while unlikely to be chance, is small enough to question its usefulness. A nonsignificant result from a sample of 40 deserves a sentence about power rather than a confident claim that the groups do not differ.
Supervisors also look for evidence that you anchored your interpretation in the literature. If prior studies of your intervention found effects around d = 0.6 and you found d = 0.15, that discrepancy is the most interesting thing in your thesis. Explaining it will earn you more credit at the defense than the significance star ever will.
What APA 7 says about reporting both
The seventh edition of the APA Publication Manual is explicit on this point. Its journal article reporting standards (JARS-Quant) require effect sizes and confidence intervals for primary outcomes, and the manual states that it is almost always necessary to include some measure of effect size alongside significance tests. This position goes back to Wilkinson and the Task Force on Statistical Inference (1999), who told authors to "always present effect sizes for primary outcomes."
In practice, an APA-7 results sentence looks like this:
Participants in the intervention group scored higher (M = 74.2, SD = 8.1) than controls (M = 70.9, SD = 8.4), t(198) = 2.83, p = .005, d = 0.40, 95% CI [0.12, 0.68].
Every element earns its place. The p-value establishes that chance is an unlikely explanation. The d tells the reader the difference amounts to four tenths of a standard deviation. The confidence interval shows the effect could plausibly be anywhere from trivial (0.12) to substantial (0.68), which tempers any grand claims in the discussion. Writing one sentence like this per hypothesis, then interpreting the effect size against your field's norms, satisfies nearly every reviewer objection before it is raised.
APA 7 expects three things for every primary result: the exact p-value, an effect size, and a confidence interval. Reporting significance alone has been below standard since at least 1999.
Putting it together in your thesis
Run your planned tests and report exact p-values. Compute an effect size for every hypothesis test, whether or not the result was significant, and add confidence intervals where your software provides them. In the discussion, interpret magnitude first and significance second, and compare your effect sizes against published findings rather than against Cohen's generic labels alone. Where the effect is significant but small, say so plainly and discuss whether it would matter to anyone outside the study. That combination of honesty and context is what examiners mean when they praise a "mature" results chapter, and it takes perhaps two extra sentences per finding to achieve.