The phrase John Hattie effect size often appears beside charts that rank educational influences, yet readers may not know what the values measure, why 0.40 is emphasized, or whether the benchmark applies to nursing education. Hattie used effect sizes to compare the estimated influence of educational factors on student achievement across large bodies of research. The framework supports useful comparative questions, but its rankings and hinge point are not universal statistical or clinical rules. This guide explains the framework, demonstrates its application to nursing education, and examines its principal limitations. (Hattie, 2009; Hattie, 2015)
What Is the John Hattie Effect Size?
An effect size estimates the magnitude and direction of a difference, relationship, or intervention effect. A standardized mean difference expresses a difference between means in standard-deviation units, helping researchers compare outcomes measured with different scales. Many Visible Learning values use standardized measures related to Cohen’s d, although the precise calculations and meanings vary across the studies and meta-analyses included in the synthesis. (Lakens, 2013; Hattie, 2009)
Effect size is not the same as statistical significance, sample size, clinical significance, practical value, or causal proof. Suppose one group of nursing students completes a simulation-based medication-safety program while another receives standard instruction. The difference between their post-intervention mean scores can be divided by a pooled standard deviation, placing the result on a common scale.
That standardized value describes the size and direction of the observed difference. By itself, however, it does not demonstrate that the simulation caused the difference, that students will retain the knowledge, or that patient medication errors will decline.
Who Is John Hattie and What Is Visible Learning?
John Hattie is an education researcher associated with the Visible Learning research program. The original Visible Learning book synthesized more than 800 meta-analyses. A later higher-education paper discussed an evidence base of approximately 1,200 meta-analyses, while the 2023 sequel drew on more than 2,100 meta-analyses. (Hattie, 2009; Hattie, 2015; Hattie, 2023)
Individual studies → meta-analyses → Hattie’s broader synthesis of meta-analyses
A primary study investigates a defined group of learners, educators, classrooms, or institutions. A meta-analysis statistically combines findings from multiple studies addressing a related question. Hattie’s work then synthesizes findings across many meta-analyses, a process sometimes described as a meta-meta-analysis or synthesis of meta-analyses.
The goal was not simply to ask whether an educational practice has any effect. It was to compare the relative magnitude of numerous influences on achievement. The resulting rankings should nevertheless not be treated as permanent, context-free, or universally applicable facts.
What Does an Effect Size of 0.40 Mean in Hattie’s Research?
Approximately 0.40 became Hattie’s hinge point because it represented the average effect across the educational research synthesized in the framework. When interpreted as a standardized mean difference, 0.40 indicates that two relevant means differ by approximately four-tenths of a standard deviation. (Hattie, 2009; Hattie, 2015)
| The value 0.40 is a contextual reference point within Hattie’s educational synthesis. It is not a universal definition of educational, statistical, practical, or clinical importance. |
The frequently repeated claim that an effect size of 0.40 always represents one year of learning requires caution. Translating standardized effects into months or years assumes that learning rates, tests, ages, subjects, study designs, intervention durations, and comparison conditions are sufficiently comparable. Those assumptions often do not hold. Researchers should interpret the measured outcome rather than automatically converting a standardized effect into a fixed period of learning.
Key interpretation
- 40 is Hattie’s benchmark for comparing educational influences within the Visible Learning synthesis.
- It is not the minimum value required for an intervention to matter.
- It does not automatically represent one year of progress.
- It does not apply universally to nursing clinical outcomes.
- It should be interpreted with the study design, population, outcome, uncertainty, and implementation context.
Understanding Hattie’s Effect-Size Barometer
The Hattie effect-size barometer visually places an estimate relative to zero and the 0.40 hinge point. Its purpose is to support discussion about the direction and approximate magnitude of an educational influence. It is not a formal hypothesis-testing system, a study-quality scale, or a universal rule for choosing instructional practices.
| Approximate effect-size range | General position | Appropriate interpretation | Important caution |
| Below 0 | Negative direction | The measured outcome favored the comparison condition or declined. | Check coding, design, measurement, bias, and context before concluding that the intervention caused harm. |
| Close to 0 | Little observed difference | Minimal standardized separation on the measured outcome. | A near-zero average may conceal meaningful subgroup differences. |
| Above 0 but below 0.40 | Positive but below the hinge point | Some positive difference smaller than Hattie’s synthesis-wide average. | It is not automatically ineffective or unimportant. |
| Around or above 0.40 | At or above the hinge point | A positive difference comparable with or greater than Hattie’s reference value. | It does not prove effectiveness, causation, quality, or transferability. |
| Substantially above 0.40 | Larger positive estimate | A potentially important educational difference. | Large estimates may reflect small samples, weak comparisons, bias, poor measurement, or short follow-up. |
These zones are interpretive guides. Their boundaries are not statistical-significance thresholds, quality grades, or clinical decision rules.
Hattie Effect Size Compared With Cohen’s Effect-Size Guidelines
Cohen’s commonly cited conventions describe standardized mean differences around 0.20 as small, 0.50 as medium, and 0.80 as large. Cohen intended these as broad conventions rather than universal laws. Their usefulness depends on whether more appropriate discipline-specific benchmarks are available. (Cohen, 1988; Lakens, 2013)
| Feature | Hattie’s 0.40 hinge point | Cohen’s conventions |
| Primary purpose | Comparing educational influences | General interpretation of standardized mean differences |
| Context | Hattie’s educational synthesis | Broad behavioral and social science guidance |
| Central reference | Approximately 0.40 | Approximately 0.20, 0.50, and 0.80 |
| Universal cutoff? | No | No |
| Application to nursing | Primarily learning-related nursing education outcomes | Potentially applicable across nursing studies when justified |
A result below 0.40 can still be meaningful when the outcome is difficult to change, the intervention is inexpensive, benefits accumulate over time, the population has substantial needs, the intervention prevents harm, or the comparison group receives an active treatment. Conversely, a large standardized effect may be unreliable or unimportant when it is based on a very small sample, a weak comparison group, selective reporting, poor measurement, high risk of bias, short follow-up, or an unusually narrow population.
The estimate must therefore be evaluated within the study context and alongside the principles of inferential statistics in nursing research.
Worked Nursing Education Example
Consider an original hypothetical study evaluating a clinical simulation intervention:
- Intervention group: n = 50
- Comparison group: n = 50
- Intervention mean: 84
- Comparison mean: 80
- Pooled standard deviation: 10
d = (84 − 80) ÷ 10
d = 0.40
The numerator is the four-point difference between the group means. The denominator expresses that difference relative to the pooled variability in scores. The appropriate interpretation is: “The intervention group scored approximately 0.40 standard deviations above the comparison group.”
This value is approximately equal to Hattie’s educational hinge point. However, the number alone does not establish that the intervention caused the difference or that the result is educationally important. Interpretation must also consider the confidence interval, p value, baseline comparability, attrition, reliability of the assessment, fidelity of the simulation, prior clinical experience, transfer to real clinical performance, and persistence over time.
Two studies can produce the same effect size but have different practical implications. One may show d = 0.40 for durable gains on a validated medication-safety assessment. Another may show d = 0.40 for an immediate change on a poorly validated confidence scale. The numerical effects are equal, but the outcomes, evidence quality, durability, and implications are not.
Can Hattie’s Effect Sizes Be Applied to Nursing Education?
The framework is most relevant when nursing research examines learning-related outcomes such as knowledge acquisition, clinical reasoning, skill performance, confidence, retention, examination scores, simulation performance, feedback, reflective learning, online instruction, continuing professional development, and patient-education knowledge outcomes.
Systematic reviews of nursing simulation commonly use standardized mean differences and confidence intervals when combining learning outcomes measured with different instruments. These reviews also demonstrate why nursing-specific evidence is more informative than a general educational ranking alone. (Cant & Cooper, 2017; Görücü et al., 2024)
Clinical outcomes such as mortality, blood pressure, infection, readmission, pain, medication errors, and quality of life require a different level of interpretation. Depending on the research question and data, researchers may need risk ratios, odds ratios, risk differences, hazard ratios, raw mean differences, standardized mean differences, or numbers needed to treat. Hattie’s 0.40 benchmark should not replace clinically grounded interpretation or guidance on choosing statistical tests in nursing research.
Nursing Education Examples of Effect-Size Interpretation
All values below are hypothetical. They are not Hattie rankings for nursing interventions.
| Nursing education intervention | Outcome | Hypothetical effect size | Appropriate interpretation |
| High-fidelity simulation | Clinical reasoning score | 0.55 | Moderate positive standardized difference requiring contextual evaluation. |
| Formative feedback | Medication-calculation accuracy | 0.42 | Slightly above the hinge point, but not automatically effective in every institution. |
| Online module | Knowledge score | 0.28 | Below 0.40 but potentially useful if inexpensive, accessible, and scalable. |
| Peer-assisted learning | Skills assessment | 0.18 | Small effect that may matter when accumulated or implemented widely. |
| Poorly designed intervention | Student confidence | −0.15 | Possible negative effect requiring investigation of delivery, measurement, and design. |
Does an Effect Size Below 0.40 Mean an Intervention Does Not Work?
No. Being below the hinge point is not equivalent to being ineffective. Interpretation depends on natural development, ordinary instructional progress, comparison conditions, cost, feasibility, baseline achievement, outcome sensitivity, exposure duration, implementation quality, population characteristics, measurement error, confidence intervals, and educational consequences.
An effect size of 0.25 from a low-cost intervention reaching thousands of nursing students may be more valuable than an effect size of 0.60 from an expensive intervention that cannot be implemented outside one institution. The comparator also matters: a simulation compared with no additional teaching may produce a larger effect than the same simulation compared with a strong existing skills course.
Why Effect-Size Rankings Can Be Misleading
Different underlying outcomes
Achievement, motivation, attendance, confidence, knowledge, and clinical performance are not interchangeable. A high effect on self-reported confidence cannot automatically be equated with an equally sized effect on objectively assessed clinical competence.
Different comparison conditions
An intervention compared with no instruction may appear stronger than one compared with high-quality standard instruction. Rankings can therefore reflect study-design differences as well as intervention performance.
Variation among studies
An average effect can conceal substantial heterogeneity across populations, institutions, durations, subject areas, and implementation approaches.
Dependence on study quality
Combining weak primary studies does not automatically produce a strong conclusion. Poor allocation, unreliable measures, high attrition, selective reporting, confounding, and short follow-up remain important after synthesis.
Overlapping categories and evidence
Educational influences may be defined differently across reviews, and meta-analyses may include some of the same primary studies, creating dependence.
Changes over time
Technologies, curricula, learner populations, assessments, and standard teaching practices evolve. Older estimates may not transfer directly to contemporary nursing education.
Publication bias
Positive or statistically significant findings may be more likely to be published, located, or included, potentially exaggerating pooled effects.
Context loss
A single average can hide why an intervention succeeded in one setting and failed in another. Staffing, instructor expertise, learner preparation, technology, class size, and implementation fidelity matter.
A confidence interval describes uncertainty around the estimated average. A prediction interval addresses the wider range of effects that may plausibly occur in a comparable future setting. A pooled mean can therefore be positive even when an intervention is less effective, or potentially harmful, in some contexts. (Deeks et al., 2024)
Major Criticisms of John Hattie’s Effect-Size Approach
Critiques of Visible Learning focus on methodology and interpretation rather than on Hattie personally. Concerns include combining different research designs, comparing effects calculated through different procedures, variation in outcome measures and intervention duration, dependence among meta-analyses, overlap in primary studies, study-quality variation, publication bias, heterogeneity, broad classifications, reliance on a single average, oversimplification through rankings, and the risk of treating correlational evidence as intervention evidence.
Researchers have argued that diverse evidence cannot always be placed on one scale without losing assumptions, context, and uncertainty. Critics also warn that visually precise rankings may appear more stable, causal, and transferable than the underlying studies justify. (Bergeron, 2017; Wecker et al., 2017; Nielsen & Klitmøller, 2021)
These criticisms do not erase Hattie’s contribution. Visible Learning helped make effect size more prominent in educational discussion, encouraged attention to magnitude rather than p values alone, and provided a starting point for evidence-informed debate. The framework is most useful as an entry point for investigation, not as a substitute for reading the underlying evidence.
How Nursing Researchers Should Evaluate a Hattie Effect Size
- What exact outcome was measured?
- What study designs contributed to the estimate?
- What was the comparison condition?
- How similar were the participants to nursing students or educators in my setting?
- Was the effect consistent across studies?
- What was the confidence interval?
- Was publication bias assessed?
- Were the primary studies of acceptable quality?
- Was the intervention implemented consistently?
- Is the outcome educationally or clinically meaningful?
- What resources are required?
- Is there more recent or discipline-specific evidence?
- Does the finding apply to undergraduate, postgraduate, clinical, or continuing nursing education?
- Are there possible harms or unintended consequences?
The rank of an intervention should not replace local evidence, professional judgment, student needs, curriculum requirements, resource assessment, or implementation planning.
How to Report Hattie-Related Effect Sizes in a Nursing Dissertation
APA-style reporting should identify the effect-size statistic, its direction and magnitude, confidence interval, sample size, comparison, outcome, statistical uncertainty, practical meaning, and relevant limitations. Reporting standards emphasize estimates and uncertainty rather than significance tests alone. (Appelbaum et al., 2018; Lakens, 2013)
Model APA 7-style paragraph
| A moderate standardized difference was observed between students who received the simulation intervention and those who received standard instruction, d = 0.40, 95% CI [0.01, 0.79]. Although this value corresponds approximately to Hattie’s educational hinge point, the result was interpreted using the study’s nursing education context, confidence interval, outcome validity, and practical implications rather than treating 0.40 as a universal effectiveness threshold. |
The precise statistics must match the student’s actual analysis. A dissertation should not report a p value without also interpreting the estimated magnitude and uncertainty. Related guidance includes reporting ANOVA effect sizes in SPSS and analyzing Likert-scale nursing data.
Hattie Effect Size Versus Statistical and Clinical Significance
| Concept | Main question |
| Statistical significance | Is the result sufficiently inconsistent with the specified null model at the chosen threshold? |
| Effect size | How large is the estimated difference or relationship? |
| Educational significance | Does the result meaningfully improve learning or performance? |
| Clinical significance | Does the result meaningfully improve patient care or health outcomes? |
A small effect may become statistically significant in a very large sample. A large observed effect may remain statistically uncertain in a small sample because its confidence interval is wide. A statistically significant improvement in nursing knowledge may not transfer to clinical practice, whereas a modest clinical improvement may matter greatly when it prevents serious harm.
Common Misinterpretations of the Hattie Effect Size
Myth 1: An effect size above 0.40 proves an intervention works
It does not. Bias, confounding, attrition, unreliable measurement, baseline imbalance, and statistical imprecision may remain even when the estimate exceeds 0.40.
Myth 2: An effect size below 0.40 should be ignored
It should not. Cost, feasibility, reach, comparison-group quality, cumulative benefit, outcome importance, and uncertainty can make a smaller effect valuable.
Myth 3: An effect size of 0.80 is exactly twice as beneficial as 0.40
Standardized effects describe differences in standard-deviation units. They should not automatically be translated into proportional educational or clinical benefits.
Myth 4: The highest-ranked influence is always the best choice
Rank does not establish quality, suitability, affordability, ethics, scalability, or local fit.
Myth 5: Hattie’s benchmark applies to all nursing and healthcare outcomes
The benchmark originated in educational synthesis. Clinical outcomes require appropriate measures, evidence, and interpretation standards.
Myth 6: Effect size establishes causation
Causal interpretation depends on the research design, allocation procedures, control of bias and confounding, temporal order, and alternative explanations.
Myth 7: A single effect size tells the whole story
A point estimate must be accompanied by confidence intervals, heterogeneity, study quality, outcome validity, implementation information, and practical interpretation.
Practical Takeaways for Nursing Students
- Hattie’s 0.40 value is an educational reference point, not a universal cutoff.
- Effect size describes magnitude, not certainty or causation.
- Values below 0.40 can still be meaningful.
- Values above 0.40 can still be biased or poorly supported.
- Nursing education findings should be evaluated using discipline-specific evidence.
- Clinical outcomes require clinically appropriate effect measures and benchmarks.
- Confidence intervals and study quality should accompany effect-size interpretation.
- Hattie’s rankings should guide questions rather than end the analysis.
Frequently Asked Questions
What is John Hattie’s effect size?
John Hattie’s framework uses standardized estimates to compare the magnitude of educational influences across a large synthesis of meta-analyses. The term does not refer to one unique formula. Many underlying estimates resemble standardized mean differences, but their calculation and meaning vary. The framework supports comparative educational discussion rather than replacing study-level evaluation. Readers should still inspect the measured outcome, design, comparison group, uncertainty, and relevance.
What does an effect size of 0.40 mean?
When treated as a standardized mean difference, 0.40 indicates that two means differ by approximately four-tenths of a standard deviation. In Hattie’s framework, it is also the hinge point used to compare an influence with the average effect in the educational synthesis. It does not automatically mean one year of learning, clinical importance, causal proof, or guaranteed success in a nursing program.
Why does Hattie use 0.40 as the hinge point?
Hattie used approximately 0.40 because it represented the average effect across the educational evidence synthesized in Visible Learning. The rationale was to compare influences with a more demanding educational reference than zero, since many educational activities produce some positive change. The hinge point is contextual to that synthesis and is not a universal threshold for every discipline or outcome.
Is 0.40 considered a large effect size?
Not under Cohen’s commonly cited general conventions, where approximately 0.20, 0.50, and 0.80 are described as small, medium, and large. Under that framework, 0.40 lies between small and medium. Those labels are conventions rather than laws. Meaningfulness depends on the outcome, comparison group, measurement quality, cost, duration, population, uncertainty, implementation, and consequences.
Is Hattie’s 0.40 effect size the same as Cohen’s d?
Not exactly. Cohen’s d is a standardized mean-difference statistic. Hattie’s 0.40 hinge point is a benchmark derived from his synthesis of educational evidence. Many Visible Learning estimates use measures related to Cohen’s d, but the hinge point is not a separate calculation formula. Researchers must identify the actual statistic they calculated.
Does an effect size below 0.40 mean an intervention is ineffective?
No. A value below 0.40 may still represent a useful, affordable, scalable, preventive, or cumulative benefit. Its value also depends on the comparison condition and how difficult the outcome is to change. Researchers should examine the confidence interval, study quality, costs, implementation, outcome importance, and plausible alternatives rather than applying a pass-or-fail rule.
What is the Hattie effect-size barometer?
The barometer is a visual device that places educational effect estimates along a continuum from negative effects through values near zero to larger positive effects, with 0.40 marked as the hinge point. It helps readers compare magnitudes. Its zones are interpretive guides, not statistical-significance thresholds, quality grades, or universal rules for selecting teaching strategies.
Can Hattie’s effect sizes be used in nursing education research?
They can inform interpretation when outcomes are educational, such as knowledge, skills, clinical reasoning, examination performance, feedback, or retention. Nursing researchers should still prioritize current, discipline-specific evidence. Hattie’s benchmark should not replace analysis of study design, confidence intervals, implementation, learner characteristics, or clinical relevance.
What are the main criticisms of John Hattie’s research?
Critics question the comparability of diverse designs, outcomes, calculations, durations, and comparison conditions. They also raise concerns about study overlap, dependence among meta-analyses, publication bias, heterogeneity, quality variation, broad categories, and context loss through rankings. The concern is not that synthesis is useless, but that a single average may appear more universal, causal, and precise than the evidence permits.
How should I report an effect size in a nursing dissertation?
Name the statistic, state its direction and magnitude, and report the confidence interval, sample size, groups, and outcome. Explain the estimate in the nursing context and acknowledge uncertainty and limitations. A p value should not substitute for magnitude. When referring to Hattie, describe 0.40 as an educational comparison point rather than a universal effectiveness threshold.
Is statistical significance the same as effect size?
No. Statistical significance evaluates the result’s compatibility with a null model and is affected by sample size and variability. Effect size estimates the magnitude of a difference or relationship. A small effect may be statistically significant in a large sample, while a large observed effect may be imprecise in a small one. Strong interpretation reports both magnitude and uncertainty.
Can Hattie’s effect size be used to interpret clinical outcomes?
It should not be used as a universal clinical benchmark. Clinical outcomes may require risk ratios, odds ratios, hazard ratios, risk differences, mean differences, standardized mean differences, or numbers needed to treat. Their importance should be judged using patient-centered thresholds, baseline risk, harms, benefits, precision, and clinical guidance, not whether an estimate exceeds Hattie’s educational hinge point.
Using Hattie’s Effect Sizes Responsibly in Nursing Education
Hattie’s work made effect-size interpretation more prominent in educational decision-making and encouraged readers to ask how much influence an educational practice may have. Its most famous value, 0.40, is best treated as a contextual benchmark within Visible Learning, not as a universal rule.
Nursing students and researchers should evaluate research design, confidence intervals, study quality, heterogeneity, comparison conditions, implementation fidelity, nursing relevance, practical significance, and clinical implications. Hattie’s rankings can guide questions, but they should not end the analysis.
Students who need support selecting, calculating, interpreting, or reporting effect sizes can request nursing dissertation data analysis help for context-specific guidance.
References
Appelbaum, M., Cooper, H., Kline, R. B., Mayo-Wilson, E., Nezu, A. M., & Rao, S. M. (2018). Journal article reporting standards for quantitative research in psychology: The APA Publications and Communications Board task force report. American Psychologist, 73(1), 3–25. https://doi.org/10.1037/amp0000191
Bergeron, P.-J. (2017). How to engage in pseudoscience with real data: A criticism of John Hattie’s arguments in Visible Learning from the perspective of a statistician. McGill Journal of Education, 52(1), 237–246. https://mje.mcgill.ca/article/view/9475
Cant, R. P., & Cooper, S. J. (2017). Use of simulation-based learning in undergraduate nurse education: An umbrella systematic review. Nurse Education Today, 49, 63–71. https://doi.org/10.1016/j.nedt.2016.11.015
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates. https://www.routledge.com/Statistical-Power-Analysis-for-the-Behavioral-Sciences/Cohen/p/book/9780805802832
Deeks, J. J., Higgins, J. P. T., Altman, D. G., McKenzie, J. E., & Veroniki, A. A. (2024). Analysing data and undertaking meta-analyses. In J. P. T. Higgins et al. (Eds.), Cochrane handbook for systematic reviews of interventions (Version 6.5). Cochrane. https://training.cochrane.org/handbook/current/chapter-10
Görücü, S., Türk, G., & Karaçam, Z. (2024). The effect of simulation-based learning on nursing students’ clinical decision-making skills: Systematic review and meta-analysis. Nurse Education Today, 140, 106270. https://doi.org/10.1016/j.nedt.2024.106270
Hattie, J. (2009). Visible learning: A synthesis of over 800 meta-analyses relating to achievement. Routledge. https://www.routledge.com/Visible-Learning-A-Synthesis-of-Over-800-Meta-Analyses-Relating-to-Achievement/Hattie/p/book/9780415476188
Hattie, J. (2015). The applicability of Visible Learning to higher education. Scholarship of Teaching and Learning in Psychology, 1(1), 79–91. https://doi.org/10.1037/stl0000021
Hattie, J. (2023). Visible learning: The sequel: A synthesis of over 2,100 meta-analyses relating to achievement. Routledge. https://www.routledge.com/Visible-Learning-The-Sequel-A-Synthesis-of-Over-2100-Meta-Analyses-Relating-to-Achievement/Hattie/p/book/9781032462034
Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t tests and ANOVAs. Frontiers in Psychology, 4, Article 863. https://doi.org/10.3389/fpsyg.2013.00863
Nielsen, K., & Klitmøller, J. (2021). Blind spots in visible learning: A critique of John Hattie as an educational theorist. Nordic Psychology, 73(3), 268–283. https://doi.org/10.1080/19012276.2021.1962731
Wecker, C., Vogel, F., & Hetmanek, A. (2017). Visionär und imposant – aber auch belastbar? Eine Kritik der Methodik von Hatties “Visible Learning.” Zeitschrift für Erziehungswissenschaft, 20(1), 21–40. https://doi.org/10.1007/s11618-016-0696-0