Nursing students frequently report a t test and p value without explaining the magnitude of the observed difference. A statistically significant result may represent a small difference with limited practical importance, while a nonsignificant result from a small sample may remain compatible with a clinically or educationally meaningful effect.
A second problem occurs when researchers calculate Cohen’s d without first identifying the study design. The standard formula for two independent groups is not automatically appropriate when the same patients, nurses, or nursing students are measured twice. Pre-test/post-test and matched-pairs studies can produce several standardized mean differences, including (d_z), (d_{av}), and (d_{rm}). These measures use different standardizing values and can therefore produce different numerical results from the same dataset.
Before calculating Cohen’s d effect size, nursing researchers should be able to answer several questions:
- Are the observations independent or paired?
- Which mean should be subtracted from which?
- Does pooled variability represent the intended comparison?
- Which denominator should be used for pre-test/post-test data?
- Should a small-sample correction be applied?
- How was the confidence interval calculated?
- Does the observed standardized effect represent a meaningful difference in the original clinical or educational scale?
This guide explains how to select, calculate, interpret, and report Cohen’s d in independent-group, paired-sample, pre-test/post-test, matched, and one-sample nursing studies. It also explains when Hedges’ g, Glass’s delta, or a different effect-size family may be preferable.
For a broader comparison of effect sizes used with categorical outcomes, correlations, regression models, ANOVA, and nonparametric analyses, see Effect Size in Nursing Research.
Quick Answer: What Is Cohen’s d?
Cohen’s d is a standardized mean-difference effect size. It expresses a difference between means relative to a selected standard deviation.
Its general form is:
[
d=
\frac{\text{difference between means}}
{\text{selected standard deviation}}
]
A Cohen’s (d=0.50) means that the means differ by approximately one-half of the standard deviation used in the denominator.
The sign shows direction:
- A positive value means the first mean in the numerator is higher.
- A negative value means the first mean is lower.
- The absolute value shows standardized magnitude.
Cohen’s d does not automatically demonstrate clinical importance, educational importance, causation, or statistical significance. It should be interpreted with the raw mean difference, confidence interval, sample size, study design, measurement quality, previous research, and the importance of the outcome (Lakens, 2013).
What Does Cohen’s d Measure?
Cohen’s d measures the separation between means in standard-deviation units. It has two components:
- The numerator, which is the observed mean difference.
- The denominator, which is the standard deviation selected to represent relevant variability.
Suppose a nurse-led patient-education group has a mean medication-knowledge score of 82, while a standard-education group has a mean score of 77. The raw difference is 5 points. If the selected standard deviation is 10 points:
[
d=\frac{82-77}{10}=0.50
]
The nurse-led education mean is one-half of a standard deviation higher.
Why standardization is useful
Standardization removes the original measurement unit. This can help researchers compare differences from studies that use different instruments.
For example:
- Study A reports a five-point difference on a 100-point knowledge test.
- Study B reports a two-point difference on a 20-point competence scale.
The raw differences cannot be compared directly. Standardized effects may provide a common scale when each study uses a defensible denominator.
However, standardization does not make instruments interchangeable. Studies may still differ in:
- Populations
- Outcome definitions
- Instrument reliability
- Score distributions
- Intervention intensity
- Follow-up periods
- Clinical settings
- Study quality
The denominator defines the effect
An estimand is the exact population quantity a researcher intends to estimate. For Cohen’s d, the intended estimand may be a difference standardized by:
- Pooled within-group variability
- Control-group variability
- Pre-intervention variability
- Difference-score variability
- Average pre-test and post-test variability
Changing the denominator changes the question being answered. Therefore, selecting a Cohen’s d formula is not merely a computational decision. It is a decision about how the effect should be defined and interpreted (Morris & DeShon, 2002).
What Cohen’s d Does Not Tell You
Cohen’s d does not independently establish:
- Whether the difference is statistically significant
- Whether the result is clinically meaningful
- Whether an intervention caused the difference
- Whether the study is free from bias
- Whether a measure is reliable or valid
- Whether a result will generalize to another population
- Whether the estimate is precise
- Whether the intervention is safe or feasible
Cohen’s d is a point estimate. A strong interpretation also considers:
- Means and standard deviations
- The raw mean difference
- Sample sizes
- A confidence interval
- Group or time-point order
- The study design
- Measurement reliability
- Missing data
- Outliers and distribution shape
- Prior evidence
- Clinical or educational context
Measurement error can attenuate standardized effects, while unusually restricted variability can sometimes make a standardized difference appear larger. The effect should therefore be interpreted in relation to the measurement process, not as a property of the intervention alone (Hedges, 1981).
When Should Nursing Researchers Use Cohen’s d?
Cohen’s d is commonly considered when the research question concerns a difference involving a continuous or approximately continuous outcome.
Relevant nursing examples include:
- Medication-adherence scores
- Nursing-knowledge scores
- Clinical-competence ratings
- Patient-satisfaction scores
- Pain scores
- Burnout scores
- Quality-of-life scores
- Simulation-performance scores
- Clinical reasoning scores
- Time-based outcomes that can be meaningfully summarized using means and standard deviations
Design-based selection table
| Research design | Nursing example | Possible effect size | Main requirement |
|---|---|---|---|
| Two independent groups | Nurse-led education versus standard education | Independent-samples Cohen’s d | Choose a defensible standardizer |
| Same participants measured twice | Knowledge before and after training | (d_z), (d_{av}), or (d_{rm}) | Name the paired variant |
| Matched observations | Matched intervention and comparison patients | Paired standardized difference | Preserve the matching |
| Small independent groups | Two small groups of nursing students | Hedges’ g | Apply a bias correction |
| Control variability is the reference | Intervention may alter variability | Glass’s delta | Justify the reference-group SD |
| One sample versus a reference | Mean competence score versus a prespecified standard | One-sample Cohen’s d | Use a meaningful reference value |
A practical decision sequence
Before calculating Cohen’s d, ask:
- Is the outcome suitable for mean-based interpretation?
- Does the question concern a difference between means?
- Are the observations independent, paired, or matched?
- What variability should define one standard-deviation unit?
- Are the sample sizes small enough for bias correction to matter?
- Can a confidence interval be calculated using a validated method?
- Can the standardized result be interpreted alongside the raw difference?
When Cohen’s d Is Not Appropriate
Cohen’s d is generally not the primary effect size for:
- Two categorical variables
- Binary outcomes such as mortality versus survival
- Risks, odds, or event rates
- Correlations between variables
- Overall effects involving three or more means in ANOVA
- Ordinal outcomes that do not support mean-based interpretation
- Time-to-event outcomes
- Count outcomes with unsuitable distributions
- Mann–Whitney U or Wilcoxon results without a justified effect-size method
- Severely skewed or contaminated data for which the mean and SD are misleading
Possible alternatives include:
- Risk ratios
- Odds ratios
- Cramér’s V
- Correlation coefficients
- Eta squared
- Regression coefficients
- Rank-biserial correlations
- Robust standardized differences
A standardized mean difference can still be calculated when data are nonnormal, but calculation does not guarantee meaningful interpretation. When outliers, heavy tails, or strong skew make the mean and SD poor summaries, robust location and variability measures may provide a better representation (Algina et al., 2006).
Cohen’s d Formula for Independent Samples
Independent-samples Cohen’s d is used when the observations in one group are unrelated to those in the other group.
Examples include:
- Intervention and control patients
- Nurses from two unrelated hospital units
- Two separate cohorts of nursing students
- Patients assigned to different education methods
A common independent-samples calculation is:
[
d_s=
\frac{M_1-M_2}{SD_{\text{pooled}}}
]
The subscript (s) can be used to show that the effect was standardized using a sample pooled SD.
The pooled standard deviation is:
[
SD_{\text{pooled}}
\sqrt{
\frac{
(n_1-1)SD_1^2+(n_2-1)SD_2^2
}{
n_1+n_2-2
}
}
]
where:
- (M_1) is the mean of Group 1
- (M_2) is the mean of Group 2
- (SD_1) is the standard deviation of Group 1
- (SD_2) is the standard deviation of Group 2
- (n_1) is the Group 1 sample size
- (n_2) is the Group 2 sample size
The formula weights each group variance by its degrees of freedom. It is therefore preferable to simply averaging the two SDs when the sample sizes differ (Lakens, 2013).
Why the unweighted average of two SDs is not the pooled SD
The following calculation:
[
\frac{SD_1+SD_2}{2}
]
gives both groups equal influence, even when one group contains substantially more participants. The pooled SD instead gives more weight to the group that provides more information about within-group variability.
Does every independent t test require pooled Cohen’s d?
No. Pooled Cohen’s d corresponds to a common within-group variance model. It should not be selected automatically simply because a researcher conducted a t test.
The decision should consider:
- Whether equal population variances are plausible
- Whether the intervention could change variability
- Whether one group provides the natural reference distribution
- Whether Student’s or Welch’s t test was used
- Which standardized difference matches the research objective
Welch’s t test does not assume equal population variances. When Welch’s test is selected because equal variances are not defensible, pairing it automatically with a pooled-SD effect can create a mismatch between the hypothesis test and effect-size definition. A nonpooled standardizer may be more appropriate, but its formula must be stated clearly (Delacre et al., 2021).
Unequal SDs do not automatically mean that Glass’s delta must be used. The denominator should follow the intended estimand rather than a mechanical variance test.
Worked Example: Independent-Samples Cohen’s d
The following values are hypothetical.
A nursing researcher compares medication-knowledge scores for patients receiving nurse-led education and patients receiving standard education.
| Group | (n) | Mean | SD |
|---|---|---|---|
| Nurse-led education | 60 | 82.40 | 8.60 |
| Standard education | 58 | 77.80 | 9.10 |
Step 1: Calculate the raw mean difference
The nurse-led group is entered first:
[
M_1-M_2=82.40-77.80=4.60
]
The nurse-led group scored 4.60 points higher.
Step 2: Calculate the pooled standard deviation
[
SD_{\text{pooled}}
\sqrt{
\frac{
(60-1)(8.60)^2+(58-1)(9.10)^2
}{
60+58-2
}
}
]
[
SD_{\text{pooled}}
\sqrt{
\frac{
59(73.96)+57(82.81)
}{
116
}
}
]
[
SD_{\text{pooled}}=8.849
]
Step 3: Calculate Cohen’s d
[
d_s=
\frac{82.40-77.80}{8.849}
]
[
d_s=0.520
]
Rounded to two decimal places:
[
d_s=0.52
]
Step 4: Verify the t statistic
For the equal-variance independent-samples test:
[
t=
\frac{M_1-M_2}
{
SD_{\text{pooled}}
\sqrt{
\frac{1}{n_1}+\frac{1}{n_2}
}
}
]
[
t=
\frac{4.60}
{
8.849
\sqrt{
\frac{1}{60}+\frac{1}{58}
}
}
]
[
t=2.823
]
Therefore:
[
t(116)=2.82,\quad p=.006
]
The 95% confidence interval for the raw mean difference is:
[
95%,CI=[1.37,\ 7.83]
]
The noncentral-t-based confidence interval for Cohen’s d is:
[
95%,CI=[0.15,\ 0.89]
]
Interpretation
The nurse-led education mean was approximately 0.52 pooled standard deviations higher than the standard-education mean. This is conventionally described as a medium standardized difference.
The interval indicates that the data are compatible with effects ranging from approximately 0.15 to 0.89 standard deviations. Therefore, the estimate should not be interpreted as an exact population effect of 0.52.
The 4.60-point raw difference must also be interpreted using:
- The range of the knowledge scale
- Its reliability
- Any established competency threshold
- The educational consequences of the difference
- The intervention’s cost and burden
APA 7 reporting example
In this hypothetical study, patients who received nurse-led education obtained higher medication-knowledge scores (M = 82.40, SD = 8.60, n = 60) than patients who received standard education (M = 77.80, SD = 9.10, n = 58), mean difference = 4.60, 95% CI [1.37, 7.83], t(116) = 2.82, p = .006, Cohen’s (d_s=0.52), 95% CI [0.15, 0.89].
Cohen’s d for Paired Samples and Pre-Test/Post-Test Data
Paired data occur when observations are linked.
Common examples include:
- The same patients measured before and after education
- Nurses assessed before and after simulation training
- Burnout scores measured at baseline and follow-up
- Matched intervention and comparison patients
- Two conditions completed by the same participant
Paired observations must not be analyzed as if they came from unrelated groups. Each post-test observation is connected to a specific pre-test observation.
Why the pre-post correlation matters
For paired measurements, the standard deviation of the difference scores can be calculated as:
[
SD_D=
\sqrt{
SD_{\text{pre}}^2+
SD_{\text{post}}^2-
2r(SD_{\text{pre}})(SD_{\text{post}})
}
]
where:
- (SD_D) is the SD of the individual difference scores
- (r) is the correlation between paired measurements
A strong positive pre-post correlation generally reduces difference-score variability. That increases the precision of the estimated mean change and can increase (d_z).
The correlation does not necessarily make the intervention effect intrinsically larger. It changes the standardizer used to express the effect.
Why There Is More Than One Paired Cohen’s d
For independent groups, pooled within-group variability is a common standardizer. For paired data, researchers can standardize the mean change using several sources of variability.
The main options answer different questions:
- How large is the average change relative to variability in individual change?
- How large is the change relative to typical variability at each time point?
- How large is the change on a scale intended to resemble an independent-groups standardized difference?
- How large is the change relative to baseline variability?
Because these are different questions, there is no single paired Cohen’s d that is correct for every purpose (Morris & DeShon, 2002).
Cohen’s (d_z)
Cohen’s (d_z) divides the mean paired difference by the SD of the difference scores:
[
d_z=
\frac{\bar{D}}{SD_D}
]
where:
- (\bar{D}) is the mean of the paired difference scores
- (SD_D) is their standard deviation
It can also be calculated from the paired t statistic:
[
d_z=
\frac{t}{\sqrt{n}}
]
where (n) is the number of complete pairs.
This effect describes the average change relative to variation in individual change. It is directly connected to the paired-samples t test.
Because (SD_D) depends on the pre-post correlation, (d_z) is influenced by that correlation. This makes it useful for describing within-sample change and power, but it may complicate comparisons with effects standardized using between-person variability (Lakens, 2013).
Cohen’s (d_{av})
A commonly used formulation of (d_{av}) divides the mean change by the arithmetic average of the two time-point SDs:
[
d_{av}
\frac{
M_{\text{post}}-M_{\text{pre}}
}{
\frac{
SD_{\text{pre}}+SD_{\text{post}}
}{2}
}
]
This approach ignores the pre-post correlation in the denominator and expresses change relative to typical variability across the two measurements.
It may be easier to compare with independent-group standardized differences because it is not inflated or reduced directly by the correlation between repeated measurements (Lakens, 2013).
Average SD is not the same as average variance
The arithmetic-average SD is:
[
\frac{SD_{\text{pre}}+SD_{\text{post}}}{2}
]
A different denominator is the square root of the average variance:
[
\sqrt{
\frac{
SD_{\text{pre}}^2+SD_{\text{post}}^2
}{2}
}
]
These values may be close when the SDs are similar, but they are not mathematically identical.
IBM SPSS labels one paired effect-size option “Average of variances” and uses the square root of the average variance. Researchers should not assume this output is identical to the arithmetic-average (d_{av}) convention unless the denominator has been verified (IBM, n.d.-b).
Cohen’s (d_{rm})
Under one widely cited repeated-measures convention:
[
d_{rm}
d_z\sqrt{2(1-r)}
]
or:
[
d_{rm}
\frac{\bar{D}}{SD_D}
\sqrt{2(1-r)}
]
This adjustment is intended to place the repeated-measures effect on a scale resembling a between-subject standardized difference.
When both measurements have equal variance:
[
SD_D=SD\sqrt{2(1-r)}
]
and therefore:
[
d_{rm}
\frac{\bar{D}}{SD}
]
Under equal variances, (d_{rm}), (d_{av}), and a conventional standardized difference using the common time-point SD coincide.
When the time-point SDs differ, they do not necessarily coincide. The researcher should identify the exact formula and avoid treating (d_{rm}) as a universally defined adjustment (Lakens, 2013).
Baseline-standardized change
A pre-test/post-test effect can also be standardized using baseline variability:
[
d_{\text{pre}}
\frac{
M_{\text{post}}-M_{\text{pre}}
}{
SD_{\text{pre}}
}
]
This is sometimes called a baseline-standardized change or Becker-type effect. It may be appropriate when baseline variability represents the natural untreated reference.
The choice must be justified. Baseline standardization should not be selected solely because it produces a preferred value.
Paired-effect comparison table
| Effect size | Denominator | Primary interpretation | Main caution |
|---|---|---|---|
| (d_z) | SD of difference scores | Change relative to variability in individual change | Strongly influenced by paired correlation |
| (d_{av}) | Arithmetic average of time-point SDs | Change relative to typical time-point variability | State the exact averaging method |
| (d_{rm}) | Correlation-adjusted repeated-measures standardizer | Change on a between-subject-comparable scale | Requires a named convention and assumptions |
| (d_{\text{pre}}) | Pre-test SD | Change relative to baseline variability | Baseline must be a defensible reference |
For paired data, never report only “Cohen’s d.” State the variant, denominator, direction, and calculation method.
Worked Example: Paired-Samples Cohen’s d
The following values are hypothetical.
Forty-five nurses complete a clinical-deterioration knowledge test before and after training.
| Measurement | (n) | Mean | SD |
|---|---|---|---|
| Pre-test | 45 | 68.20 | 10.00 |
| Post-test | 45 | 75.80 | 9.20 |
The pre-post correlation is:
[
r=.60
]
Step 1: Calculate the mean change
Post-test minus pre-test is used:
[
\bar{D}=75.80-68.20=7.60
]
A positive effect therefore represents increased knowledge.
Step 2: Calculate the SD of the difference scores
[
SD_D=
\sqrt{
10.00^2+9.20^2-
2(.60)(10.00)(9.20)
}
]
[
SD_D=
\sqrt{
100+84.64-110.40
}
]
[
SD_D=
\sqrt{74.24}
]
[
SD_D=8.616
]
Step 3: Calculate (d_z)
[
d_z=
\frac{7.60}{8.616}
]
[
d_z=0.882
]
Rounded:
[
d_z=0.88
]
Step 4: Verify the paired t statistic
[
t=
\frac{\bar{D}}
{SD_D/\sqrt{n}}
]
[
t=
\frac{7.60}
{8.616/\sqrt{45}}
]
[
t=5.917
]
Therefore:
[
t(44)=5.92,\quad p<.001
]
The raw mean-change interval is:
[
95%,CI=[5.01,\ 10.19]
]
The noncentral-t-based interval for (d_z) is:
[
95%,CI=[0.53,\ 1.22]
]
Step 5: Calculate (d_{av})
Using the arithmetic average of the two SDs:
[
d_{av}
\frac{7.60}
{
(10.00+9.20)/2
}
]
[
d_{av}
\frac{7.60}{9.60}
]
[
d_{av}=0.792
]
Rounded:
[
d_{av}=0.79
]
Step 6: Calculate (d_{rm})
Using the stated Cohen–Lakens repeated-measures convention:
[
d_{rm}
0.882
\sqrt{
2(1-.60)
}
]
[
d_{rm}
0.882\sqrt{0.80}
]
[
d_{rm}=0.789
]
Rounded:
[
d_{rm}=0.79
]
Why the values differ
The result is (d_z=0.88), while (d_{av}) and (d_{rm}) are approximately 0.79.
The values differ because:
- (d_z) uses the SD of individual change scores.
- (d_{av}) uses average time-point variability.
- (d_{rm}) adjusts (d_z) for dependence under a specified convention.
None of the values should be presented without its label. Selecting among them depends on the intended interpretation and comparison.
APA 7 reporting example
Knowledge scores increased from pre-test (M = 68.20, SD = 10.00) to post-test (M = 75.80, SD = 9.20), mean change = 7.60, 95% CI [5.01, 10.19], t(44) = 5.92, p < .001, Cohen’s (d_z=0.88), 95% CI [0.53, 1.22]. The effect was standardized using the SD of the paired difference scores.
Cohen’s d for One-Sample Studies
One-sample Cohen’s d compares a sample mean with a fixed reference value:
[
d=
\frac{
M-\mu_0
}{
SD
}
]
where:
- (M) is the sample mean
- (\mu_0) is the reference value
- (SD) is the sample SD
Hypothetical nursing example
Suppose an established nursing-competence standard is 72. A sample of 40 nurses has:
[
M=78.40,\quad SD=9.60
]
Then:
[
d=
\frac{78.40-72}{9.60}
]
[
d=0.67
]
The sample mean is approximately 0.67 sample standard deviations above the reference value.
The reference value must be meaningful and selected before examining the results. It might represent:
- A validated competency threshold
- A prespecified historical standard
- A theoretically justified neutral value
- An externally established population value
Researchers should not select a convenient reference after viewing the sample mean.
Cohen’s d vs Hedges’ g
Sample Cohen’s d has a small positive bias as an estimator of the corresponding population standardized difference. The bias is most relevant in small samples.
Hedges’ g applies a correction factor:
[
g=J(df)d
]
The exact correction is:
[
J(df)
\frac{
\Gamma(df/2)
}{
\sqrt{df/2},
\Gamma((df-1)/2)
}
]
where (\Gamma) is the gamma function.
A commonly used approximation is:
[
J(df)
\approx
1-\frac{3}{4df-1}
]
For two independent groups:
[
df=n_1+n_2-2
]
The approximate correction is extremely close to the exact value in many practical applications, but validated software may use the exact gamma-function calculation (Hedges, 1981).
Hypothetical small-sample calculation
Suppose:
[
n_1=12,\quad n_2=11,\quad d=0.824
]
Then:
[
df=12+11-2=21
]
The exact correction factor is approximately:
[
J(21)=0.964
]
Therefore:
[
g=(0.964)(0.824)
]
[
g=0.794
]
Rounded:
[
g=0.79
]
Cohen’s d and Hedges’ g comparison
| Feature | Cohen’s d | Hedges’ g |
|---|---|---|
| Purpose | Standardized mean difference | Bias-corrected standardized mean difference |
| Small-sample correction | No | Yes |
| Difference in large samples | Usually minimal | Approaches Cohen’s d |
| Common use | Primary studies and general reporting | Small samples and meta-analysis |
| Interpretation | Standard-deviation units | Standard-deviation units |
Hedges’ g does not correct:
- Low statistical power
- Confounding
- Selection bias
- Measurement error
- Missing data
- Poor intervention fidelity
- Wide confidence intervals
- An inappropriate effect-size denominator
It is also not selected merely because the p value is nonsignificant.
Cohen’s d vs Glass’s Delta
Glass’s delta standardizes a mean difference using the SD of a designated reference group:
[
\Delta=
\frac{
M_{\text{intervention}}-
M_{\text{control}}
}{
SD_{\text{control}}
}
]
It may be useful when:
- The intervention could change outcome variability
- The control group represents the untreated distribution
- Pooling intervention and control variability would not match the research question
- Baseline variability provides the most defensible scale
For example, an education program may increase the average knowledge score while reducing variation among participants. Pooling post-intervention and control SDs would combine variability from distributions affected differently by the intervention.
Unequal SDs alone do not prove that Glass’s delta is correct. The control or reference group must provide a substantively meaningful denominator (Hedges, 1981).
Cohen’s d Effect Size Interpretation
The sign and absolute value of Cohen’s d convey different information.
A value of zero
[
d=0
]
The sample means are equal. This does not prove that the population effect is exactly zero.
Positive values
A positive value means that the first mean in the numerator is higher.
Negative values
A negative value means that the first mean is lower.
A negative effect is not automatically weak or harmful. For outcomes such as pain, burnout, falls, symptom severity, or medication errors, a lower intervention mean may represent improvement.
Absolute magnitude
The absolute value describes standardized separation:
[
|d|
]
For example, (d=-0.70) and (d=0.70) have the same absolute magnitude but opposite directions.
Conventional benchmarks
Cohen (1988) proposed widely used reference values.
Approximate conventional benchmarks—not universal clinical rules
| Absolute Cohen’s d | Conventional description |
|---|---|
| 0.20 | Small |
| 0.50 | Medium |
| 0.80 | Large |
These values are not natural boundaries. A result of 0.49 is not meaningfully different from 0.50 merely because one falls below a category label.
Conventional labels should be secondary to:
- Previous effects in comparable studies
- Raw-unit differences
- Confidence intervals
- Instrument reliability
- Outcome severity
- Patient relevance
- Intervention cost
- Feasibility
- Safety
- Duration of benefit
A small average effect may matter when an outcome is common or serious and the intervention is safe and inexpensive. A large standardized effect may have limited value when the outcome is unimportant, the sample is highly selected, or the interval is extremely wide (Lakens, 2013).
Can Cohen’s d Be Greater Than 1?
Cohen’s d is not restricted to values between −1 and +1.
A value of:
[
d=1.30
]
means that the means differ by 1.30 times the standard deviation used in the denominator.
A value greater than 1 can be valid. However, an unusually large result should prompt checks for:
- Data-entry errors
- Incorrect coding
- Extremely small SDs
- Restricted samples
- Outliers
- Small-sample instability
- Ceiling or floor effects
- Noncomparable groups
- Use of an SE instead of an SD
- An incorrect denominator
A large point estimate should be interpreted cautiously when its confidence interval is wide.
Why the Sign of Cohen’s d Matters
Consider:
[
d=
\frac{
M_{\text{intervention}}-
M_{\text{control}}
}{
SD_{\text{pooled}}
}
]
Under this order:
- (d>0) means the intervention mean is higher.
- (d<0) means the intervention mean is lower.
Reversing the order reverses the sign:
[
\frac{
M_{\text{control}}-
M_{\text{intervention}}
}{
SD_{\text{pooled}}
}
=-d
]
The absolute magnitude is unchanged.
A complete report should state:
- Which mean was subtracted from which
- Which group or time point scored higher
- Whether a higher score indicates improvement or worsening
Reporting only the absolute value removes information needed to interpret the result.
Cohen’s d and P-Values
A p value and Cohen’s d answer different questions.
A p value evaluates how compatible the observed result is with a specified null model. It does not measure the size or importance of the difference.
Cohen’s d estimates standardized magnitude. It does not determine statistical significance by itself.
Sample size affects the relationship:
- A large sample may produce a small p value for a small effect.
- A small sample may produce a nonsignificant result for a moderate point estimate.
- Large samples usually produce narrower intervals.
- Small samples often produce unstable effects and wide intervals.
Both significant and nonsignificant results should include effect sizes and confidence intervals when appropriate (Sullivan & Feinn, 2012).
For a detailed explanation of statistical significance, see P-Values in Nursing Research.
Cohen’s d and Clinical Significance
Statistical magnitude and clinical importance are not interchangeable.
A clinical interpretation should consider:
- The raw mean difference
- The original measurement unit
- The confidence interval
- A validated minimally important difference
- Patient safety
- Adverse effects
- Intervention burden
- Costs
- Staffing requirements
- Feasibility
- Patient preferences
- Duration of benefit
- Generalizability
A Cohen’s (d=0.50) does not automatically indicate a clinically meaningful intervention.
For example, a standardized difference of 0.50 on a knowledge scale might correspond to:
- Two additional correct answers
- Ten additional correct answers
- Crossing a competency threshold
- No meaningful change in clinical behavior
The meaning depends on the instrument and application.
A minimally important difference should be used only when validated for the relevant:
- Instrument
- Population
- Condition
- Outcome
- Follow-up period
Do not borrow an MCID from an unrelated population simply because the same scale was used.
Confidence Intervals for Cohen’s d
A point estimate alone is incomplete. A confidence interval communicates the precision and directional uncertainty of the effect.
How to read the interval
A narrow interval suggests relatively high precision. A wide interval indicates that the data remain compatible with a broad range of effects.
An interval including zero does not prove the true effect is zero. It indicates that zero remains one of the values reasonably compatible with the data and model.
Three hypothetical interpretations
Relatively precise moderate effect
[
d=0.52,\quad 95%,CI=[0.35,\ 0.69]
]
The data support a comparatively narrow range of positive effects.
Imprecise large point estimate
[
d=0.82,\quad 95%,CI=[-0.04,\ 1.67]
]
The point estimate is conventionally large, but the interval includes effects from almost zero in the opposite direction to very large positive effects.
Nonsignificant paired effect
[
d_z=0.34,\quad 95%,CI=[-0.21,\ 0.87]
]
The data do not establish a null effect. They remain compatible with negative, negligible, moderate, and potentially large positive effects.
Noncentral-t confidence intervals
A validated method estimates confidence limits for the noncentrality parameter of the relevant t distribution and transforms those limits to the Cohen’s d scale.
For an equal-variance independent-samples effect:
[
d=
\lambda
\sqrt{
\frac{1}{n_1}+
\frac{1}{n_2}
}
]
where (\lambda) is the noncentrality parameter.
For paired (d_z):
[
d_z=
\frac{\lambda}{\sqrt{n}}
]
This method accounts for the relationship among the observed t statistic, degrees of freedom, sample size, and standardized effect. It is generally preferable to unsupported shortcut formulas (Algina et al., 2006).
Reproducibility note for this article
All hypothetical confidence intervals in this article were calculated using unrounded input values.
- Raw mean-difference intervals use the central t distribution.
- Cohen’s d and (d_z) intervals use inversion of the noncentral t distribution.
- The Hedges’ g interval is obtained by applying the exact (J(df)) correction to the corresponding Cohen’s d limits.
Different validated software implementations may differ slightly because of rounding, correction conventions, or CI algorithms. Researchers should report the software, version, function, and options used.
How Sample Size Affects Cohen’s d
Sample size is not written explicitly in the basic formula:
[
d=
\frac{\text{mean difference}}{SD}
]
However, sample size still affects the estimate.
Small samples produce less stable estimates of:
- Means
- Standard deviations
- Correlations
- Difference-score variability
- Confidence limits
A large observed effect from a tiny sample may result partly from sampling fluctuation. Small-sample standardized differences are also positively biased on average, which is why Hedges’ correction can be useful (Hedges, 1981).
A large sample can estimate a small effect precisely. Precision does not make the effect clinically important, but it may allow researchers to exclude larger effects.
The interpretation should distinguish:
- Magnitude: the point estimate
- Precision: the interval width
- Evidence against a null model: the p value
- Clinical importance: the effect in context
Calculating Cohen’s d From a t-Test
Independent-samples Student t test
For an equal-variance independent-samples t test:
[
d_s=
t
\sqrt{
\frac{1}{n_1}+
\frac{1}{n_2}
}
]
This formula produces pooled-SD Cohen’s (d_s).
It assumes:
- Independent observations
- A Student equal-variance t statistic
- Correct analyzed sample sizes
- A pooled-SD estimand
Do not automatically apply it to a Welch t statistic.
Paired-samples t test
For paired observations:
[
d_z=
\frac{t}{\sqrt{n}}
]
This produces (d_z), not (d_{av}) or (d_{rm}).
Additional information is needed for the other variants:
- Pre-test SD
- Post-test SD
- Difference-score SD
- Paired correlation
- Exact standardizer definition
One-sample t test
For a one-sample test:
[
d=
\frac{t}{\sqrt{n}}
]
This resembles the paired formula because a paired t test is mathematically a one-sample test of the individual difference scores.
How to Calculate Cohen’s d in SPSS
IBM introduced the ES subcommand for T-TEST in SPSS Statistics release 27. Current documentation shows effect-size estimation for one-sample, independent-samples, and paired-samples tests (IBM, n.d.-a).
Before running the analysis
Check:
- The group codes
- The group order
- Whether observations are independent or paired
- The analyzed sample sizes
- Missing-data handling
- Whether higher scores represent improvement or worsening
- Whether Student’s or Welch’s test is appropriate
Independent-samples SPSS output
SPSS can display:
- Group sample sizes
- Means
- SDs
- Mean differences
- Confidence intervals
- Student and Welch test results
- Effect-size estimates
Do not assume that the same denominator is appropriate for both the pooled and unequal-variance test rows. Record which test and effect-size definition you report.
Paired-samples SPSS standardizers
IBM documentation identifies three paired options:
- Standard deviation of the difference
- Corrected standard deviation of the difference
- Average of variances
The default difference-score option corresponds conceptually to (d_z). The average-of-variances option uses:
[
\sqrt{
\frac{
SD_1^2+SD_2^2
}{2}
}
]
This is not identical to the arithmetic-average-SD formula often labeled (d_{av}) (IBM, n.d.-b).
Record:
- SPSS version
- Procedure
- Test type
- Standardizer option
- Group or pair order
- Confidence level
- Whether Cohen’s d or Hedges’ correction was reported
Need help calculating Cohen’s d from your SPSS output? Upload your dataset, output, research questions, and rubric through our SPSS Data Analysis Help page for support with effect-size selection, calculation, confidence intervals, and APA reporting.
Cohen’s d in Excel, Jamovi, JASP, and R
| Software | Possible approach | Essential check |
|---|---|---|
| Excel | Enter a transparent formula using means, SDs, and sample sizes | Confirm the design and denominator |
| Jamovi | Request the effect size and its confidence interval in the relevant t test | Verify the effect definition and group order |
| JASP | Select effect-size and interval options in the t-test analysis | Check whether the reported method matches the design |
| R | Use a validated package or explicitly coded formula | Record package version, function, and arguments |
Jamovi’s independent-samples procedure can request an effect size and an effect-size confidence interval, but users should still verify the analysis type and output definition (The jamovi project, n.d.).
In R, the effectsize package provides independent, one-sample, paired, repeated-measures, Hedges’ g, and Glass’s delta options. It also documents noncentrality-based intervals for many standardized differences (Ben-Shachar et al., 2020).
Software packages may use different defaults for paired data. Never compare two outputs based only on the label “Cohen’s d.” Compare their denominators and correction methods.
How to Report Cohen’s d in APA 7
A strong quantitative result should allow the reader to understand the design, direction, magnitude, and uncertainty of the effect.
APA quantitative-reporting guidance supports reporting effect estimates and confidence intervals rather than relying solely on threshold-based significance conclusions (Appelbaum et al., 2018).
Include:
- Test name
- Group or time-point order
- Means
- SDs
- Sample sizes
- Raw mean difference
- Confidence interval for the raw difference
- Test statistic
- Degrees of freedom
- Exact p value unless p < .001
- Named effect-size variant
- Effect-size value
- Effect-size confidence interval
- Interpretation of direction
- Contextual interpretation
All examples below are hypothetical.
1. Significant independent-samples result
Patients receiving nurse-led education had higher knowledge scores (M = 82.40, SD = 8.60, n = 60) than patients receiving standard education (M = 77.80, SD = 9.10, n = 58), mean difference = 4.60, 95% CI [1.37, 7.83], t(116) = 2.82, p = .006, Cohen’s (d_s=0.52), 95% CI [0.15, 0.89].
2. Nonsignificant independent-samples result
The intervention group (M = 71.00, SD = 10.20, n = 25) did not differ statistically significantly from the comparison group (M = 68.50, SD = 9.80, n = 24), mean difference = 2.50, 95% CI [−3.25, 8.25], t(47) = 0.87, p = .386, Cohen’s (d_s=0.25), 95% CI [−0.31, 0.81]. The interval included negative, negligible, moderate, and potentially large positive effects.
3. Significant paired-samples result
Knowledge scores increased from pre-test (M = 68.20, SD = 10.00) to post-test (M = 75.80, SD = 9.20), mean change = 7.60, 95% CI [5.01, 10.19], t(44) = 5.92, p < .001, Cohen’s (d_z=0.88), 95% CI [0.53, 1.22]. The effect was standardized using the SD of the difference scores.
4. Nonsignificant paired result with a wide interval
In a sample of 14 nurses, scores increased from pre-test (M = 62.00, SD = 11.00) to post-test (M = 66.00, SD = 10.50), mean change = 4.00, 95% CI [−2.80, 10.80], t(13) = 1.27, p = .226, Cohen’s (d_z=0.34), 95% CI [−0.21, 0.87]. The pre-post correlation was .40, and the interval indicated substantial uncertainty.
5. Hedges’ g in a small sample
The intervention group (M = 79.00, SD = 8.00, n = 12) had higher scores than the comparison group (M = 72.00, SD = 9.00, n = 11), mean difference = 7.00, 95% CI [−0.37, 14.37], t(21) = 1.98, p = .062, Hedges’ (g_s=0.79), 95% CI [−0.04, 1.61]. Although the point estimate was conventionally large, its confidence interval was highly uncertain.
APA numerical conventions
Use a leading zero for statistics that can exceed 1:
- Correct: (d=0.54)
- Incorrect: (d=.54)
Do not report:
[
p=.000
]
Report:
[
p<.001
]
Use spaces around operators in narrative statistical reporting:
[
d=0.54,\quad p=.021
]
When formatted in prose:
Cohen’s d = 0.54, p = .021.
Common Cohen’s d Mistakes
Using an independent formula for paired data
Pre-test and post-test scores from the same participants are dependent. Treating them as unrelated discards the covariance between measurements.
Reporting paired Cohen’s d without naming the variant
A result reported only as “Cohen’s d = 0.70” cannot be interpreted or reproduced without its denominator.
Using the unweighted average of SDs as the pooled SD
The arithmetic average does not properly weight groups with different sample sizes.
Treating arithmetic-average SD and average variance as identical
[
\frac{SD_1+SD_2}{2}
]
is not identical to:
[
\sqrt{
\frac{
SD_1^2+SD_2^2
}{2}
}
]
Ignoring group or time-point order
The sign cannot be interpreted without knowing which mean was entered first.
Reporting only the absolute effect
Removing the sign conceals direction.
Assuming a negative value means no effect
A negative effect can be large and beneficial when lower outcomes are preferable.
Treating benchmarks as rigid categories
The values 0.20, 0.50, and 0.80 are approximate conventions, not clinical thresholds.
Claiming that a medium effect is clinically meaningful
Clinical meaning requires outcome-specific interpretation.
Omitting the raw mean difference
Readers need the result in its original measurement unit.
Omitting the confidence interval
A point estimate cannot communicate precision.
Confusing SD and SE
Cohen’s d normally uses an SD-based denominator, not the standard error.
Applying the independent t-to-d formula to Welch or paired data
Each design and test requires the appropriate conversion.
Assuming Hedges’ g fixes a weak study
The correction addresses small-sample bias in the standardized estimate, not broader design problems.
Comparing (d_z) directly with pooled independent d
The values use different standardizers.
Relying on software labels
Always document the calculation method and denominator.
Using inconsistent hypothetical numbers
Means, SDs, sample sizes, t values, p values, intervals, and effect sizes must be mathematically compatible.
Cohen’s d Calculation and Reporting Checklist
Before submitting your nursing dissertation, confirm:
- The outcome supports mean-based interpretation.
- The research design is correctly identified.
- Independent observations are not treated as paired.
- Paired observations are not treated as independent.
- Group or time-point order is documented.
- Means are reported.
- SDs are reported.
- Analyzed sample sizes are reported.
- The raw mean difference is reported.
- The paired difference-score SD is available where needed.
- The paired correlation is reported where required.
- The Cohen’s d variant is named.
- The denominator is defined.
- The unequal-variance issue has been considered.
- Small-sample correction has been considered.
- A validated confidence interval is included.
- Conventional labels are used cautiously.
- Clinical or educational meaning is discussed separately.
- APA statistical spacing is consistent.
- Software and version are recorded.
- The function, procedure, or formula is documented.
- All numerical results have been independently checked.
When to Seek Statistical Support
Additional statistical review may be useful when:
- Group independence is unclear
- Participants have repeated measurements
- Matching was used
- The paired variant is uncertain
- Only a t statistic and p value are available
- Sample sizes are small
- Group SDs differ substantially
- Welch’s test was used
- The intervention may have changed variability
- SPSS reports several standardizers
- Two software packages produce different results
- A noncentral-t confidence interval is required
- Hedges’ g has been requested
- The direction of the effect is confusing
- Chapter 4 requires APA 7 revision
Conclusion
Correct Cohen’s d effect size reporting begins with the research design and intended standardizer.
For two independent groups, pooled Cohen’s (d_s) may be appropriate when pooled within-group variability matches the intended estimand. It should not be used automatically when population variances differ or when another reference distribution is more meaningful.
For paired and pre-test/post-test studies, researchers must identify the reported variant. (d_z), (d_{av}), (d_{rm}), and baseline-standardized change answer related but different questions. Reporting only “Cohen’s d” hides the denominator and prevents meaningful comparison.
A complete nursing research interpretation should include:
- Design
- Group or time-point order
- Means and SDs
- Sample sizes
- Raw mean difference
- Named standardized effect
- Confidence interval
- Small-sample correction where appropriate
- Contextual interpretation
- A separate assessment of clinical significance
Conventional small, medium, and large labels can provide orientation, but they cannot determine whether an intervention is clinically valuable, safe, feasible, or worth implementing.
Get your Cohen’s d calculation reviewed before submission. Share your dataset, statistical output, research questions, methodology, rubric, and supervisor feedback through our Dissertation Data Analysis Help service for focused nursing research support.
References
Algina, J., Keselman, H. J., & Penfield, R. D. (2006). Confidence interval coverage for Cohen’s effect size statistic. Educational and Psychological Measurement, 66(6), 945–960. https://doi.org/10.1177/0013164406288161
American Psychological Association. (2020). Publication manual of the American Psychological Association (7th ed.). https://www.apa.org/pubs/books/publication-manual-7th-edition-paperback
Appelbaum, M., Cooper, H., Kline, R. B., Mayo-Wilson, E., Nezu, A. M., & Rao, S. M. (2018). Journal article reporting standards for quantitative research in psychology: The APA Publications and Communications Board task force report. American Psychologist, 73(1), 3–25. https://doi.org/10.1037/amp0000191
Ben-Shachar, M. S., Lüdecke, D., & Makowski, D. (2020). effectsize: Estimation of effect size indices and standardized parameters. Journal of Open Source Software, 5(56), Article 2815. https://doi.org/10.21105/joss.02815
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates. https://www.routledge.com/Statistical-Power-Analysis-for-the-Behavioral-Sciences/Cohen/p/book/9780805802832
Cumming, G. (2014). The new statistics: Why and how. Psychological Science, 25(1), 7–29. https://doi.org/10.1177/0956797613504966
Delacre, M., Lakens, D., Ley, C., Liu, L., & Leys, C. (2021). Why Hedges’ (g_s^) based on the non-pooled standard deviation should be reported with Welch’s t-test* [Preprint]. PsyArXiv. https://doi.org/10.31234/osf.io/tu6mp
Hedges, L. V. (1981). Distribution theory for Glass’s estimator of effect size and related estimators. Journal of Educational Statistics, 6(2), 107–128. https://doi.org/10.3102/10769986006002107
IBM. (n.d.-a). T-TEST. IBM SPSS Statistics documentation. Retrieved July 28, 2026, from https://www.ibm.com/docs/en/spss-statistics/32.0.0?topic=reference-t-test
IBM. (n.d.-b). Paired-samples t test. IBM SPSS Statistics documentation. Retrieved July 28, 2026, from https://www.ibm.com/docs/en/spss-statistics/30.0.0?topic=tests-paired-samples-t-test
Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t tests and ANOVAs. Frontiers in Psychology, 4, Article 863. https://doi.org/10.3389/fpsyg.2013.00863
Morris, S. B., & DeShon, R. P. (2002). Combining effect size estimates in meta-analysis with repeated measures and independent-groups designs. Psychological Methods, 7(1), 105–125. https://doi.org/10.1037/1082-989X.7.1.105
Sullivan, G. M., & Feinn, R. (2012). Using effect size—or why the p value is not enough. Journal of Graduate Medical Education, 4(3), 279–282. https://doi.org/10.4300/JGME-D-12-00156.1
The jamovi project. (n.d.). Independent samples t-test. Retrieved July 28, 2026, from https://jamovi.readthedocs.io/si/latest/jmv/jmv_ttestIS/
FAQs
What does a Cohen’s d of 0.50 mean?
A Cohen’s (d=0.50) means that the two means differ by approximately one-half of the standard deviation used in the calculation. It is conventionally described as a medium effect, but its practical meaning depends on the raw difference, confidence interval, outcome, population, and study design.
Which Cohen’s d should I use for pre-test and post-test data?
Pre-test and post-test data are paired. Possible variants include (d_z), which uses the SD of difference scores; (d_{av}), which uses average time-point variability; and (d_{rm}), which applies a repeated-measures adjustment. Report the exact variant and denominator rather than writing only “Cohen’s d.”
Can Cohen’s d be negative?
Yes. The sign depends on the order of subtraction. A negative value means that the first mean was lower than the second. It does not indicate that the effect is weak or invalid.
Should Cohen’s d be reported when the p-value is nonsignificant?
Yes, when a standardized mean difference is appropriate. The effect estimate and its confidence interval show the magnitude and uncertainty of the result. A nonsignificant finding may still be compatible with meaningful positive or negative effects.
When should I use Hedges’ g instead of Cohen’s d?
Hedges’ g is useful when sample sizes are small or a bias-corrected standardized mean difference is required. It reduces the small-sample upward bias of Cohen’s d but does not correct low power, confounding, unreliable measurement, or other design weaknesses.