Nursing students frequently collect questionnaire responses using 4-point, 5-point, or 7-point options, yet the analysis often becomes confusing once data collection ends. Should responses be treated as ordinal or continuous? Is a mean appropriate, or should the median be reported? Can several questions be combined into one score? Should the analysis use a t test, Mann–Whitney U test, analysis of variance, or ordinal regression?
These decisions affect more than the choice of statistical test. Incorrect coding, missed reverse-scored items, inappropriate scale construction, or unjustified category collapsing can distort scores and affect an entire results chapter. A statistically significant result cannot correct a questionnaire that was scored incorrectly.
Likert scale analysis therefore begins with the instrument—not SPSS. Researchers must understand what each item measures, whether items form validated scales or subscales, how missing responses are handled, and what higher scores represent. Only then should they select descriptive and inferential statistics.
This guide explains how to prepare, analyze, interpret, and report Likert-scale data in nursing research. It distinguishes individual Likert items from composite scales, presents practical SPSS guidance, provides an internally consistent nursing example, and demonstrates APA 7 reporting.
What Is a Likert Scale?
A Likert scale measures a respondent’s position toward a statement or question using ordered response categories. The categories move from a lower to a higher degree of agreement, satisfaction, confidence, frequency, or importance.
Common formats include:
- Four-point scale: Strongly disagree, disagree, agree, strongly agree
- Five-point scale: Strongly disagree to strongly agree
- Seven-point scale: Strongly disagree through a neutral midpoint to strongly agree
Four-point formats remove the midpoint and may encourage respondents to indicate a direction. Five-point formats are easy to understand and widely used. Seven-point formats provide greater differentiation but may increase response burden when distinctions between adjacent categories are unclear.
Nursing and healthcare questionnaires may use:
- Agreement: Strongly disagree to strongly agree
- Satisfaction: Very dissatisfied to very satisfied
- Confidence: Not confident to extremely confident
- Frequency: Never to always
- Importance: Not important to extremely important
For example, a nursing education questionnaire might ask students to rate “I feel confident performing a sterile dressing change.” A patient-experience survey might ask respondents to rate satisfaction with discharge education. A medication-adherence questionnaire could ask how frequently a patient forgets a prescribed dose.
Although these items use ordered response options, response categories alone do not make a questionnaire a validated scale. Validation requires evidence that the items measure the intended construct reliably and appropriately in the target population. Boateng et al. (2018) describe scale development as a process involving item development, assessment of dimensionality, reliability testing, and validity evaluation.
Likert Item vs. Likert Scale
The terms Likert item and Likert scale are often used interchangeably, but they describe different units of measurement.
A single Likert item is one question or statement with ordered response categories. A multi-item Likert scale contains several related items intended to measure the same construct. A scale may also contain separate subscales, each measuring a different dimension.
A total score is calculated by adding eligible item scores. An average composite score is calculated by finding their mean. Both approaches require a defensible scoring rule.
| Feature | Single Likert item | Multi-item Likert scale |
|---|---|---|
| Meaning | One response to one statement | Combined responses to related items measuring a construct |
| Number of questions | One | Usually two or more |
| Measurement level | Ordinal | Composite may be treated as approximately continuous when justified |
| Appropriate summaries | Frequencies, percentages, mode, median, IQR | Mean and SD may be used when construction and assumptions support them; median and IQR remain alternatives |
| Potential statistical tests | Mann–Whitney U, Wilcoxon, Kruskal–Wallis, Friedman, Spearman correlation, ordinal regression | t tests, ANOVA, Pearson correlation, or linear regression when justified; nonparametric alternatives otherwise |
| Nursing research example | “I am confident assessing a patient’s pain” | Average score across eight items measuring pain-assessment confidence |
A validated scale is more than a collection of similarly worded questions. Its proposed structure, scoring method, reliability, and validity should have been examined. If a validated instrument has separate emotional exhaustion, depersonalization, and professional accomplishment subscales, combining all items into one burnout score may contradict its intended structure.
Are Likert-Scale Data Ordinal or Continuous?
Individual Likert responses are ordinarily treated as ordinal. The categories have a meaningful order, but the distance between adjacent categories cannot automatically be assumed to be equal. The difference between “disagree” and “neutral” may not represent the same psychological distance as the difference between “agree” and “strongly agree.”
That does not mean means and parametric tests are always incorrect.
When several related items are combined, the resulting score has more possible values and may behave more like an approximately continuous variable. Treating such a composite score as continuous may be defensible when:
- The items measure the same construct or validated subscale.
- Reverse-scored items have been corrected.
- The scoring manual supports a total or mean score.
- The composite has enough possible values.
- Its distribution is not severely distorted by floor or ceiling effects.
- Extreme outliers are absent or appropriately addressed.
- Parametric model assumptions are adequately satisfied.
- The interpretation concerns the composite construct rather than one response category.
Harpe (2015) recommends distinguishing individual rating items from aggregated scales when choosing an analytical approach. Sullivan and Artino (2013) similarly explain that item-level and scale-level data may require different summaries.
Research also shows that some parametric methods are reasonably robust under departures from strict normality. Norman (2010) argues that parametric procedures can perform well with many Likert-derived scores, while de Winter and Dodou (2010) found that the independent-samples t test and Mann–Whitney procedure often had similar Type I error rates and power for simulated five-category items.
These findings do not create a universal rule. Instrument validation, the number of items, response distribution, sample size, research design, scoring instructions, and analytical objective still matter. The researcher should justify the decision rather than write that “Likert data are always ordinal” or “a five-point scale is automatically continuous.”
Step 1: Review the Questionnaire and Scoring Instructions
Before entering data, identify:
- Every individual item
- The construct measured by each item
- Any subscales
- Response options and their order
- Positively and negatively worded items
- Items requiring reverse scoring
- Missing-response rules
- The minimum number of completed items
- Instructions for calculating totals or averages
- Published cutoff values, if any
- What higher and lower scores mean
Create a scoring table showing the variable name, item wording, response codes, subscale, scoring direction, and missing-data rule.
Do not invent a scoring rule for an established instrument. If its manual instructs researchers to calculate three separate subscales, do not create one overall score merely because a single score is easier to analyze. Likewise, do not create “low,” “moderate,” and “high” categories unless the instrument provides validated cutoffs or the categories have a clear, prespecified justification.
Step 2: Code Likert-Scale Responses
Assign numerical values in the same order as the response categories.
| Response | Numerical code |
|---|---|
| Strongly disagree | 1 |
| Disagree | 2 |
| Neither agree nor disagree | 3 |
| Agree | 4 |
| Strongly agree | 5 |
Use clear variable names and retain full question wording in variable labels. In SPSS, value labels should indicate what each number represents.
Missing values require separate codes or system-missing entries. A response such as “not applicable” should not automatically receive the neutral score of 3. Neutral means the participant has a midpoint position; not applicable means the item may not apply to that person. Treating the two as equivalent can produce misleading scores.
Where possible, code items so higher values have a consistent substantive meaning. For example, higher values might always represent stronger confidence. If an established instrument intentionally uses mixed direction, retain its original data and create new reverse-scored variables for analysis.
Keep an unedited copy of the raw dataset. Detailed variable setup is covered in SPSS Data Entry for Nursing Research.
Step 3: Reverse-Code Negatively Worded Items
Negatively worded items must be aligned with the direction of the remaining scale before scores are combined.
Suppose higher scores should indicate greater confidence, but one item states, “I feel unsure when patients ask questions about their medication.” Agreeing with that statement represents lower confidence. Its score must therefore be reversed.
The general formula is:
Reversed score = highest possible response + lowest possible response − original score
For a 1-to-5 item:
| Original score | Reversed score |
|---|---|
| 1 | 5 |
| 2 | 4 |
| 3 | 3 |
| 4 | 2 |
| 5 | 1 |
Create a new variable rather than overwriting the original. Compare its frequencies with the original variable to verify that the transformation worked.
Incorrect or omitted reverse scoring may cause:
- Negative corrected item-total correlations
- Unexpectedly low Cronbach’s alpha
- Misleading composite scores
- Reversal of the intended interpretation
- Incorrect conclusions about group differences or relationships
A negative item-total correlation should prompt a coding check before the item is removed.
Step 4: Clean and Screen the Data
Data cleaning should occur before reliability testing or composite-score calculation.
Check for:
- Values outside the permitted response range
- Inconsistent numerical coding
- Duplicate records
- Missing responses
- Incomplete questionnaires
- Contradictory responses
- Unusual response sequences
- Straight-lining across many items
- Floor and ceiling effects
- Direct identifiers that should be removed or protected
Straight-lining means selecting the same response repeatedly. It can indicate low engagement, but it can also reflect a participant’s genuine position. Flag such records for review rather than deleting them automatically. Apply prespecified data-quality criteria consistently.
Floor effects occur when many participants select the lowest score; ceiling effects occur when responses cluster at the highest score. Severe effects reduce the instrument’s ability to distinguish respondents and may limit change detection.
Likert-scale data-cleaning checklist
- Preserve an untouched raw-data file.
- Confirm that all response codes fall within the permitted range.
- Separate neutral, missing, and not-applicable responses.
- Check every reverse-scored item.
- Review duplicate and incomplete records.
- Calculate missingness by item and participant.
- Inspect frequencies before computing composites.
- Review possible floor and ceiling effects.
- Document exclusions and corrections.
- Protect confidentiality and remove unnecessary identifiers.
Step 5: Assess Reliability Before Combining Items
Reliability concerns the consistency of scores produced by a measurement procedure. Cronbach’s alpha is commonly used to examine the internal consistency of items intended to measure the same construct.
Review:
- Cronbach’s alpha: A summary of interrelationships among the included items
- Corrected item-total correlation: The relationship between an item and the total formed from the remaining items
- Alpha if item deleted: An indication of how alpha changes when an item is removed
Alpha is affected by the number of items, their correlations, and assumptions about the measurement model. A high alpha does not prove that a scale is valid or unidimensional. A very high value may even indicate redundant items. Tavakol and Dennick (2011) explain these interpretive limitations, while McNeish (2018) cautions against relying on alpha automatically when its assumptions are unsuitable.
Analyze distinct subscales separately. Do not run alpha on unrelated questionnaire items simply because they share the same response options. Items should not be deleted solely to increase alpha; first examine item content, coding, reverse scoring, dimensionality, and the instrument’s validated structure.
The complete procedure is available in the SPSS Reliability Analysis guide.
Step 6: Calculate a Total or Average Scale Score
Combine items only when theory, instrument documentation, and available psychometric evidence support doing so.
A sum score adds eligible item values. An average score divides that total by the number of answered items. Sum and average scores preserve the same respondent ordering when everyone answers the same number of items, but their interpretations differ.
Consider a hypothetical five-item nursing confidence scale scored from 1 to 5. One nurse gives scores of 4, 5, 4, 2, and 4. Item 4 is negatively worded, so the original value of 2 becomes 4.
- Sum score: 4 + 5 + 4 + 4 + 4 = 21
- Average score: 21 ÷ 5 = 4.20
The average is easy to interpret because it remains on the original 1-to-5 metric. However, this calculation is valid only if the scale’s scoring rule permits it.
Missing responses complicate scoring. A researcher might require at least four of five items before calculating a participant’s average, but that rule should come from the instrument manual or an explicitly justified analysis plan. Unrelated items and separate subscales should not be combined.
Unsure how to score your questionnaire? Our Dissertation Data Analysis Help service can help you review the instrument, identify subscales, manage reverse-coded items, and prepare an appropriate analysis plan.
Step 7: Select Descriptive Statistics
For an individual Likert item, report frequencies and percentages because they preserve the response categories. The mode, median, and interquartile range may also be informative.
Hypothetical item-level results
Item: “I can use teach-back to confirm patient understanding”
| Response | n | % |
|---|---|---|
| Strongly disagree | 4 | 3.3 |
| Disagree | 12 | 10.0 |
| Neither agree nor disagree | 21 | 17.5 |
| Agree | 58 | 48.3 |
| Strongly agree | 25 | 20.8 |
| Total | 120 | 99.9 |
Note. Percentages do not total 100 because of rounding.
The median response was 4, corresponding to “agree,” with an interquartile range of 3–4. Overall, 69.1% selected agree or strongly agree. The table is more informative than reporting only a mean because it shows the distribution across categories.
Hypothetical composite-scale results
| Variable | n | M | SD | Median | IQR | Minimum | Maximum |
|---|---|---|---|---|---|---|---|
| Patient-education confidence | 120 | 3.82 | 0.61 | 3.80 | 3.40–4.20 | 2.20 | 5.00 |
The average confidence score was 3.82 on a 1-to-5 scale. The standard deviation of 0.61 indicates moderate dispersion around the mean. Because the score was formed from five related items and its analytical treatment was justified, the mean and standard deviation are reported alongside the median, IQR, and range.
Further guidance is available in Descriptive Data Analysis in Nursing Research.
Step 8: Choose the Correct Statistical Test
Test selection depends on the research question, measurement level, number of groups, independence of observations, sample size, distribution, outliers, and model assumptions.
| Research objective | Type of Likert variable | Study design | Potential test | Key consideration |
|---|---|---|---|---|
| Describe one item | Individual ordinal item | One sample | Frequencies, percentages, median, IQR | Preserve response categories |
| Compare two independent groups | Individual ordinal item | Independent groups | Mann–Whitney U | Ordinal outcome |
| Compare three or more independent groups | Individual ordinal item | Independent groups | Kruskal–Wallis | Ordinal outcome |
| Compare two related measurements | Individual ordinal item | Paired observations | Wilcoxon signed-rank | Repeated responses |
| Compare three or more related measurements | Individual ordinal item | Repeated observations | Friedman test | Repeated ordinal responses |
| Examine an ordinal relationship | Individual items or ranked variables | Association | Spearman correlation | Monotonic relationship |
| Compare suitable composite scores | Multi-item composite | Independent or paired groups | t test or ANOVA when justified | Check assumptions |
| Predict an ordinal outcome | Ordered response categories | Regression | Ordinal logistic regression | Check proportional-odds assumption |
| Predict a suitable composite score | Multi-item composite | Regression | Linear regression when justified | Check model assumptions |
Nonparametric approaches
The Mann–Whitney U test compares ranks between two independent groups. It should not automatically be described as a test of medians; that interpretation requires similarly shaped distributions.
The Kruskal–Wallis test extends rank-based comparison to three or more independent groups. A significant result indicates that at least one group differs, but adjusted post hoc comparisons are needed to identify where.
The Wilcoxon signed-rank test examines paired differences. It is suitable for two related ordinal measurements when its assumptions concerning the paired differences are reasonable.
The Friedman test compares three or more related ordinal measurements, such as confidence scores collected before training, immediately afterward, and at follow-up.
Spearman’s rank correlation evaluates the direction and strength of a monotonic relationship. It does not establish causation.
Parametric approaches
An independent-samples t test compares mean composite scores between two independent groups. A paired-samples t test compares two related mean scores. ANOVA compares mean scores across three or more groups.
Pearson correlation may be appropriate for two defensible continuous composite scores when the relationship is approximately linear and influential outliers are absent.
Linear regression may model a suitable composite outcome if linearity, independence, residual, variance, and influence assumptions are assessed. Ordinal logistic regression is often more appropriate when the outcome remains an ordered category. Its proportional-odds assumption must be examined.
For broader test-selection guidance, consult Statistical Tests in Nursing Research and When to Use Statistical Tests in Nursing Research. Additional explanations appear in Inferential Data Analysis in Nursing Research and the guide to Spearman Correlation Analysis in Nursing Research.
Should Researchers Combine Likert Response Categories?
Researchers sometimes combine strongly agree with agree and strongly disagree with disagree. This can simplify presentation or address very small category counts, but it also:
- Removes information
- Reduces variability
- May reduce statistical power
- Conceals differences between moderate and strong positions
- Can change the apparent pattern of results
Do not collapse categories after examining the results merely to obtain statistical significance. Prespecify and justify the decision whenever possible. Report the original categories in descriptive tables if they remain important, and explain exactly which categories were combined.
How to Conduct Likert Scale Analysis in SPSS
IBM SPSS Statistics supports data preparation, descriptive analysis, reliability assessment, nonparametric procedures, regression, and other relevant analyses.
A concise workflow is:
- Create and label each questionnaire variable.
- Define numerical response codes.
- Define missing values.
- Run frequencies to detect coding errors.
- Reverse-score applicable items into new variables.
- Assess reliability for each proposed scale or subscale.
- Calculate approved total or average scores.
- Produce appropriate descriptive statistics.
- Examine distributions, outliers, and assumptions.
- Select the statistical test matching the research question.
- Interpret effect sizes, confidence intervals, and p values.
- Prepare an APA-style table and results paragraph.
Students maintaining preliminary data in spreadsheets can consult Using Excel for Data Analysis in Nursing. Excel can assist with structured data entry and initial checks, but scoring rules and analyses still require careful validation.
Need help analyzing Likert-scale data in SPSS? Upload your questionnaire, dataset, codebook, and research questions through our SPSS Data Analysis Help page for support with coding, reliability testing, test selection, output interpretation, and APA reporting.
Factor Analysis and Likert-Scale Questionnaires
Factor analysis may be relevant when developing a new questionnaire, evaluating whether items reflect expected dimensions, or examining a scale in a substantially different population.
Reliability and factor analysis answer different questions:
- Reliability evaluates score consistency under specified assumptions.
- Factor analysis examines the pattern of relationships among items and the questionnaire’s potential dimensional structure.
Cronbach’s alpha does not prove that all items measure one factor. Questionnaire development also requires conceptual definition, content validation, appropriate item generation, structural evaluation, reliability evidence, and validity evidence. These stages are described by Boateng et al. (2018).
Established instruments should usually be scored according to their validated structure. Researchers evaluating structure can review SPSS Factor Analysis and How to Do Exploratory Factor Analysis in SPSS.
Worked Nursing Research Example
The following example and all numerical results are hypothetical.
Research question and hypothesis
Research question: Do nurses who completed structured patient-education training report higher confidence than nurses who did not complete the training?
Hypothesis: Nurses who completed the training will have a higher mean patient-education confidence score.
Sample and items
The sample contains 120 registered nurses:
- Training group: n = 62
- No-training group: n = 58
The hypothetical scale includes five items:
- I can explain the purpose of medication in plain language.
- I can use teach-back to assess patient understanding.
- I can adapt education to a patient’s health-literacy needs.
- I feel unsure when patients ask questions about their treatment.
- I document patient education accurately.
Responses range from 1 (strongly disagree) to 5 (strongly agree). Item 4 is reverse-scored. Higher composite scores indicate greater confidence.
Coding and missing-data rule
Responses are coded from 1 to 5. Item 4 is reversed using (6-\text{original score}). The participant’s mean score is calculated when at least four of the five items are complete. Otherwise, the composite score is missing.
Reliability and descriptive results
The five-item scale demonstrated acceptable internal consistency, Cronbach’s α = .84. Corrected item-total correlations were positive, and no item was removed.
Across all 120 nurses, the mean confidence score was 3.82 (SD = 0.61). The training group had a mean of 3.99 (SD = 0.55), while the no-training group had a mean of 3.64 (SD = 0.62).
Statistical test and findings
Because the outcome was a prespecified multi-item composite, observations were independent, and distributional checks supported mean comparison, an independent-samples t test was used.
The estimated mean difference was 0.35 points. A Welch independent-samples t test indicated that the training group reported higher confidence, t(113.30) = 3.26, p = .001, 95% CI [0.14, 0.56], Cohen’s d = 0.60.
In plain language, nurses who completed training reported moderately higher patient-education confidence. However, the observational comparison does not prove that training caused the difference. Other characteristics may distinguish the groups.
Statistical significance also does not automatically establish clinical or educational significance. Researchers should consider whether a 0.35-point difference is meaningful in practice, how confidence relates to observed competence, and whether similar results appear in other settings.
APA 7 results paragraph
The five-item Patient-Education Confidence Scale demonstrated acceptable internal consistency, Cronbach’s α = .84. Nurses who completed structured training (n = 62, M = 3.99, SD = 0.55) reported higher confidence than nurses who had not completed the training (n = 58, M = 3.64, SD = 0.62). A Welch independent-samples t test indicated that the 0.35-point mean difference was statistically significant, t(113.30) = 3.26, p = .001, 95% CI [0.14, 0.56], Cohen’s d = 0.60. The results suggest a moderate association between training completion and patient-education confidence; they do not establish a causal effect.
How to Present Likert-Scale Results
Use presentation methods that answer the research question:
- Frequency and percentage tables for individual items
- Median and IQR for ordinal distributions
- Mean and SD for justified composite scores
- Diverging stacked bar charts for multiple items
- Effect sizes and confidence intervals for group differences
- Focused APA-style results statements
Diverging stacked bars place negative categories on one side, positive categories on the other, and neutral responses near the center. They make response patterns across several items easier to compare.
Pie charts become difficult to interpret when several items or categories must be compared. Similarly, overcrowded tables can obscure the important result. Do not copy every SPSS output table into a dissertation. Select, reorganize, and label results so each table addresses a research question.
APA 7 Reporting Examples
All examples below are hypothetical. APA guidance recommends reporting exact p values unless p < .001 and encourages effect sizes and confidence intervals where appropriate (American Psychological Association, n.d.).
- Item-level result: Of the 120 nurses, 58 (48.3%) agreed and 25 (20.8%) strongly agreed that they could use teach-back effectively. The median response was 4 (IQR = 3–4).
- Composite result: Patient-education confidence scores averaged 3.82 (SD = 0.61, 95% CI [3.71, 3.93]).
- Reliability: The five-item confidence scale demonstrated acceptable internal consistency, Cronbach’s α = .84, 95% CI [.79, .88].
- Mann–Whitney result: Confidence ratings were higher in Group A (n = 45, median = 4, IQR = 3–4) than in Group B (n = 42, median = 3, IQR = 2–4), U = 702.00, z = −2.31, p = .021, r = .25.
- Wilcoxon result: Post-training confidence ratings (median = 4, IQR = 3–4) were higher than pretraining ratings (median = 3, IQR = 2–3), z = −3.42, p < .001, r = .55.
- Kruskal–Wallis result: Confidence differed across the three clinical units, H(2) = 8.74, p = .013, ε² = .07. Adjusted pairwise comparisons indicated that the difference was between the emergency and outpatient units.
- Spearman correlation: Greater evidence-based practice confidence was associated with more frequent guideline use, rs = .42, 95% CI [.25, .56], p < .001, n = 110.
- Independent-samples t test: Nurses who completed training (M = 3.99, SD = 0.55) reported higher composite confidence than nurses who did not (M = 3.64, SD = 0.62), t(113.30) = 3.26, p = .001, 95% CI [0.14, 0.56], d = 0.60.
Common Mistakes in Likert Scale Analysis
Common analytical errors include:
- Treating one question as a validated scale
- Combining unrelated items
- Mixing separate subscales
- Forgetting reverse coding
- Coding “not applicable” as neutral
- Ignoring the instrument’s missing-data rules
- Calculating totals before reviewing scoring instructions
- Running Cronbach’s alpha on unrelated variables
- Assuming all Likert data are continuous
- Assuming parametric tests are never permitted
- Choosing tests because they produce significance
- Reporting only p values
- Ignoring effect sizes and confidence intervals
- Collapsing categories without justification
- Copying unedited SPSS tables into a dissertation
- Interpreting association as causation
- Inventing cutoff scores
- Ignoring the original scoring manual
- Claiming statistical significance proves clinical importance
Likert Scale Analysis Checklist
- State the research question.
- Determine whether the outcome is one item or a multi-item scale.
- Review the instrument’s scoring instructions.
- Confirm response coding and value labels.
- Define missing and not-applicable responses.
- Identify and reverse-score negative items.
- Confirm the scale and subscale structure.
- Assess reliability where appropriate.
- Calculate the approved sum or average score.
- Select suitable descriptive statistics.
- Examine relevant statistical assumptions.
- Match the statistical test to the design.
- Report an effect size.
- Report a confidence interval where appropriate.
- Follow APA 7 statistical-reporting conventions.
- Interpret findings transparently without causal overstatement.
When to Seek Statistical Support
Additional statistical support may be appropriate when:
- Questionnaire scoring is unclear.
- The instrument contains several subscales.
- Reverse-scored items are confusing.
- Cronbach’s alpha is unexpectedly low.
- Missing data are extensive.
- A supervisor questions the selected test.
- Parametric and nonparametric approaches appear plausible.
- Repeated measurements are involved.
- Ordinal logistic regression is required.
- SPSS output is difficult to interpret.
- APA tables or the results chapter require revision.
These problems may require instrument review, data screening, or a revised analysis plan rather than simply running another test. Focused assistance is available through the SPSS Data Analysis Help service and Dissertation Data Analysis Help service.
Conclusion
Accurate Likert scale analysis requires more than assigning numbers to questionnaire responses. Nursing researchers must distinguish individual items from multi-item scales, follow official scoring instructions, handle missing and reverse-coded responses correctly, and assess reliability before combining eligible items.
The statistical method should reflect the research question, variable type, study design, score construction, distribution, and relevant assumptions. Frequencies and percentages usually provide the clearest description of individual items. Means and parametric tests may be defensible for appropriately constructed composite scores, but the decision must be justified rather than assumed.
Transparent APA 7 reporting should include sample sizes, descriptive results, test statistics, exact p values, effect sizes, and confidence intervals where appropriate. Statistical significance should never be interpreted as proof of causation or automatic evidence of clinical importance.
Get your Likert-scale analysis reviewed before submission. Share your questionnaire, dataset, research questions, university guidelines, and supervisor feedback through our Dissertation Data Analysis Help service for focused nursing research support.
References
American Psychological Association. (n.d.). Numbers and statistics guide. APA Style. https://apastyle.apa.org/instructional-aids/numbers-statistics-guide.pdf
Boateng, G. O., Neilands, T. B., Frongillo, E. A., Melgar-Quiñonez, H. R., & Young, S. L. (2018). Best practices for developing and validating scales for health, social, and behavioral research: A primer. Frontiers in Public Health, 6, Article 149. https://doi.org/10.3389/fpubh.2018.00149
de Winter, J. C. F., & Dodou, D. (2010). Five-point Likert items: t test versus Mann–Whitney–Wilcoxon. Practical Assessment, Research & Evaluation, 15, Article 11. https://doi.org/10.7275/bj1p-ts64
Harpe, S. E. (2015). How to analyze Likert and other rating scale data. Currents in Pharmacy Teaching and Learning, 7(6), 836–850. https://doi.org/10.1016/j.cptl.2015.08.001
IBM. (n.d.). IBM SPSS Statistics. Retrieved July 17, 2026, from https://www.ibm.com/products/spss-statistics
McNeish, D. (2018). Thanks coefficient alpha, we’ll take it from here. Psychological Methods, 23(3), 412–433. https://doi.org/10.1037/met0000144
Norman, G. (2010). Likert scales, levels of measurement and the “laws” of statistics. Advances in Health Sciences Education, 15(5), 625–632. https://doi.org/10.1007/s10459-010-9222-y
Sullivan, G. M., & Artino, A. R., Jr. (2013). Analyzing and interpreting data from Likert-type scales. Journal of Graduate Medical Education, 5(4), 541–542. https://doi.org/10.4300/JGME-5-4-18
Tavakol, M., & Dennick, R. (2011). Making sense of Cronbach’s alpha. International Journal of Medical Education, 2, 53–55. https://doi.org/10.5116/ijme.4dfb.8dfd
FAQs
1. What is the best method for analyzing Likert-scale data?
The best method depends on whether the outcome is an individual ordinal item or a properly constructed composite scale. Frequencies, percentages, medians, and nonparametric tests are often appropriate for individual items. Means and parametric tests may be defensible for multi-item composite scores when their construction and assumptions support that treatment.
2. Can I calculate a mean for a five-point Likert scale?
You should be cautious when calculating a mean for one item because the response categories are ordinal. A mean may be more defensible for a composite formed from several related items, especially when the scoring manual supports it and the score has an acceptable distribution.
3. Which statistical test should I use for Likert-scale data?
Mann–Whitney U, Kruskal–Wallis, Wilcoxon signed-rank, Friedman, and Spearman correlation are common options for ordinal outcomes. Justified composite scores may be analyzed using t tests, ANOVA, Pearson correlation, or linear regression when their assumptions are satisfied.
4. Do I need to reverse-score negatively worded Likert items?
Yes, when the negatively worded item runs in the opposite direction from the intended composite. Reverse scoring aligns it with the remaining items. Always retain the original variable and create a separate reversed version.
5. How should Likert-scale results be reported in APA 7?
Report sample sizes, appropriate descriptive statistics, the test statistic, degrees of freedom where applicable, exact p values unless p < .001, effect sizes, and confidence intervals where relevant. Clearly state whether the analysis concerns an individual item or a composite scale.