How to Report Quantitative and Statistical Data Analysis

How to Report Quantitative and Statistical Data Analysis

Quantitative data analysis refers to the statistical procedures used to evaluate research questions and hypotheses using numerical data. In a scientific article, it is not enough to list the statistical tests performed. Authors should also explain why the analyses were selected, how assumptions were evaluated, and how the resulting estimates were interpreted.

The analysis plan should be aligned with the research objective, data structure, measurement level of variables, study design, and hypotheses.

What Should the Statistical Analysis Section Explain?

Readers should be able to determine:

  • Which statistical software was used?
  • How were the data checked before analysis?
  • How were missing data handled?
  • Were outliers assessed?
  • How were distributional and other assumptions evaluated?
  • Which analysis addressed each research question?
  • What significance level was used?
  • Were effect sizes and confidence intervals reported?

How Should Statistical Software Be Reported?

The statistical software and, where relevant, its version should be identified.

Examples include:

  • IBM SPSS Statistics
  • R
  • JASP
  • Jamovi
  • Stata
  • SAS
  • AMOS
  • Mplus
  • SmartPLS
  • MATLAB
  • Python-based analytical tools

Example: Data were analyzed using IBM SPSS Statistics 29.

Data Screening Before Analysis

Before statistical testing, the data should be examined for analytical suitability. Depending on the study, screening may include:

  • Data-entry errors
  • Missing data
  • Outliers
  • Distributional characteristics
  • Variable coding
  • Correct recoding of reverse-scored items
  • Correct calculation of total and subscale scores

How Should Missing Data Be Handled?

The extent and pattern of missing data should be evaluated. Simply deleting every incomplete case is not always appropriate.

Possible approaches include:

  • Listwise deletion
  • Analysis-specific available cases
  • Simple imputation methods
  • Multiple imputation
  • Maximum-likelihood-based approaches

The selected approach should be compatible with the amount and pattern of missingness and with the analytical method.

What Is an Outlier?

An outlier is an observation that differs substantially from the general pattern of the data. Its presence does not automatically justify deletion.

Outliers may result from:

  • Data-entry errors
  • Genuine unusual observations
  • Distinct subgroups
  • Measurement error

How Can Outliers Be Evaluated?

Depending on the analysis, researchers may use:

  • Standardized scores
  • Box plots
  • Mahalanobis distance
  • Cook's distance
  • Leverage statistics
  • Standardized residuals

If observations are excluded, the decision should have a scientific and statistical justification.

How Should Normality Be Evaluated?

Normality should generally not be assessed through a single statistical test alone.

Evidence may include:

  • Histograms
  • Q-Q plots
  • Skewness
  • Kurtosis
  • Shapiro-Wilk test
  • Other appropriate distributional diagnostics

In large samples, formal normality tests can become highly sensitive to minor departures, so graphical and distributional evidence should also be considered.

Parametric and Nonparametric Tests

The choice between parametric and nonparametric analysis should not depend solely on whether a variable is “normally distributed.”

Other test assumptions and the structure of the data also matter. Where assumptions are not adequately met, options may include:

  • Data transformation
  • Robust statistical methods
  • Bootstrap procedures
  • Appropriate nonparametric tests

How Should Descriptive Statistics Be Reported?

Descriptive statistics summarize the basic characteristics of the data.

For continuous variables, relevant statistics may include:

  • Mean
  • Standard deviation
  • Median
  • Interquartile range
  • Minimum and maximum

Categorical variables are commonly reported using counts and percentages.

Example: Mean participant age was 34.8 ± 8.6 years.

For a skewed distribution:

Median age was 33 years, with an interquartile range of 28–41 years.

How Are Frequencies and Percentages Reported?

Categorical data may commonly be presented as:

n (%)

Example: Of the participants, 184 (61.3%) were women.

Which Tests Are Used to Compare Two Groups?

For two independent groups and a continuous outcome, an independent-samples t test may be used when its assumptions are sufficiently met.

When assumptions are not appropriate, alternatives such as the Mann-Whitney U test may be considered.

For two measurements obtained from the same participants, a paired-samples t test may be appropriate, with the Wilcoxon signed-rank test serving as a possible alternative in relevant circumstances.

How Are Three or More Groups Compared?

One-way ANOVA may be used to compare three or more independent groups when its assumptions are appropriate.

When the overall test is significant, suitable post hoc comparisons may be conducted to identify which groups differ.

The Kruskal-Wallis test may be considered where a nonparametric approach is appropriate.

How Is Homogeneity of Variance Evaluated?

Homogeneity of variance is relevant to some group-comparison procedures. Levene's test may be used as one diagnostic.

When equal variances cannot reasonably be assumed, methods such as Welch's t test or Welch ANOVA may be more suitable.

How Are Categorical Variables Compared?

A chi-square test may be used to assess associations between categorical variables.

When expected cell frequencies are small, Fisher's exact test or another appropriate method may be required.

How Should Correlation Analysis Be Selected?

Pearson correlation may be appropriate for evaluating linear relationships between continuous variables when relevant assumptions are satisfied.

Spearman correlation may be appropriate for ordinal data or monotonic relationships when Pearson assumptions are not suitable.

A correlation coefficient describes association; it does not demonstrate causality.

How Should a Correlation Result Be Reported?

Example: X was moderately and positively associated with Y (r = .42, p < .001).

A confidence interval for the correlation coefficient may also be reported where appropriate.

When Is Regression Analysis Used?

Regression analysis may be used to examine how one or more predictors are associated with or predict an outcome.

Depending on the model, reporting may include:

  • Regression coefficients
  • Standard errors
  • Standardized coefficients where useful
  • Confidence intervals
  • R² and adjusted R²
  • Overall model test

What Assumptions Should Be Evaluated in Regression?

Relevant assumptions depend on the model. In linear regression, considerations may include:

  • Linearity
  • Residual distribution
  • Homoscedasticity
  • Independence
  • Multicollinearity
  • Highly influential observations

How Is Multicollinearity Evaluated?

Strong overlap among predictors can reduce the stability and interpretability of regression coefficients.

Possible diagnostics include:

  • Variance Inflation Factor (VIF)
  • Tolerance
  • Correlations among predictors

How Is Logistic Regression Reported?

Logistic regression may be used when the outcome is binary.

Results may include:

  • Regression coefficient
  • Standard error
  • Odds ratio
  • 95% confidence interval
  • p value

Example: X was associated with higher odds of the outcome (OR = 1.72, 95% CI [1.21, 2.44], p = .002).

Confounders in Multivariable Models

Potential confounders should be selected on the basis of the research question, theory, previous evidence, and study design.

Automatically including only variables that were significant in univariable analysis is not appropriate for every study.

How Should Mediation Analysis Be Reported?

Mediation analysis evaluates whether the relationship between X and Y operates partly through a mediator M.

Reporting may include:

  • X → M path
  • M → Y path
  • Direct effect
  • Indirect effect
  • Total effect where relevant
  • Bootstrap confidence interval

The indirect effect itself should be evaluated rather than relying only on the statistical significance of individual component paths.

How Should Moderation Analysis Be Reported?

Moderation analysis evaluates whether the relationship between X and Y differs according to levels of Z, usually through an interaction term.

Reporting may include:

  • Main effects
  • X × Z interaction coefficient
  • Confidence interval
  • p value
  • Simple slopes or conditional effects

How Should Structural Equation Modeling Be Reported?

In structural equation modeling, the measurement and structural components should be defined clearly.

Depending on the study, authors may report:

  • Estimation method
  • Model fit indices
  • Standardized path coefficients
  • Standard errors
  • p values
  • Confidence intervals
  • Explained variance
  • Direct and indirect effects

Why Should Effect Size Be Reported?

Statistical significance alone does not indicate whether an effect is scientifically, clinically, or practically important.

Possible effect-size measures include:

  • Cohen's d
  • Hedges' g
  • Eta squared
  • Partial eta squared
  • r
  • Odds ratio
  • Risk ratio

Why Are Confidence Intervals Important?

Confidence intervals provide information about the precision and uncertainty of an estimate.

For example:

OR = 1.72, 95% CI [1.21, 2.44]

contains more information than a p value alone.

How Should p Values Be Reported?

Exact p values may be reported where appropriate.

Example: p = .032

For very small values:

p < .001

Reporting p = .000 is incorrect.

Are Results with p > .05 Unimportant?

No. Statistically non-significant findings may still be scientifically informative.

Interpretation should also consider:

  • Effect size
  • Confidence interval
  • Sample size
  • Measurement precision
  • Theoretical relevance

How Should Multiple Comparisons Be Addressed?

Testing many hypotheses can increase the probability of false-positive findings.

Depending on the analytical plan, approaches may include:

  • Bonferroni adjustment
  • Holm procedure
  • False Discovery Rate methods
  • Other appropriate corrections

Adjustment is not automatically required for every analysis; the decision should reflect the structure of primary and secondary hypotheses.

When Is Bootstrap Analysis Useful?

Bootstrap methods repeatedly resample the observed data to estimate the sampling distribution of a statistic.

They may be useful for:

  • Indirect effects
  • Confidence intervals
  • Some situations involving weak distributional assumptions

The number of bootstrap samples and confidence-interval method may be reported where relevant.

One-Tailed and Two-Tailed Tests

A one-tailed test should generally be used only when a directional hypothesis was justified before analysis.

Switching to a one-tailed test after observing the results is inappropriate.

Distinguish Confirmatory and Exploratory Analyses

Analyses planned in advance to test primary hypotheses should, where possible, be distinguished from exploratory analyses conducted after inspection of the data.

This distinction improves transparency.

Statistical Significance Is Not the Same as Scientific Importance

With very large samples, even very small effects may reach statistical significance. With small samples, potentially important effects may fail to cross a conventional significance threshold.

Results should therefore not be interpreted only as “significant” or “non-significant.”

How Should Statistical Analysis Be Written in the Methods Section?

A useful sequence is:

Software → Data screening → Descriptive analyses → Assumption checks → Tests linked to research questions → Effect sizes / confidence intervals → Significance threshold

Example Statistical Analysis Paragraph

Example: Data were analyzed using IBM SPSS Statistics 29. Continuous variables were summarized using means and standard deviations or, where appropriate, medians and interquartile ranges; categorical variables were presented as counts and percentages. Distributional characteristics were evaluated using graphical methods and skewness and kurtosis. Independent-samples t tests were used for two-group comparisons where assumptions were met, and chi-square tests were used for categorical variables. Relationships among variables were examined using Pearson or, where appropriate, Spearman correlations. Multivariable relationships were evaluated using regression models, with estimates presented alongside 95% confidence intervals. Two-sided p values below .05 were considered statistically significant.

This paragraph illustrates reporting structure only. Authors should report only analyses that were actually conducted in their own study.

How Should Statistical Results Be Written in the Results Section?

Results should not be limited to statements such as “a significant difference was found.”

Depending on the analysis, report:

Test statistic + degrees of freedom + p value + effect size + confidence interval

Example: Scores were lower in the intervention group than in the control group, t(198) = -2.84, p = .005, d = 0.40.

Common Statistical Analysis Mistakes

  • Selecting tests without reference to the research question
  • Using a formal normality test as the only basis for choosing parametric methods
  • Deleting outliers without justification
  • Failing to explain how missing data were handled
  • Conducting large numbers of unnecessary tests
  • Reporting p = .000
  • Reporting only p values
  • Ignoring effect sizes
  • Omitting non-significant findings
  • Inferring causality from correlation
  • Automatically interpreting regression prediction as causal effect
  • Failing to evaluate multicollinearity and other model assumptions
  • Evaluating mediation only through significance of separate paths
  • Judging an SEM model using only one fit index
  • Changing hypotheses after seeing the statistical results

Check Before Submission

  • Are analyses aligned with the research questions and hypotheses?
  • Are data screening and missing-data procedures explained?
  • Is outlier assessment described?
  • Were relevant analytical assumptions evaluated?
  • Are descriptive statistics appropriate for the data structure?
  • Are test names and software reported accurately?
  • Are p values reported correctly?
  • Are effect sizes provided where relevant?
  • Are confidence intervals provided where relevant?
  • Were multiple-testing issues considered where necessary?
  • Are non-significant findings also reported?
  • Does causal language remain within the limits of the study design?

Related Guides