How to Report Data Collection Instruments, Validity, and Reliability
Clear reporting of data collection instruments is essential for evaluating what a study actually measured, how measurement scores should be interpreted, and how consistently those measurements were obtained.
When a study uses a scale, test, questionnaire, observation form, interview guide, instrument-based measurement, or another data collection tool, the source, structure, scoring procedure, and relevant measurement properties should be described adequately in the Methods section.
What Is a Data Collection Instrument?
A data collection instrument is a tool used to measure or record a variable, characteristic, experience, behavior, or outcome in a study.
Common examples include:
- Scales
- Questionnaires
- Achievement or knowledge tests
- Psychometric tests
- Observation forms
- Interview guides
- Checklists
- Clinical measurement instruments
- Laboratory and device-based measurements
- Institutional records and databases
What Information Should Be Reported for a Scale or Test?
When an established scale or test is used, authors should report, where applicable:
- Full name of the instrument
- Original developer or source
- Year of development
- Relevant adaptation study
- Number of items
- Subscales or dimensions
- Response format
- Scoring procedure
- Meaning of higher or lower scores
- Key validity and reliability evidence from previous studies
- Reliability or measurement-quality findings in the present sample
How Should the Source of a Scale Be Reported?
The original development study should be cited appropriately. If a translated or culturally adapted version is used, the relevant adaptation study should also be cited.
Providing only the name of an instrument is usually insufficient.
How Should the Structure of an Instrument Be Described?
Authors should identify the number of items, whether the scale is unidimensional or multidimensional, and what each subscale represents.
Example: The instrument consists of 18 items across three dimensions. Items are rated on a five-point Likert-type scale ranging from 1 to 5.
How Should Scoring Be Explained?
Readers should be able to understand how scores are derived.
Where relevant, report:
- Total score range
- Subscale scores
- Reverse-scored items
- Cut-off scores
- Meaning of high or low scores
What Is Validity?
Validity refers to the body of evidence supporting the interpretation and use of scores obtained from a measurement instrument.
Validity should not be treated as a permanent property that is established once and then assumed to apply in every population and context. Intended use, sample, setting, and measurement purpose all matter.
What Is Reliability?
Reliability concerns the consistency of measurement scores and the extent to which they are affected by measurement error.
A reliable instrument is not automatically valid. A measure may produce highly consistent scores while measuring the wrong construct.
What Is Internal Consistency?
Internal consistency evaluates how well items intended to measure the same construct function together.
Common indices include:
- Cronbach's alpha
- McDonald's omega
How Should Cronbach's Alpha Be Reported?
Cronbach's alpha is widely used to assess internal consistency in multi-item scales.
Example: Cronbach's alpha for the scale in the present study was .86.
Where subscales are analyzed separately, reporting reliability coefficients for the individual subscales may be more informative than providing only one total-scale coefficient.
Is There a Universal Cutoff for Cronbach's Alpha?
No single cutoff is universally appropriate. Values such as .70 are often referenced, but interpretation should consider the study purpose, number of items, construct, and intended use of the measurement.
Very high alpha values do not automatically indicate superior quality and may sometimes suggest excessive item redundancy.
What Is McDonald's Omega?
Omega is another coefficient used to evaluate internal consistency. In some measurement models, particularly when item loadings differ substantially, omega may provide more appropriate information than alpha.
Researchers may report omega alongside or instead of alpha where methodologically justified.
What Is Test-Retest Reliability?
Test-retest reliability evaluates the stability of scores over time by administering the same instrument to the same or similar participants at more than one time point.
The interval between measurements should be appropriate for the characteristic being measured.
What Is Inter-Rater Reliability?
When several evaluators rate or code the same material, agreement between raters may need to be assessed.
Possible statistics include:
- Cohen's kappa
- Fleiss' kappa
- Intraclass correlation coefficient (ICC)
- Percentage agreement
The appropriate statistic depends on the type of data and study design.
What Is Content Validity?
Content validity evaluates whether the items adequately represent the domain of the construct being measured.
Expert review is commonly used during instrument development.
Experts may assess:
- Item relevance
- Coverage of the construct
- Clarity
- Cultural appropriateness
Can a Content Validity Index Be Reported?
Studies using expert ratings may report item-level or scale-level content validity indices.
Authors should explain the method used and the number of experts involved.
What Is Construct Validity?
Construct validity concerns whether an instrument behaves as expected in relation to the theoretical construct it is intended to measure.
Evidence may include:
- Factor analysis
- Convergent validity
- Discriminant validity
- Hypothesis testing
- Expected group differences
What Is Exploratory Factor Analysis?
Exploratory Factor Analysis (EFA) is used to investigate the latent structure underlying relationships among items.
It is particularly useful during new scale development or when the underlying structure is not yet well established.
What Should Be Reported Before or During EFA?
Depending on the study, authors may report:
- KMO statistic
- Bartlett's test of sphericity
- Factor extraction method
- Rotation method
- Method used to determine the number of factors
- Item loadings
- Cross-loadings
- Explained variance
How Should the Number of Factors Be Determined?
Factor retention should not rely only on the eigenvalue-greater-than-one rule. Multiple sources of evidence may be considered together, such as:
- Parallel analysis
- Scree plot
- Theoretical expectations
- Interpretability of the resulting factors
What Is Confirmatory Factor Analysis?
Confirmatory Factor Analysis (CFA) tests how well a predefined measurement model fits the data.
It may be used in scale development, adaptation, or when a previously proposed factor structure is evaluated in a new sample.
What Should Be Reported in CFA?
Depending on the study, report:
- Specified factor structure
- Standardized factor loadings
- Estimation method
- Model fit indices
- Error covariances, if any, together with their justification
- Alternative model comparisons where relevant
How Should Model Fit Be Reported?
It is generally preferable to evaluate several fit indices rather than relying on a single statistic.
Commonly reported measures include:
- χ² and degrees of freedom
- CFI
- TLI
- RMSEA with confidence interval
- SRMR
Fixed cutoff values should not be treated as universally definitive rules. Model complexity, sample size, and estimation method should also be considered.
What Is Convergent Validity?
Convergent validity assesses whether indicators expected to measure the same or closely related constructs show sufficient convergence.
In latent-variable models, evidence may include:
- Factor loadings
- Average Variance Extracted (AVE)
- Composite Reliability (CR)
What Is Discriminant Validity?
Discriminant validity assesses whether theoretically distinct constructs can be distinguished empirically.
Possible approaches include:
- HTMT
- Fornell-Larcker criterion
- Evaluation of correlations among latent factors
What Is HTMT?
The Heterotrait-Monotrait Ratio (HTMT) is one approach to evaluating discriminant validity, particularly in latent-variable models.
HTMT values may be compared with justified thresholds or evaluated using bootstrap confidence intervals.
The selected decision rule should be explained and supported by an appropriate source.
What Is Criterion Validity?
Criterion validity examines whether scores from an instrument show the expected relationship with an external criterion.
This may include:
- Concurrent validity
- Predictive validity
What Is Known-Groups Validity?
A measure may also be evaluated by testing whether it distinguishes groups expected to differ theoretically.
For example, expected score differences between clinical and non-clinical groups may contribute to validity evidence.
What Should Be Reported in Scale Adaptation Studies?
Adapting an instrument developed in another language or culture requires more than a literal translation.
Depending on the study, the process may include:
- Translation
- Back-translation
- Expert review
- Cultural adaptation
- Pilot testing
- Factor-structure evaluation
- Reliability analysis
- Convergent and discriminant validity
- Criterion validity
What Should Be Reported in Scale Development Studies?
Scale development commonly involves several stages, which may include:
- Conceptual definition of the construct
- Item generation
- Expert evaluation
- Pilot testing
- Item analysis
- EFA
- CFA
- Reliability analysis
- Convergent, discriminant, and criterion validity
Using the same sample for every exploratory and confirmatory step is not always ideal. Where feasible, independent samples for EFA and CFA may provide a stronger design.
Is a Questionnaire the Same as a Scale?
No. Not every questionnaire is a psychometric scale.
A questionnaire may simply collect factual or descriptive information, whereas a scale typically consists of multiple items designed to measure a latent construct and is evaluated psychometrically.
How Should a Researcher-Developed Information Form Be Reported?
When a demographic or researcher-developed form is used, authors should state what information it collects.
Example: The Personal Information Form consisted of four questions assessing age, educational level, professional experience, and work unit.
Conventional psychometric validity and reliability analysis is not always appropriate for such factual information forms.
Reliability in Knowledge or Achievement Tests
For tests containing dichotomously scored items, coefficients such as KR-20 may be appropriate.
Depending on the test, item difficulty and item discrimination may also be reported.
Can Cronbach's Alpha Be Used for a Single-Item Measure?
No. Internal consistency coefficients require multiple items. Cronbach's alpha is not meaningful for a single-item measure.
Which Reliability Coefficient Should Be Used for Likert-Type Items?
The choice should depend on the data structure, measurement model, and assumptions. Although Cronbach's alpha is widely used, it is not automatically the optimal coefficient in every situation.
Ordinal-data approaches and alternative reliability coefficients may sometimes be more suitable.
How Should Reverse-Scored Items Be Handled?
Reverse-scored items must be recoded correctly before total or subscale scores are calculated.
Incorrect reverse scoring can substantially distort total scores, factor structure, and reliability estimates.
Should Items Be Removed to Increase Alpha?
Items should not be removed automatically simply because deletion increases Cronbach's alpha.
Item-removal decisions should consider:
- Theoretical meaning
- Item-total relationship
- Factor loading
- Cross-loadings
- Content coverage
- Previous validity evidence
Is Reporting the Original Alpha Sufficient?
No. Reliability coefficients reported in the original development or adaptation study are informative, but the reliability of scores in the current sample should also be reported in most applications.
Reliability may vary across samples and administration conditions.
What Is Measurement Invariance?
When scale scores are compared across groups, researchers may need to assess whether the construct is measured in a comparable way across those groups.
This may be relevant across:
- Gender groups
- Cultures
- Languages
- Countries
- Time points
Depending on the analysis, configural, metric, scalar, and strict invariance may be evaluated.
Is Permission Required to Use a Scale?
This depends on the copyright and licensing conditions of the instrument.
Some measures may be freely available, whereas others require permission from the developer or rights holder.
Authors are responsible for checking the applicable conditions of use.
Should All Scale Items Be Reproduced in the Article?
Not always. Reproducing all items may be inappropriate depending on copyright status and the purpose of the manuscript.
For a newly developed scale, providing items as supplementary material may be useful, subject to copyright and publication policies.
How Are Validity and Reliability Addressed in Qualitative Research?
It is not always appropriate to transfer quantitative concepts of validity and reliability directly into qualitative research.
Depending on the methodological approach, quality may be discussed in terms of:
- Credibility
- Transferability
- Dependability
- Confirmability
- Reflexivity
How Can Credibility Be Strengthened in Qualitative Research?
Where consistent with the study design, possible strategies include:
- Prolonged engagement
- Data triangulation
- Researcher triangulation
- Participant validation
- Peer debriefing
- Examination of negative or deviant cases
Is Member Checking Required in Every Qualitative Study?
No. Member checking is not a universal requirement across all qualitative methodologies. Quality strategies should be justified in relation to the specific methodological approach.
Is Inter-Coder Agreement Required in Every Qualitative Study?
No. In some interpretive qualitative approaches, complete agreement between researchers is not considered the primary indicator of analytical quality.
Inter-coder agreement should be used when it is meaningful for the chosen analytical approach, and the method of evaluation should be described clearly.
How Should an Interview Guide Be Reported?
For a semi-structured interview guide, authors may describe:
- How the guide was developed
- Number of primary questions
- Whether expert review was conducted
- Whether pilot interviews were undertaken
- Whether follow-up or probing questions were used
How Should Laboratory and Device-Based Measurements Be Reported?
Studies using laboratory or instrument-based measurement may report the manufacturer and model, calibration procedures, measurement protocol, and relevant technical specifications.
Sufficient technical information should be provided to support reproducibility.
How Should a Data Collection Instrument Be Described in the Methods Section?
A useful sequence is:
Instrument name → Developer/source → Structure → Scoring → Previous measurement evidence → Measurement findings in the present study
Example of Scale Reporting
Example: The X Scale was used in the study. The instrument was developed by Yılmaz and Demir and contains 18 items across three subscales. Items are scored from 1 (strongly disagree) to 5 (strongly agree), with higher scores indicating higher levels of X. The original study supported a three-factor structure and reported internal consistency coefficients ranging from .78 to .89 across subscales. In the present study, Cronbach's alpha for the total scale was .87.
Example of CFA Reporting
Example: The proposed three-factor structure was evaluated using confirmatory factor analysis. Model fit was assessed using χ²/df, CFI, TLI, RMSEA, and SRMR, and standardized factor loadings and interfactor relationships were also examined.
Common Mistakes
- Naming an instrument without explaining its structure
- Failing to cite the original development study
- Using an adapted version without citing the adaptation study
- Failing to report reliability in the present sample
- Presenting Cronbach's alpha as evidence of validity
- Treating a high alpha value as sufficient proof of validity
- Calculating alpha for a single-item measure
- Reporting only total-scale alpha when subscales are analyzed separately
- Determining the number of factors solely by the eigenvalue > 1 rule
- Relying on only one fit index in CFA
- Adding correlated errors based only on modification indices without substantive justification
- Using HTMT or other criteria without explaining the decision rule
- Deleting items solely to increase alpha
- Applying quantitative validity-reliability language mechanically to qualitative studies
- Assuming member checking is mandatory in every qualitative study
Check Before Submission
- Are all data collection instruments clearly described?
- Is the original source of each scale or test cited?
- If an adapted version was used, is the adaptation study cited?
- Are item number, dimensions, and scoring procedures explained?
- Is the meaning of high and low scores clear?
- Are reliability coefficients reported for the current sample?
- Are validity analyses appropriate to the study purpose?
- If EFA or CFA was conducted, are methodological details reported adequately?
- Are model fit indices evaluated together rather than individually?
- Are convergent and discriminant validity assessed where relevant?
- Are item deletions or model modifications substantively justified?
- In qualitative research, are quality criteria consistent with the chosen methodology?