Reliability and Validity — Cronbach's Alpha, AVE, HTMT in Plain English
Reviewers want to know two things about your scales: do they measure what they claim to measure (validity), and do they measure it consistently (reliability)? This guide unpacks the five tests that publishable papers report — Cronbach's alpha, composite reliability, average variance extracted, the Fornell–Larcker criterion, and HTMT.
Reliability first — Cronbach's alpha
Reliability is internal consistency: do the items in a scale agree with each other? Cronbach's alpha (α) is the most common measure. It ranges from 0 to 1, with rules of thumb:
- α < 0.60 — unacceptable; the items aren't measuring the same thing.
- 0.60 ≤ α < 0.70 — questionable; acceptable for exploratory research with a sentence of justification.
- 0.70 ≤ α < 0.80 — acceptable; most journals accept this without comment.
- α ≥ 0.80 — good. Above 0.95 is often a sign of redundant items.
How to fix a low alpha. Check the item-total correlations. Items with corrected item-total correlations below 0.30 are dragging the scale down. Smart Form's analytics dashboard shows alpha-if-item-deleted for every item — so you can see exactly which to drop.
The newer alternative — Composite reliability (CR)
Cronbach's alpha assumes all items contribute equally to the construct. Composite reliability (sometimes called ω) accounts for differing factor loadings, which is more realistic. Most CFA software reports it. Same thresholds as alpha: above 0.70 is acceptable, above 0.80 is good.
Validity — the four flavours
Reliability tells you the scale is consistent; validity tells you it's measuring the right thing. There are four kinds, in order of difficulty:
- Content validity — do the items cover the full meaning of the construct? Established by expert review and is conceptual, not statistical.
- Face validity — do the items look right to respondents? Pilot tests reveal this.
- Convergent validity — do items in the same scale converge with each other? Tested by Average Variance Extracted (AVE).
- Discriminant validity — do items in different scales discriminate from each other? Tested by Fornell–Larcker and increasingly HTMT.
Convergent validity — Average Variance Extracted (AVE)
AVE captures how much variance in a construct is explained by its indicators, on average. Rule: AVE ≥ 0.50 means the construct explains more variance than measurement error.
If AVE falls below 0.50 but CR is above 0.60, some methodologists (Fornell & Larcker, 1981) consider this acceptable. Many reviewers won't.
Discriminant validity — the Fornell–Larcker criterion
For each pair of constructs, the square root of each AVE should be greater than the correlation between the two constructs. Visually: the diagonal of the construct correlation matrix (with √AVE on the diagonal) should be larger than every off-diagonal value in its row and column.
The newer, stricter test — HTMT
Henseler, Ringle and Sarstedt (2015) showed that Fornell–Larcker often fails to detect discriminant validity problems. They proposed the heterotrait-monotrait ratio (HTMT) as a more sensitive test.
Rules of thumb for HTMT:
- HTMT < 0.85 — strict threshold for conceptually distinct constructs.
- HTMT < 0.90 — liberal threshold for conceptually similar constructs.
- HTMT ≥ 0.90 — discriminant validity fails; the two scales are too similar.
Increasingly, HTMT is the test reviewers want to see in addition to Fornell–Larcker. Both have their place.
Common method bias — the silent killer
Even if your scales are reliable and valid in isolation, the correlations between them can be inflated by common method bias (CMB) — the artificial covariance that arises when all variables are measured at the same time, from the same source, using the same response format.
Tests you can run:
- Harman's single-factor test — load all items on a single factor. If the first factor explains less than 50% of the variance, CMB is not a major concern. (Increasingly seen as weak.)
- Common latent factor — add a method factor to your CFA and re-run. Significant change in loadings indicates CMB.
- Marker variable approach — include an unrelated marker scale and partial out its correlation with your focal constructs.
What to report in your methods section
A defensible psychometrics paragraph reports:
- Cronbach's α for each scale.
- Composite reliability (CR) for each scale.
- AVE for each scale.
- √AVE on the diagonal of the construct correlation matrix (Fornell–Larcker).
- HTMT ratios for each construct pair.
- One CMB test (Harman or common latent factor).
Smart Form's live Reliability and Validity dashboard calculates all of these automatically as responses arrive — Cronbach's α, CR, AVE, Fornell–Larcker, HTMT, plus Harman's single-factor test — so you can spot problems while there's still time to fix them.
See it live
Smart Form computes reliability and validity automatically from the responses to any published questionnaire — no SPSS, no Mplus, no R required for the basics. Use it to pilot, screen, and report.
Final thought
Reviewers don't expect perfection — they expect transparency. Report the thresholds you used, acknowledge anything that fell short of strict criteria, and explain why your scales are still adequate for the study's conclusions. A paper with α = 0.68 and an honest discussion will get further than a paper with α = 0.92 and no discriminant validity test at all.