Measurement is the quiet engine of empirical research. Before any hypothesis can be tested, any policy evaluated, or any intervention judged effective, an unobservable theoretical construct such as patient satisfaction, service quality, organizational commitment, or quality of life has to be turned into numbers. The discipline that governs this translation is measurement science, and the health of every downstream finding depends on how well it is done. Among the threats to that process, one that has drawn growing attention over the past two decades is redundancy: the situation in which a measurement instrument carries the same information more than once, whether at the level of items, indicators, constructs, or whole scales.
Redundancy looks like a technical nuisance, but it is not. It sits at the root of a wide range of problems, from the artificial inflation of reliability coefficients, to the erosion of discriminant validity, to bias in parameter estimates, and finally to the needless proliferation of constructs that fragment a research field (Boyle, 1991; MacKenzie et al., 2011; Podsakoff et al., 2016). This post lays out what redundancy is, why it matters, how to detect it, and what to do about it, with particular attention to the realities of health management research.
What Redundancy Actually Is
Redundancy is the presence, within a measurement instrument, of units that carry overlapping information to the point of being interchangeable (Boyle, 1991; DeVellis, 2017). The overlap is not merely a matter of high correlation between items. Read more carefully, it describes a single, undivided portion of the theoretical domain being sampled repeatedly from the same direction. Boyle (1991), in his foundational note, argued that high internal consistency cannot be treated as a virtue in itself, because a high correlation may simply mean that the items are asking almost the same thing. That argument has since been extended into a standard component of scale development practice.
It also helps to distinguish redundancy from its close neighbors. Multicollinearity refers to high linear association among predictor variables and destabilizes parameter estimates in regression. Redundancy is broader: it includes high correlation but also conceptual overlap, so two indicators can be statistically distinct yet conceptually redundant, and the reverse can also hold. The halo effect, a response bias in which a global judgment contaminates ratings of specific attributes, can be a source of redundancy but is not identical to it (Campbell & Fiske, 1959). Common method bias, the systematic variance produced when all items share a single measurement method, likewise inflates inter-item correlations without being the same thing as theoretical redundancy. Keeping these apart is what prevents a researcher from reaching for the wrong remedy.
The Four Levels of Redundancy
Redundancy cannot be discussed at a single level of analysis. It surfaces in different forms at different layers of the measurement operation, and each layer carries its own diagnostic tools and intervention strategies.
Item redundancy is the most familiar form. It occurs when the items of a scale restate the same content in slightly different words. A classic example is a satisfaction scale that includes “I am satisfied with this service,” “This service met my expectations,” and “I was pleased with this service” together. The high correlation appears to raise reliability, but the rise signals that the same region of the construct has been sampled several times rather than that the domain has been better represented. Cattell (1973) described this as a narrow band problem, in which a high mean inter-item correlation points to a constricted theoretical domain. Streiner (2003) made the same point pragmatically, noting that alpha values above 0.90 are often a sign of item surplus rather than excellence.
Indicator redundancy is conceptually close to item redundancy but operationally distinct. An item is typically a single survey question, whereas an indicator is the measurement unit used to represent a latent construct in parameter-based models such as structural equation modeling. In reflective measurement models, indicators are treated as manifestations of the latent construct, so high correlation among them is expected and welcome. In formative models, by contrast, the indicators constitute the construct, and high correlation violates the model’s core assumption and biases parameter estimates (Diamantopoulos & Winklhofer, 2001; Diamantopoulos et al., 2008). Hair et al. (2022), working within partial least squares structural equation modeling, recommend assessing multicollinearity among formative indicators using the variance inflation factor, where values above 5 indicate serious redundancy and values of 3 or higher already warrant attention.
Construct redundancy arises when latent constructs published under different names in fact represent largely the same theoretical territory. Le et al. (2010), in an influential empirical investigation, showed that closely related constructs such as job satisfaction, organizational commitment, and work engagement share much of their variance and that their discriminant validity falls below accepted methodological standards. Shaffer et al. (2016) extended this work, arguing that construct proliferation is a substantial problem in organizational research and that evidence of discriminant distinctiveness should be supplied before any new construct is proposed. Henseler et al. (2015) developed the heterotrait-monotrait ratio of correlations to diagnose this issue more sensitively than the older Fornell-Larcker criterion, with values below 0.85 (or 0.90 for conceptually proximate constructs) taken as evidence against redundancy. Rönkkö and Cho (2022) later showed that even the heterotrait-monotrait ratio can mislead under certain conditions, which is why the field now agrees that construct redundancy should be judged against a set of criteria rather than any single one.
Scale redundancy is the highest level of overlap. It occurs when different scales built to measure the same construct coexist without anyone asking whether each provides incremental value over the others. In health research, quality of life is measured with the SF-36, the EQ-5D, the WHOQOL-BREF, and many others used in parallel; in organizational research, job satisfaction is measured with the Minnesota Satisfaction Questionnaire, the Job Descriptive Index, and the Job in General scale, often treated as interchangeable. Drolet and Morrison (2001) and Bergkvist and Rossiter (2007) examined the comparative predictive validity of multiple-item and single-item measures and found that for some concrete constructs the incremental value of the longer multiple-item scale is quite limited. Scale redundancy is therefore not only surplus; it is a source of inefficiency that raises research cost and respondent burden, a consideration that matters greatly in health settings where respondent fatigue lowers completion rates.

How Different Theoretical Frameworks See Redundancy
The way redundancy is defined and assessed depends heavily on the measurement theory in which one is working.
Within Classical Test Theory, an observed score is modeled as the sum of a true score and an error score, and reliability is gauged mainly through internal consistency coefficients, above all Cronbach’s alpha (Cronbach, 1951; Lord & Novick, 1968). Because alpha is sensitive both to the mean inter-item correlation and to the number of items, a high value can mean either a genuinely consistent scale or serious redundancy, and the two cannot be separated by looking at alpha alone (Cortina, 1993). Streiner (2003) treats the 0.70 to 0.90 range as ideal and reads values above 0.90 as a warning sign. Within this framework, item redundancy is diagnosed through item-total correlations, the alpha-if-item-deleted statistic, and the mean inter-item correlation, with Cattell’s (1973) acceptable range of roughly 0.15 to 0.50 serving as a complementary guide.
Item Response Theory reframes the problem around the assumption of local independence: once the latent trait level is held constant, responses to items should be independent (Embretson & Reise, 2000). When that assumption is violated, parameter estimates become biased, standard errors shrink, and information functions are inflated, and one major cause of such violation is item redundancy. Yen (1984) proposed the Q3 statistic, which computes the correlation among residual responses not explained by the model, with values above 0.20 signaling local dependence between an item pair.
Structural equation modeling splits into covariance-based and variance-based traditions, and each diagnoses redundancy in its own way. In covariance-based modeling, after a confirmatory factor analysis is fitted, modification indices are inspected; a high index between two indicators (commonly 10 or above) suggests a model-unspecified correlation between their error terms that often reflects indicator redundancy. Brown (2015) notes that the researcher then has two options, either removing one indicator or, with explicit theoretical justification, freeing a covariance between the error terms, the second of which becomes data mining when no theory supports it. In variance-based modeling, Hair et al. (2022) recommend examining variance inflation factors for formative constructs and outer loadings together with the heterotrait-monotrait ratio for reflective ones.
Network psychometrics offers a more recent alternative that questions the latent variable paradigm altogether (Borsboom & Cramer, 2013; Epskamp et al., 2018). Here items are modeled not as reflections of a common latent construct but as nodes that interact directly, and redundancy appears as excessively high partial correlations between items. Christensen et al. (2023) proposed Unique Variable Analysis, a network-based method that detects locally dependent (that is, redundant) items and recommends combining them into a single composite or removing them. Its advantage is that it reads redundancy directly from the data structure without prior assumptions about the number or shape of latent constructs.
Why Redundancy Demands a Response
Recognizing redundancy is not enough; a researcher needs to understand its concrete consequences before treating it as a priority.
The first consequence is artificially high reliability. A high alpha has long been read as a mark of quality, yet it can arise either from genuine consistency or from serious redundancy, and the two cannot be told apart by the coefficient alone (Boyle, 1991; Cortina, 1993; Streiner, 2003). In health management this shows up routinely. A patient satisfaction scale whose items all ask nearly the same thing can reach an alpha of 0.95, creating an impression of a precise and sensitive instrument while the remaining dimensions of the domain, such as communication, met expectations, or continuity of care, go undersampled.
The second consequence is the erosion of discriminant validity. Construct redundancy weakens the empirical separability of two constructs that are assumed to be conceptually distinct (Campbell & Fiske, 1959). The Fornell and Larcker (1981) criterion requires the average variance extracted by a construct to exceed its squared correlation with other constructs, but Henseler et al. (2015) showed this criterion to be insufficiently sensitive in models with many variables, which is why they offered the heterotrait-monotrait ratio as an alternative. Beyond the technical question, Le et al. (2010) demonstrated that job satisfaction and organizational commitment overlap so heavily in predictive power that treating them as subcomponents of a broader positive work attitude may be more productive than treating them as separate concepts.
The third consequence is bias in parameter estimates, which is especially acute in formative measurement models. Because formative indicators constitute the construct, high correlation among them is not expected, and a high variance inflation factor destabilizes the outer weights, so that small changes in the sample can flip their sign and magnitude. Diamantopoulos and Winklhofer (2001) recommend that formative indicators first be defined comprehensively, then tested for multicollinearity, with any indicator whose variance inflation factor exceeds 3 reconsidered. In reflective models high indicator correlation is consistent with the model, but correlated errors still point to unmodeled redundancy, and although freeing such a covariance on the basis of a high modification index is methodologically legitimate, it must be backed by theory or it risks overfitting the sample-specific error structure.
The fourth consequence is respondent burden and degraded response quality. Long questionnaires lead respondents to develop fatigue and answer later items with less care (Galesic & Bosnjak, 2009). Items that ask the same thing in different words can shake respondents’ trust and strengthen social desirability tendencies. Bergkvist and Rossiter (2007) showed that single-item measures can match multiple-item ones in predictive validity for concrete constructs, which means the marginal contribution of a long scale does not always justify the burden it imposes.
Detecting Redundancy
Detection requires the evaluation of a set of indicators rather than any single number, and it combines quantitative and qualitative evidence.
On the quantitative side, the Classical Test Theory toolkit includes the mean inter-item correlation, the item-total correlation, the alpha-if-item-deleted value, the ratio of alpha to the number of items, and the structure of the inter-item correlation matrix that accompanies a high alpha. A mean inter-item correlation above 0.50 points to the narrow band problem. In exploratory factor analysis, an item loading very highly (above 0.90) on a single factor and clustering with another item on that same factor suggests the two may be substitutes, and Tabachnick and Fidell (2019) advise separate scrutiny of cases where loadings are extreme and a factor is defined by only two items. In the structural equation modeling context, the relevant indicators fall into four categories: measurement-model diagnostics such as standardized residual correlations and modification indices; multicollinearity diagnostics such as the variance inflation factor, tolerance, and the condition index; discriminant validity diagnostics such as the Fornell-Larcker criterion, the heterotrait-monotrait ratio, and cross-loadings; and structural-model diagnostics such as the stability of path coefficients and the plausibility of R-squared values.
The following thresholds offer a practical orientation. Cronbach’s alpha above 0.90 can indicate item redundancy (Cronbach, 1951; Streiner, 2003). An item-total correlation above 0.70 should be questioned in light of the number of items (Nunnally & Bernstein, 1994). A variance inflation factor above 3 is a warning and above 5 is serious in variance-based modeling (Hair et al., 2022). A heterotrait-monotrait ratio below 0.85 is acceptable, with 0.90 reserved for proximate constructs (Henseler et al., 2015). A modification index above 10 invites theoretical questioning (Brown, 2015; Kline, 2023). A Q3 value above 0.20 indicates local dependence (Yen, 1984), and Unique Variable Analysis flags locally dependent items in the partial correlation network and recommends merging them (Christensen et al., 2023).
On the qualitative side sits expert review, which is as important as the numbers and is too often skipped. The content validity ratio and content validity index, proposed by Lawshe (1975) and refined by Polit et al. (2007), draw on expert judgment of how well items represent the conceptual domain. Experts can also be asked directly about conceptual overlap among items, which catches a boundary that numbers miss: two items can be statistically distinct yet judged to represent the same conceptual region, and two items can correlate highly yet, in expert judgment, tap different domains with the high correlation being an ordinary empirical artifact. Triangulating quantitative and qualitative evidence is what makes a redundancy diagnosis methodologically sound.
What to Do When Redundancy Is Found
The right intervention depends on the level and nature of the redundancy detected.
At the item level, the most common response is to remove the items judged redundant, but elimination should not be blind. DeVellis (2017) recommends a protocol in which quantitative and qualitative evidence are weighed together: items with very high item-total correlations are placed on a candidate list, item pairs with high modification indices are identified, expert review decides which member of each pair is conceptually broader, and the narrower item is removed. The subtle point is that removal must be followed by a recheck of whether the scale still represents its theoretical domain, because, as Cattell (1973) warned, cutting items to reduce redundancy can deepen the narrow band problem. Elimination should therefore proceed alongside a map of the theoretical domain, with the question of whether each removed item’s subregion is still sampled by the remaining items asked repeatedly.
At the indicator and model level, simple item removal can be insufficient and the measurement model itself may need rethinking. One option is to move to a bifactor formulation, which models a general construct and relatively independent specific subconstructs at the same time, so that indicators labeled redundant in a first-order model may in fact represent a distinct subdomain (Reise, 2012). A second option is exploratory structural equation modeling, which relaxes the strict zero-loading constraints of confirmatory factor analysis and allows indicators to load to a small degree on more than one factor, modeling the cross-loadings common in real data in a theoretically legitimate way (Asparouhov & Muthén, 2009; Marsh et al., 2014). A third option is to revisit the reflective-formative distinction itself: a construct modeled reflectively but showing low indicator correlations may belong in a formative specification, and a formative one with unacceptably high correlations and variance inflation factors may belong in a reflective specification, a decision that must rest on theory and not only on the empirical evidence (Diamantopoulos et al., 2008).
At the construct level, where two constructs overlap empirically to a large degree, a simple model revision will not suffice and the rethinking has to be conceptual. Le et al. (2010) and Shaffer et al. (2016) suggest three possible paths: reformulating the two constructs as subdimensions of a single higher-order construct; developing additional indicators that separate them and demonstrating that those indicators function in discriminant validity tests; or arguing that one construct can be theoretically reduced to the other and removing the reduced construct from the literature. The last path tends to meet resistance, since retiring a construct can feel like questioning the careers built on it, yet scientific progress requires that constructs be revised or, when warranted, abandoned in light of empirical evidence. Podsakoff et al. (2016) stress that clear, mutually distinct conceptual definitions are the most basic step in preventing construct redundancy in the first place.
At the scale level, in fields where several scales for the same construct run in parallel, consolidation is the remedy. The process involves a systematic review that maps the domains the existing scales represent, a meta-analytic assessment of the incremental value each scale offers over the others, and the gradual withdrawal of scales that add no incremental value, with researchers steered toward the evidence-based instruments that remain. Health management is precisely the field where this need is most visible, given the dozens of parallel scales for patient satisfaction, service quality, and health-related quality of life that compromise the comparability of findings.
Special Relevance for Health Management Research
Redundancy carries particular weight in health management because of the field’s own dynamics. In patient-centered measurement, respondent fatigue and the time constraints of clinical settings are pragmatic limits that directly shape scale length. In health-related quality of life measurement, the abundance of parallel scales, including the SF-36, the EQ-5D, the WHOQOL-BREF, and the Nottingham Health Profile, creates both a comparability problem and a concrete instance of scale redundancy. Similar issues appear in patient safety culture measurement, where the AHRQ Hospital Survey on Patient Safety Culture, the Manchester Patient Safety Framework, and the Safety Attitudes Questionnaire are used side by side without systematic evaluation of the incremental value each provides. Because health services are high-risk by nature, the methodological soundness of measurement instruments takes on special importance, and redundancy assessment deserves a permanent place in scale development protocols.
Cross-cultural adaptation introduces a further, often overlooked, twist. Two Turkish items may capture distinct conceptual nuances in their English original, and when translation flattens those nuances the items can converge and produce redundancy that did not exist before. Beaton et al. (2000), in their adaptation protocol, recommend evaluating the conceptual overlap of target-language items in addition to forward-back translation and expert panel review. A familiar example is SERVQUAL, where reliability and assurance are clearly distinct in English but tend to converge semantically in Turkish, so that the limited discriminant validity of these two dimensions in Turkish adaptations is a sign of language-specific redundancy and a cue either to restructure the scale or to develop an original instrument.
Closing Thoughts
Three conclusions stand out. First, redundancy cannot be discussed at a single level of analysis; item overlap, indicator multicollinearity, construct proliferation, and parallel scale multiplicity are related but conceptually distinct phenomena, each with its own diagnostics and remedies. Second, redundancy assessment is qualitative as well as quantitative, and indicators such as the variance inflation factor, the heterotrait-monotrait ratio, and Q3 fall short without expert judgment, because high correlation does not always mean conceptual redundancy and low correlation does not always mean distinctiveness. Third, the field is evolving quickly, with network psychometrics, Bayesian methods (Muthén & Asparouhov, 2012), exploratory structural equation modeling, and artificial-intelligence-assisted item screening extending the reach of classical measurement theory, and these new tools are best positioned as complements to the classical ones rather than replacements. For research communities working in health management, building domain-specific scale development and evaluation protocols around these principles is a contribution worth making, both methodologically and theoretically.
References
Asparouhov, T., & Muthén, B. (2009). Exploratory structural equation modeling. Structural Equation Modeling: A Multidisciplinary Journal, 16(3), 397–438. https://doi.org/10.1080/10705510903008204
Beaton, D. E., Bombardier, C., Guillemin, F., & Ferraz, M. B. (2000). Guidelines for the process of cross-cultural adaptation of self-report measures. Spine, 25(24), 3186–3191. https://doi.org/10.1097/00007632-200012150-00014
Bergkvist, L., & Rossiter, J. R. (2007). The predictive validity of multiple-item versus single-item measures of the same constructs. Journal of Marketing Research, 44(2), 175–184. https://doi.org/10.1509/jmkr.44.2.175
Borsboom, D., & Cramer, A. O. J. (2013). Network analysis: An integrative approach to the structure of psychopathology. Annual Review of Clinical Psychology, 9(1), 91–121. https://doi.org/10.1146/annurev-clinpsy-050212-185608
Boyle, G. J. (1991). Does item homogeneity indicate internal consistency or item redundancy in psychometric scales? Personality and Individual Differences, 12(3), 291–294. https://doi.org/10.1016/0191-8869(91)90115-R
Brown, T. A. (2015). Confirmatory factor analysis for applied research (2nd ed.). Guilford Press.
Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105. https://doi.org/10.1037/h0046016
Cattell, R. B. (1973). Personality and mood by questionnaire. Jossey-Bass.
Christensen, A. P., Garrido, L. E., & Golino, H. (2023). Unique variable analysis: A network psychometrics method to detect local dependence. Multivariate Behavioral Research, 58(6), 1165–1182. https://doi.org/10.1080/00273171.2023.2194606
Cortina, J. M. (1993). What is coefficient alpha? An examination of theory and applications. Journal of Applied Psychology, 78(1), 98–104. https://doi.org/10.1037/0021-9010.78.1.98
Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334. https://doi.org/10.1007/BF02310555
DeVellis, R. F. (2017). Scale development: Theory and applications (4th ed.). Sage Publications.
Diamantopoulos, A., Riefler, P., & Roth, K. P. (2008). Advancing formative measurement models. Journal of Business Research, 61(12), 1203–1218. https://doi.org/10.1016/j.jbusres.2008.01.009
Diamantopoulos, A., & Winklhofer, H. M. (2001). Index construction with formative indicators: An alternative to scale development. Journal of Marketing Research, 38(2), 269–277. https://doi.org/10.1509/jmkr.38.2.269.18845
Drolet, A. L., & Morrison, D. G. (2001). Do we really need multiple-item measures in service research? Journal of Service Research, 3(3), 196–204. https://doi.org/10.1177/109467050133001
Embretson, S. E., & Reise, S. P. (2000). Item response theory for psychologists. Lawrence Erlbaum Associates.
Epskamp, S., Borsboom, D., & Fried, E. I. (2018). Estimating psychological networks and their accuracy: A tutorial paper. Behavior Research Methods, 50(1), 195–212. https://doi.org/10.3758/s13428-017-0862-1
Fornell, C., & Larcker, D. F. (1981). Evaluating structural equation models with unobservable variables and measurement error. Journal of Marketing Research, 18(1), 39–50. https://doi.org/10.1177/002224378101800104
Galesic, M., & Bosnjak, M. (2009). Effects of questionnaire length on participation and indicators of response quality in a web survey. Public Opinion Quarterly, 73(2), 349–360. https://doi.org/10.1093/poq/nfp031
Hair, J. F., Hult, G. T. M., Ringle, C. M., & Sarstedt, M. (2022). A primer on partial least squares structural equation modeling (PLS-SEM) (3rd ed.). Sage Publications.
Henseler, J., Ringle, C. M., & Sarstedt, M. (2015). A new criterion for assessing discriminant validity in variance-based structural equation modeling. Journal of the Academy of Marketing Science, 43(1), 115–135. https://doi.org/10.1007/s11747-014-0403-8
Kline, R. B. (2023). Principles and practice of structural equation modeling (5th ed.). Guilford Press.
Lawshe, C. H. (1975). A quantitative approach to content validity. Personnel Psychology, 28(4), 563–575. https://doi.org/10.1111/j.1744-6570.1975.tb01393.x
Le, H., Schmidt, F. L., Harter, J. K., & Lauver, K. J. (2010). The problem of empirical redundancy of constructs in organizational research: An empirical investigation. Organizational Behavior and Human Decision Processes, 112(2), 112–125. https://doi.org/10.1016/j.obhdp.2010.02.003
Lord, F. M., & Novick, M. R. (1968). Statistical theories of mental test scores. Addison-Wesley.
MacKenzie, S. B., Podsakoff, P. M., & Podsakoff, N. P. (2011). Construct measurement and validation procedures in MIS and behavioral research: Integrating new and existing techniques. MIS Quarterly, 35(2), 293–334. https://doi.org/10.2307/23044045
Marsh, H. W., Morin, A. J. S., Parker, P. D., & Kaur, G. (2014). Exploratory structural equation modeling: An integration of the best features of exploratory and confirmatory factor analysis. Annual Review of Clinical Psychology, 10(1), 85–110. https://doi.org/10.1146/annurev-clinpsy-032813-153700
Muthén, B., & Asparouhov, T. (2012). Bayesian structural equation modeling: A more flexible representation of substantive theory. Psychological Methods, 17(3), 313–335. https://doi.org/10.1037/a0026802
Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
Podsakoff, P. M., MacKenzie, S. B., & Podsakoff, N. P. (2016). Recommendations for creating better concept definitions in the organizational, behavioral, and social sciences. Organizational Research Methods, 19(2), 159–203. https://doi.org/10.1177/1094428115624965
Polit, D. F., Beck, C. T., & Owen, S. V. (2007). Is the CVI an acceptable indicator of content validity? Appraisal and recommendations. Research in Nursing & Health, 30(4), 459–467. https://doi.org/10.1002/nur.20199
Reise, S. P. (2012). The rediscovery of bifactor measurement models. Multivariate Behavioral Research, 47(5), 667–696. https://doi.org/10.1080/00273171.2012.715555
Rönkkö, M., & Cho, E. (2022). An updated guideline for assessing discriminant validity. Organizational Research Methods, 25(1), 6–47. https://doi.org/10.1177/1094428120968614
Shaffer, J. A., DeGeest, D., & Li, A. (2016). Tackling the problem of construct proliferation: A guide to assessing the discriminant validity of conceptually related constructs. Organizational Research Methods, 19(1), 80–110. https://doi.org/10.1177/1094428115598239
Streiner, D. L. (2003). Starting at the beginning: An introduction to coefficient alpha and internal consistency. Journal of Personality Assessment, 80(1), 99–103. https://doi.org/10.1207/S15327752JPA8001_18
Tabachnick, B. G., & Fidell, L. S. (2019). Using multivariate statistics (7th ed.). Pearson.
Yen, W. M. (1984). Effects of local item dependence on the fit and equating performance of the three-parameter logistic model. Applied Psychological Measurement, 8(2), 125–145. https://doi.org/10.1177/014662168400800201
