top of page

EXTERNAL VALIDATION OF THE RECURSIVE RELIABILITY EFFECT

  • Writer: Don Gaconnet
    Don Gaconnet
  • Jun 6
  • 13 min read

Convergent Evidence from the McMaster Diagnostic Interview Meta-Analysis

(Duncan et al., 2026) and the Domain Misidentification Simulation



Don L. Gaconnet, CSE III

Founder & Principal Investigator

LifePillar Institute for Structural Identity Sciences

Lake Geneva, Wisconsin

ORCID: 0009-0001-6174-8384

DOI: 10.13140/RG.2.2.20199.61602

Correspondence: don@lifepillar.org


June 2026

Preprint — LifePillar Institute for Structural Identity Sciences


Copyright © Don L. Gaconnet, June 2026. All rights reserved.



Abstract

The Recursive Reliability Effect (RRE) predicts that human systems under structural load cannot accurately self-assess, that the degradation of self-assessment accuracy is recursive rather than linear, and that any assessment methodology beginning from the subject’s self-report inherits this structural error as primary input data (Gaconnet, 2026a). A direct corollary of this prediction is that interview-based assessment instruments—which depend on the subject’s verbal narrative as their primary data source—will demonstrate degraded reliability proportional to the structural load carried by the subject population.


On May 28, 2026, Duncan et al. published a meta-analysis in JAMA Network Open examining the test-retest reliability of standardized diagnostic interviews (SDIs) across 57 studies, 26 countries, and 8,146 participants. The finding: SDIs demonstrated only moderate pooled test-retest reliability (κ = 0.69), with substantial heterogeneity across disorder categories and reduced consistency for conditions relying on subjective experience.


This paper presents the Duncan et al. finding as independent external validation of a specific RRE prediction: that interview-based instruments will produce inconsistent results because they depend on a data source—the subject’s self-report—that is structurally unreliable under load. The convergence is analyzed as a dual-source error problem: interview-based assessment carries two independent, simultaneously operating sources of structural error—tool inconsistency (Duncan et al., 2026) and source unreliability (Gaconnet, 2026b).


Implications for professional risk assessment in high-stakes fiduciary, investment, and organizational contexts are discussed. The case for assessment methodology that bypasses both error sources through independent biometric measurement is presented as a scientific implication consistent with both findings. Specific instrumentation methodology is proprietary.


Keywords: recursive reliability effect, self-assessment degradation, diagnostic interview reliability, test-retest reliability, dual-source error, structural load, inverse reliability, Monte Carlo validation, independent measurement, key person risk, executive assessment, cognitive due diligence


1. Introduction

1.1 The Prediction

The Recursive Reliability Effect (Gaconnet, 2026a) establishes that self-assessment accuracy degrades as a function of structural load, that this degradation is recursive rather than linear (each failed self-assessment compounds into the next cycle), and that the degradation is most severe in the population where accurate assessment carries the highest stakes—a property termed inverse reliability.


A direct corollary follows: any assessment instrument that depends on the subject’s self-report as primary input data will inherit the RRE as a structural error embedded in its measurement. Interview-based instruments—including standardized diagnostic interviews, behavioral interviews, structured clinical interviews, and psychometric assessments administered through verbal response—all share this dependency. The RRE predicts that these instruments will demonstrate degraded reliability, and that the degradation will be most pronounced for conditions that rely most heavily on the subject’s subjective experience rather than observable behaviors.


This prediction was published before the Duncan et al. (2026) meta-analysis appeared. The present paper documents the convergence.


1.2 The Independent Confirmation

Duncan et al. (2026) conducted a systematic review and meta-analysis of test-retest reliability of standardized diagnostic interviews (SDIs) for common adult psychiatric disorders, published in JAMA Network Open on May 28, 2026. The study analyzed 57 studies involving 8,146 participants from 26 countries. The pooled estimate of SDI test-retest reliability was κ = 0.69—moderate by conventional standards. Reliability varied substantially across disorder categories, ranging from κ = 0.55 for nonaffective psychoses to κ = 0.74 for bipolar disorders among mental disorders, and from κ = 0.59 for hallucinogen use to κ = 0.81 for opioid use among substance use disorders.


The finding that reliability was higher for substance use disorders (κ = 0.72) than for mental disorders (κ = 0.65) is structurally predicted by the RRE. Substance use disorders involve more observable behavioral markers and more concrete temporal anchors. Mental disorders—particularly those relying on subjective experience, internal states, and self-reported symptom severity—depend more heavily on the subject’s self-report. The RRE predicts that the greater the dependency on self-report, the greater the reliability degradation. Duncan et al. confirmed this pattern independently.


1.3 Scope and Proprietary Disclosure

This paper presents the convergent evidence analysis and the dual-source error framework as a theoretical contribution. The formal derivation of the Recursive Reliability Effect from the Law of Recursion and the Law of Obligated Systems within the Structural Identity Sciences framework is published in the cited works (Gaconnet, 2026a; 2026c; 2026d) and is not reproduced here. The specific instrumentation that implements the recursion-breaking principle—the Structural Identity Profiler, its scoring methodology, its biometric integration architecture, and its assessment protocols—is proprietary to the LifePillar Institute for Structural Identity Sciences and available through professional services engagement (dongaconnet.com).


2. The McMaster Finding: Tool Inconsistency

The Duncan et al. (2026) meta-analysis tested the most fundamental property of a measurement instrument: does it produce the same result when applied to the same subject under similar conditions? For standardized diagnostic interviews—the instruments the clinical and assessment fields have treated as the gold standard for psychiatric diagnosis—the answer is: not consistently.


The pooled kappa of 0.69 is conventionally classified as “moderate” agreement (Landis & Koch, 1977). For context, a kappa of 1.0 indicates perfect agreement, and a kappa of 0.40 or below indicates poor to fair agreement. A pooled kappa of 0.69 means that the same individual, assessed with the same standardized interview under similar conditions, has a meaningful probability of receiving a different diagnostic classification on reassessment.


The heterogeneity finding is equally significant. Reliability was not uniformly moderate; it varied substantially by disorder category, with conditions relying on observable behaviors and concrete timelines showing higher consistency than conditions requiring self-reported subjective experience. This pattern—lower reliability for assessments more dependent on self-report—is the precise pattern the RRE predicts.


The senior author stated the finding directly: the field should reconsider treating standardized diagnostic interviews as a gold standard of assessment (McMaster University press release, May 28, 2026). The study recommended combining standardized tools with contextual and phenomenological information, rather than relying on a single interview.


3. The Domain Misidentification Finding: Source Unreliability

A separate body of evidence documents not the tool’s inconsistency but the source’s unreliability. A 10,000-case Monte Carlo simulation (Gaconnet, 2026b) modeled a near-capacity population—individuals operating under sustained structural obligation that approaches or exceeds their system’s carrying capacity—and measured the gap between self-reported structural state and instrument-measured structural reality.


The simulation produced the following headline findings with 95% Wilson score confidence intervals:


Domain Mismatch Rate: 81.4% (95% CI: 80.7–82.2%). Four out of five subjects misidentified which domain their primary structural failure lived in. The error was systematic and directional—subjects consistently displaced the problem into domains that the performance architecture could manage, away from domains where the structural failure actually resided.

Depth Minimization Rate: 73.0% (95% CI: 72.1–73.9%). Nearly three out of four subjects placed their problem closer to the surface than the instrument measured.


Compound Risk Rate: 61.1% (95% CI: 60.1–62.0%). More than six out of ten subjects were simultaneously wrong about both the domain and the depth of their structural failure.


The most consequential finding is inverse reliability: self-report accuracy degraded as structural severity increased. The deepest structural failures produced the most extreme self-report distortion, with depth minimization rates exceeding 94% in the most severe categories. The cases where accurate assessment would have the greatest impact are the cases where self-report is most structurally wrong.


The simulation methodology—Monte Carlo with 10,000 cases—follows the same validation standard used in aerospace structural load analysis and pharmaceutical pharmacokinetic modeling. The confidence intervals at N = 10,000 are tight: all headline findings fall within ±2.3% of the stated rate.


Limitations: The simulation is based on a population model derived from the author’s body of work, not empirical measurement of living subjects. Empirical validation with the target population is the next research phase. The simulation validates the instrument’s logic, computation, and classification system. It does not constitute empirical prevalence validation. This limitation is stated in every publication reporting these findings.


4. The Dual-Source Error Framework

4.1 Two Independent Sources of Error

The convergence of the Duncan et al. (2026) and Gaconnet (2026b) findings produces a structural insight that has not been articulated in the assessment methodology literature: interview-based assessment carries two independent, simultaneously operating sources of structural error.


Source 1 — Tool inconsistency. The diagnostic interview does not produce consistent results when applied to the same subject under similar conditions (Duncan et al., 2026). The instrument itself introduces measurement variability independent of the subject’s state.


Source 2 — Source unreliability. The subject producing the input data for the interview cannot accurately self-report under the load conditions that would make the report matter (Gaconnet, 2026a; 2026b). The data entering the instrument is structurally wrong at the population level.


These error sources are independent. Tool inconsistency operates even when the subject’s self-report is accurate—the instrument introduces its own variability through framing effects, interviewer variation, and the inherent ambiguity of mapping subjective experience onto categorical criteria. Source unreliability operates even with a perfectly consistent tool—the data entering the tool is wrong because the system under load cannot accurately assess its own state.


The dual-source error structure means that improving either source in isolation does not resolve the measurement problem. A more reliable interview administered to an unreliable source still produces unreliable output. A reliable source assessed with an inconsistent tool still produces inconsistent results. The structural correction requires bypassing both sources simultaneously.


4.2 The Prediction-Confirmation Relationship

The RRE (Gaconnet, 2026a) was formalized and published before the Duncan et al. (2026) meta-analysis appeared. The RRE predicts that interview-based instruments will demonstrate degraded reliability because they depend on self-report as primary data. Duncan et al. confirmed that standardized diagnostic interviews demonstrate degraded reliability. The prediction preceded the confirmation.


Furthermore, the RRE predicts that the degradation will be most severe for conditions relying most heavily on subjective experience. Duncan et al. confirmed that reliability was lower for mental disorders (κ = 0.65) than for substance use disorders (κ = 0.72), and that the authors attributed this pattern to the greater subjectivity involved in assessing mental disorders. The structural explanation provided by the RRE—that self-report is least reliable when the subject matter most depends on the subject’s internal state—is consistent with the observed pattern.


This constitutes a prediction-confirmation relationship between an independently derived theoretical framework and a large-scale empirical meta-analysis. The RRE made a testable prediction about the behavior of interview-based instruments. The prediction was confirmed by an independent research program that had no knowledge of the RRE and no contact with its author.


5. Implications for Professional Risk Assessment

The dual-source error framework has direct implications for any professional context where assessment of a human system under load determines high-stakes decisions.

In private equity leadership due diligence, the behavioral interview is the primary assessment instrument for evaluating founder and executive capacity. Seventy-three percent of PE-backed CEOs are replaced during the hold period, with 55% of that turnover unplanned (AlixPartners, 2026; Heidrick & Struggles, 2026). The behavioral interview used to assess founder readiness shares the same structural dependency the Duncan et al. study documented—the founder’s verbal narrative. The dual-source error framework predicts that this interview is simultaneously inconsistent (tool error) and operating on unreliable input data (source error) in the near-capacity founder population.


In fiduciary and legal risk assessment, attorneys evaluating client capacity rely on diagnostic evaluations built on the same standardized interview methodology. The standard of care in fiduciary practice is independent verification—the forensic accountant reads the books, not the CFO’s description of the books. The diagnostic interview does not meet this standard of independence because it depends on the client’s verbal narrative as its primary data source.

In organizational governance, boards and family offices overseeing executive performance rely on assessment tools rooted in interview methodology for succession planning, fitness evaluations, and key person risk assessment. The dual-source error framework predicts that the executive whose structural integrity matters most to the organization is the executive whose self-report is most structurally unreliable—the inverse reliability property.


6. The Case for Assessment That Bypasses Both Error Sources

If interview-based assessment carries two independent sources of structural error—tool inconsistency and source unreliability—then the structural correction is assessment methodology that depends on neither. This requires an instrument that (a) does not use the interview as its primary data collection method, and (b) does not depend on the subject’s self-report as primary input data.


The human factors workload assessment literature arrived at this conclusion decades ago in a different domain. Hart and Staveland (1988) and Webster et al. (2018) recommend external physiological measures—EEG, heart rate variability, galvanic skin response—because self-report is unreliable under operational stress. The cognitive load literature (Sweller, Ayres, & Kalyuga, 2011) independently confirmed that metacognitive monitoring degrades under load, supporting the same conclusion: when the internal state matters most, the internal report is least reliable, and external measurement is required.


The Structural Identity Profiler (Gaconnet, 2026d) implements this principle through four-channel biometric integration: EEG, heart-rate variability, facial affect analysis, and voice prosody analysis. The instrument does not begin from the subject’s verbal narrative. It reads the structural state of the system directly through independent biometric channels that bypass the conscious performance layer. The specific architecture, scoring methodology, and assessment protocols of the instrument are proprietary. The principle underlying its design—that accurate assessment under load requires external measurement independent of self-report—is a scientific implication confirmed by the convergent evidence presented in this paper and consistent with the established literature across cognitive load theory, human factors, and clinical self-assessment research.


The instrument has been validated across 28,400 simulated cases in three independent validation programs: the PE Divergence Simulation (10,000 cases), the Clinical Trial Simulations (8,400 cases), and the Organizational Monte Carlo (10,000 cases). The validation standard is engineering, not clinical. Empirical validation with living subjects in the target population is the next research phase.


7. Limitations

The Duncan et al. (2026) meta-analysis examined clinical diagnostic interviews in psychiatric populations. The application of its findings to executive assessment and leadership due diligence contexts represents an inference from the shared structural dependency—reliance on the subject’s verbal narrative—rather than a direct experimental confirmation in the executive population. The inference is supported by the structural analysis but has not been empirically validated through a direct comparison study.


The domain misidentification findings (Gaconnet, 2026b) are based on a 10,000-case Monte Carlo simulation using a population model, not empirical measurement of living subjects. The simulation validates the instrument’s logic and classification system but does not constitute empirical prevalence validation. Specific percentages may shift when the model is calibrated against empirical data from the initial assessment cohort.


The dual-source error framework is presented as a theoretical contribution. The claim that both sources of error operate simultaneously and independently has not been tested in a single experimental design that measures both tool inconsistency and source unreliability within the same assessment context. Such a study would constitute a direct empirical test of the framework and is proposed as a future research direction.


The convergence between the RRE prediction and the Duncan et al. confirmation is presented as consistent with the prediction-confirmation model, not as formal hypothesis testing. The RRE predicted a qualitative pattern (degraded reliability correlated with self-report dependency); Duncan et al. confirmed a qualitative pattern consistent with this prediction. A stronger test would require quantitative predictions about specific kappa values, which the RRE does not make in its current form.


8. Falsification Criteria

The dual-source error framework and the prediction-confirmation relationship presented in this paper are subject to the following falsification criteria:


Criterion 1. If a future meta-analysis demonstrates that standardized diagnostic interview reliability is not correlated with the degree of self-report dependency—that is, if conditions requiring maximal subjective self-report show equal or higher reliability than conditions with observable behavioral markers—the RRE prediction about interview reliability is falsified.


Criterion 2. If empirical measurement of the near-capacity executive population demonstrates that self-report accuracy does NOT degrade as structural severity increases (i.e., inverse reliability is not observed in living subjects), the inverse reliability property of the RRE is falsified.


Criterion 3. If an instrument-based assessment that does not rely on self-report or interview data produces LOWER accuracy than interview-based assessment in the near-capacity population, the claim that bypassing both error sources improves measurement is falsified.


Criterion 4. If the domain misidentification rate in empirical measurement of the target population falls below 50%—that is, if self-report in this population is more often correct than incorrect about domain—the specific rates reported in the simulation will require revision, though the qualitative finding (self-report is unreliable) would persist unless the rate falls below random chance.


Each criterion is genuinely testable. Empirical measurement of the target population—the planned next phase of research—will produce data relevant to Criteria 2, 3, and 4.


9. Conclusion

The Duncan et al. (2026) meta-analysis provides independent external validation of a prediction derived from the Recursive Reliability Effect: interview-based assessment instruments demonstrate degraded reliability, and the degradation follows the structural pattern predicted by the RRE—greatest for conditions most dependent on self-report.


The convergence produces the dual-source error framework: interview-based assessment carries two independent, simultaneously operating sources of structural error. The tool does not produce consistent results (Duncan et al., 2026). The source cannot produce accurate input data under load (Gaconnet, 2026a; 2026b). Both sources of error are independent. Both operate simultaneously. Any assessment methodology that depends on the interview as its primary instrument is subject to both.


The structural correction requires assessment that bypasses both error sources. The human factors literature, the cognitive load literature, and the clinical self-assessment literature converge on the same implication: when the internal state matters most, external measurement is required. The Structural Identity Profiler implements this principle through four independent biometric channels that read the system’s structural state without depending on the subject’s verbal narrative.


The dual-source error framework, the prediction-confirmation relationship, and the falsification criteria presented in this paper constitute a testable theoretical contribution to the assessment methodology literature. Empirical validation with the target population is the next research phase.



References

AlixPartners. (2026). Eleventh annual private equity leadership survey: Expectation and execution. AlixPartners.

Davis, D. A., Mazmanian, P. E., Fordis, M., Van Harrison, R., Thorpe, K. E., & Perrier, L. (2006). Accuracy of physician self-assessment compared with observed measures of competence. JAMA, 296(9), 1094–1102.

Duncan, L. J., Xie, W., et al. (2026). Test-retest reliability of standardized diagnostic interviews for common adult psychiatric disorders. JAMA Network Open. DOI: 10.1001/jamanetworkopen.2026.15039.

Ehrlinger, J., Johnson, K., Banner, M., Dunning, D., & Kruger, J. (2008). Why the unskilled are unaware: Further explorations of (absent) self-insight among the incompetent. Organizational Behavior and Human Decision Processes, 105(1), 98–121.

Eva, K. W., & Regehr, G. (2005). Self-assessment in the health professions: A reformulation and research agenda. Academic Medicine, 80(10), S46–S54.

Felitti, V. J., et al. (1998). Relationship of childhood abuse and household dysfunction to many of the leading causes of death in adults: The adverse childhood experiences (ACE) study. American Journal of Preventive Medicine, 14(4), 245–258.

Gaconnet, D. L. (2026a). The Recursive Reliability Effect: Self-assessment degradation in human systems under structural load. LifePillar Institute for Structural Identity Sciences. DOI: 10.17605/OSF.IO/MVYZT.

Gaconnet, D. L. (2026b). Self-report unreliability under structural load: Domain misidentification in high-obligation systems. SSRN Working Paper.

Gaconnet, D. L. (2026c). Cognitive due diligence: Independent structural measurement as the missing pillar in private equity leadership assessment. SSRN Working Paper.

Gaconnet, D. L. (2026d). The Structural Identity Profiler: Instrument architecture, validation portfolio, and professional services framework. SSRN Working Paper.

Hart, S. G., & Staveland, L. E. (1988). Development of NASA-TLX: Results of empirical and theoretical research. Advances in Psychology, 52, 139–183.

Heidrick & Struggles. (2026). Closing the leadership gap in private equity. Heidrick & Struggles.

Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174.

Paulhus, D. L. (1991). Measurement and control of response bias. In J. P. Robinson, P. R. Shaver, & L. S. Wrightsman (Eds.), Measures of personality and social psychological attitudes (pp. 17–59). Academic Press.

Podsakoff, P. M., MacKenzie, S. B., Lee, J. Y., & Podsakoff, N. P. (2003). Common method biases in behavioral research: A critical review. Journal of Applied Psychology, 88(5), 879–903.

Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285.

Sweller, J., Ayres, P., & Kalyuga, S. (2011). Cognitive load theory. Springer.

Webster, C. S., et al. (2018). Psychophysiological measurement of cognitive load in anaesthesia: A systematic review. British Journal of Anaesthesia, 120(6), 1145–1158.



———


Don L. Gaconnet, CSE III

Cognitive Systems Engineer III

Founder & Principal Investigator, LifePillar Institute for Structural Identity Sciences

ORCID: 0009-0001-6174-8384 · SSRN: 7657314

Lake Geneva, Wisconsin · don@lifepillar.org


© 2026 Don L. Gaconnet. All rights reserved. The Structural Identity Profiler, the diagnostic engine, the nine-law architecture, and all associated methodologies are proprietary trade secrets of Don L. Gaconnet and the LifePillar Institute for Structural Identity Sciences. No part of the instrument’s operational architecture may be reproduced, reverse-engineered, or derived from this publication.


 
 
 

Comments


Commenting on this post isn't available anymore. Contact the site owner for more info.

© 2026 Don L. Gaconnet. All Rights Reserved.
LifePillar Institute for Structural Identity Sciences
This page constitutes the canonical source for Structural Identity Sciences (formerly published as Recursive Sciences) and its component frameworks: Echo-Excess Principle (EEP), Cognitive Field Dynamics (CFD), Collapse Harmonics Theory (CHT), and Identity Collapse Therapy (ICT).
Founder: Don L. Gaconnet | ORCID: 0009-0001-6174-8384 | DOI: 10.5281/zenodo.15758805
Academic citation required for all derivative work.

bottom of page