A novel indirect method for deriving reference intervals through iterative data cleaning guided by self-organizing maps of multi-test patterns (SOM-clean).

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: A novel indirect method for deriving reference intervals through iterative data cleaning guided by self-organizing maps of multi-test patterns (SOM-clean).
Συγγραφείς: Ichihara K; Faculty of Health Sciences, Department of Clinical Laboratory Sciences, Yamaguchi University Graduate School of Medicine, Minami-Kogushi 1-1-1, Ube, 755-0001, Japan. Electronic address: ichihara@yamaguchi-u.ac.jp., Yamashita T; Department of Clinical Pharmacology, Tokai University School of Medicine, Isehara, 259-1193, Japan., Borai A; King Abdullah International Medical Research Center, King Saud Bin Abdulaziz University for Health Sciences, King Abdulaziz Medical City, Ministry of National Guard Health Affairs. P. O. Box 9515, Jeddah 21423, Saudi Arabia.
Πηγή: Computer methods and programs in biomedicine [Comput Methods Programs Biomed] 2026 May 01; Vol. 278, pp. 109279. Date of Electronic Publication: 2026 Feb 05.
Τύπος έκδοσης: Journal Article
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: Elsevier Scientific Publishers Country of Publication: Ireland NLM ID: 8506513 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 1872-7565 (Electronic) Linking ISSN: 01692607 NLM ISO Abbreviation: Comput Methods Programs Biomed Subsets: MEDLINE
Imprint Name(s): Publication: Limerick : Elsevier Scientific Publishers
Original Publication: Amsterdam : Elsevier Science Publishers, c1984-
Ιατρικοί όροι (MeSH): Clustering Algorithms* , Software*, Humans ; Algorithms ; Cluster Analysis ; Data Mining ; Reference Values ; Saudi Arabia
Περίληψη: Background & Objective: Most existing methods for indirectly deriving reference intervals (RIs) from routine laboratory databases use univariate approaches with limited or no rigorous data cleaning. Recognizing the potential of multivariate data-mining strategies, we developed novel software-SOM-clean-that employs self-organizing map (SOM) clustering for iterative exclusion of records exhibiting atypical multi-test patterns.
Methods: We retrieved records for 22 major health-screening tests (HSTs) from a Saudi Arabian laboratory participating in a RI study. After excluding records from frequently tested individuals and those with <10 HST results, 37,285 records remained for analysis. Initial crude RIs were calculated parametrically using a two-parameter Box-Cox power transformation. All transformed values were standardized against these RIs to generate uniform-scale values, so that any result within RI limits fell between ±1.96. The self-organizing map (m × m cells, m = 5-8) was initialized with normal random values, and records were clustered into cells with highest similarity. Cells' patterns were updated by records assigned to each of them. This learning process of the map was repeated until equilibrium. Subsequently, cells exhibiting atypical features were excluded, and RIs were recalculated using records from the remaining cells. This process was repeated iteratively until all RIs stabilized.
Results: Histograms of retrieved results frequently exhibited peaks differing in shape and location from those in the direct study (n = 880). The goodness-of-fit (GOF) of SOM-clean RIs was assessed by skewness, kurtosis, and Kolmogorov-Smirnov test P-values after transformation, as well as by the bias ratio of reference limits compared with the direct study. GOF depended on map size and criteria for identifying atypical cells; the software therefore incorporated an all-inclusive search for optimal conditions referencing the direct study RIs. By using the optimal settings, SOM-clean achieved excellent GOF of RIs simultaneously across nearly all HSTs, indicating conformity of the estimated RIs to the healthy status. In comparison, RIs derived using a representative indirect method (refineR) were generally broader or biased, particularly for tests with highly skewed distributions.
Conclusion: SOM-clean represents a practical and robust parametric tool for estimating RIs indirectly from routine laboratory data employing a novel multivariate-based data cleaning scheme.
(Copyright © 2026 The Author(s). Published by Elsevier B.V. All rights reserved.)
Competing Interests: Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Contributed Indexing: Keywords: Data mining; Laboratory information system; Power transformation; Reference interval; Two-parameter Box-Cox formula; Uniform scale
Entry Date(s): Date Created: 20260220 Date Completed: 20260708 Latest Revision: 20260708
Update Code: 20260708
DOI: 10.1016/j.cmpb.2026.109279
PMID: 41719685
Βάση Δεδομένων: MEDLINE
Περιγραφή
ISSN:1872-7565
DOI:10.1016/j.cmpb.2026.109279