Academic Journal

SHAP as a Data Reduction Technique for Highly Imbalanced Big Data.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: SHAP as a Data Reduction Technique for Highly Imbalanced Big Data.
Συγγραφείς: Hancock III, John T., Bauder, Richard A., Khoshgoftaar, Taghi M.
Πηγή: International Journal on Artificial Intelligence Tools; Jun-Aug2025, Vol. 34 Issue 4/5, p1-27, 27p
Θεματικοί όροι: Data reduction, Feature selection, Anomaly detection (Computer security), Outlier detection, Machine learning, Data quality, Big data
Περίληψη: Fraud detection through the classification of highly imbalanced Big Data is an exciting area of Machine Learning research. On the one hand, in certain fraud detection application domains, the use of One-Class classifiers is an overlooked opportunity. On the other hand, for researchers faced with the task of building Machine Learning models for identifying fraud, when only legitimate transaction data is available, One-Class Classifiers are indispensable. We investigate the efficacy of SHapley Additive exPlanations (SHAP) as a feature selection technique for One-Class classification tasks. In this study we utilize authentic data from the Credit Card fraud and Medicare insurance fraud application domains. Our contribution is to show that researchers can use SHAP in conjunction with One-Class Classifiers to do feature selection on highly imbalanced datasets, and then build models, with the selected features, that yield performance similar to, or better than, models built using all features. Our results in Big Medicare data fraud detection show that an over 90% data reduction through feature selection can nevertheless coincide with the best performance in terms of Area under the Precision Recall Curve. [ABSTRACT FROM AUTHOR]
Copyright of International Journal on Artificial Intelligence Tools is the property of World Scientific Publishing Company and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index