Academic Journal
Machine Learning–Based Prediction of Breast Cancer in Women: Insights From Feature Selection of Clinical and Lifestyle Data.
| Title: | Machine Learning–Based Prediction of Breast Cancer in Women: Insights From Feature Selection of Clinical and Lifestyle Data. |
|---|---|
| Authors: | Allahqoli, Leila, Behzadi, Mohammad Hassan, Aghamohammadi, Seyedeh Zahra, Hakimi, Sevil, Salehiniya, Hamid, Fallahi, Arezoo, Rahmani, Azam, Shahabinia, Zahra, Mazidimoradi, Afrooz, Momenimovahed, Zohre, Ghiyasvand, Mohammadmatin, Ribeiro, Ivana |
| Source: | International Journal of Breast Cancer; 5/18/2026, Vol. 2026, p1-11, 11p |
| Subject Terms: | Machine learning, Feature selection, Breast tumors, Medical databases, Risk assessment, Health behavior, Random forest algorithms, Prediction models |
| Abstract: | Objective: This study is aimed at developing and evaluating a machine learning–based model for breast cancer classification using integrated clinical, demographic, reproductive, and lifestyle data. Methods: A retrospective machine learning framework was developed using data from a case–control study conducted in Tehran. The dataset included demographic, clinical, reproductive, lifestyle, and screening‐related variables. Data preprocessing was performed within machine learning pipelines to ensure data quality and prevent data leakage. Duplicate records and noninformative variables were removed. Missing values were imputed using median values for numerical variables and the most frequent value for categorical variables. Categorical features were encoded appropriately, and Min–Max normalization was applied where required. Feature selection was conducted using mutual information (MI) and analysis of variance (ANOVA) within a stratified cross‐validation framework. The data were split into training (80%) and test (20%) sets. Several supervised learning algorithms, including Gaussian Naive Bayes (GNB), K‐nearest neighbors (KNN), decision tree (DT), random forest (RF), support vector machine (SVM), logistic regression (LR), and artificial neural network (ANN), were trained and evaluated. Model performance was assessed using accuracy, precision, recall (sensitivity), F1‐score, and receiver operating characteristic–area under the curve (ROC‐AUC), with stratified 5‐fold cross‐validation and final evaluation on an independent test set. Results: Significant differences were observed between breast cancer patients and healthy controls across multiple demographic and clinical variables. Patients were generally older and more likely to be widowed, belong to higher socioeconomic classes, and be housewives, whereas higher education levels and employment were more frequent among healthy individuals (p < 0.001). Reproductive factors, including age at first marriage and breastfeeding duration, also showed significant differences. Feature selection reduced 414 initial variables to 40 key predictors. The most influential features included genetic factors (BRCA1/2 mutations and family history), reproductive and hormonal characteristics (age at menarche, menopause, and infertility), lifestyle behaviors (dietary patterns and physical activity), anthropometric measures (BMI and weight at age 30), and screening‐related variables (mammography, ultrasound, and biopsy). All models demonstrated strong and stable performance with minimal differences between cross‐validation and test results, indicating good generalization. RF achieved the highest performance (accuracy: 0.9897, precision: 0.9946, recall: 0.9840, F1‐score: 0.9892), followed by SVM and LR, whereas ANN showed the lowest overall performance. Conclusion: Machine learning models can effectively classify breast cancer using multidimensional patient data. Ensemble methods, particularly RF, demonstrated superior accuracy and robustness, highlighting their ability to capture complex nonlinear relationships. The identified predictors are consistent with established clinical and epidemiological risk factors, supporting the validity of the proposed models. These findings suggest that machine learning approaches hold strong potential for personalized risk assessment and early detection of breast cancer; however, external validation across diverse populations is necessary to confirm generalizability. [ABSTRACT FROM AUTHOR] |
| Copyright of International Journal of Breast Cancer is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Biomedical Index |
| FullText | Links: – Type: other Text: Availability: 0 |
|---|---|
| Header | DbId: edm DbLabel: Biomedical Index An: 193836928 RelevancyScore: 1061 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 1060.76306152344 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Machine Learning–Based Prediction of Breast Cancer in Women: Insights From Feature Selection of Clinical and Lifestyle Data. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Allahqoli%2C+Leila%22">Allahqoli, Leila</searchLink><br /><searchLink fieldCode="AR" term="%22Behzadi%2C+Mohammad+Hassan%22">Behzadi, Mohammad Hassan</searchLink><br /><searchLink fieldCode="AR" term="%22Aghamohammadi%2C+Seyedeh+Zahra%22">Aghamohammadi, Seyedeh Zahra</searchLink><br /><searchLink fieldCode="AR" term="%22Hakimi%2C+Sevil%22">Hakimi, Sevil</searchLink><br /><searchLink fieldCode="AR" term="%22Salehiniya%2C+Hamid%22">Salehiniya, Hamid</searchLink><br /><searchLink fieldCode="AR" term="%22Fallahi%2C+Arezoo%22">Fallahi, Arezoo</searchLink><br /><searchLink fieldCode="AR" term="%22Rahmani%2C+Azam%22">Rahmani, Azam</searchLink><br /><searchLink fieldCode="AR" term="%22Shahabinia%2C+Zahra%22">Shahabinia, Zahra</searchLink><br /><searchLink fieldCode="AR" term="%22Mazidimoradi%2C+Afrooz%22">Mazidimoradi, Afrooz</searchLink><br /><searchLink fieldCode="AR" term="%22Momenimovahed%2C+Zohre%22">Momenimovahed, Zohre</searchLink><br /><searchLink fieldCode="AR" term="%22Ghiyasvand%2C+Mohammadmatin%22">Ghiyasvand, Mohammadmatin</searchLink><br /><searchLink fieldCode="AR" term="%22Ribeiro%2C+Ivana%22">Ribeiro, Ivana</searchLink> – Name: TitleSource Label: Source Group: Src Data: International Journal of Breast Cancer; 5/18/2026, Vol. 2026, p1-11, 11p – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Feature+selection%22">Feature selection</searchLink><br /><searchLink fieldCode="DE" term="%22Breast+tumors%22">Breast tumors</searchLink><br /><searchLink fieldCode="DE" term="%22Medical+databases%22">Medical databases</searchLink><br /><searchLink fieldCode="DE" term="%22Risk+assessment%22">Risk assessment</searchLink><br /><searchLink fieldCode="DE" term="%22Health+behavior%22">Health behavior</searchLink><br /><searchLink fieldCode="DE" term="%22Random+forest+algorithms%22">Random forest algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction+models%22">Prediction models</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Objective: This study is aimed at developing and evaluating a machine learning–based model for breast cancer classification using integrated clinical, demographic, reproductive, and lifestyle data. Methods: A retrospective machine learning framework was developed using data from a case–control study conducted in Tehran. The dataset included demographic, clinical, reproductive, lifestyle, and screening‐related variables. Data preprocessing was performed within machine learning pipelines to ensure data quality and prevent data leakage. Duplicate records and noninformative variables were removed. Missing values were imputed using median values for numerical variables and the most frequent value for categorical variables. Categorical features were encoded appropriately, and Min–Max normalization was applied where required. Feature selection was conducted using mutual information (MI) and analysis of variance (ANOVA) within a stratified cross‐validation framework. The data were split into training (80%) and test (20%) sets. Several supervised learning algorithms, including Gaussian Naive Bayes (GNB), K‐nearest neighbors (KNN), decision tree (DT), random forest (RF), support vector machine (SVM), logistic regression (LR), and artificial neural network (ANN), were trained and evaluated. Model performance was assessed using accuracy, precision, recall (sensitivity), F1‐score, and receiver operating characteristic–area under the curve (ROC‐AUC), with stratified 5‐fold cross‐validation and final evaluation on an independent test set. Results: Significant differences were observed between breast cancer patients and healthy controls across multiple demographic and clinical variables. Patients were generally older and more likely to be widowed, belong to higher socioeconomic classes, and be housewives, whereas higher education levels and employment were more frequent among healthy individuals (p < 0.001). Reproductive factors, including age at first marriage and breastfeeding duration, also showed significant differences. Feature selection reduced 414 initial variables to 40 key predictors. The most influential features included genetic factors (BRCA1/2 mutations and family history), reproductive and hormonal characteristics (age at menarche, menopause, and infertility), lifestyle behaviors (dietary patterns and physical activity), anthropometric measures (BMI and weight at age 30), and screening‐related variables (mammography, ultrasound, and biopsy). All models demonstrated strong and stable performance with minimal differences between cross‐validation and test results, indicating good generalization. RF achieved the highest performance (accuracy: 0.9897, precision: 0.9946, recall: 0.9840, F1‐score: 0.9892), followed by SVM and LR, whereas ANN showed the lowest overall performance. Conclusion: Machine learning models can effectively classify breast cancer using multidimensional patient data. Ensemble methods, particularly RF, demonstrated superior accuracy and robustness, highlighting their ability to capture complex nonlinear relationships. The identified predictors are consistent with established clinical and epidemiological risk factors, supporting the validity of the proposed models. These findings suggest that machine learning approaches hold strong potential for personalized risk assessment and early detection of breast cancer; however, external validation across diverse populations is necessary to confirm generalizability. [ABSTRACT FROM AUTHOR] – Name: Abstract Label: Group: Ab Data: <i>Copyright of International Journal of Breast Cancer is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edm&AN=193836928 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1155/ijbc/8510185 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 11 StartPage: 1 Subjects: – SubjectFull: Machine learning Type: general – SubjectFull: Feature selection Type: general – SubjectFull: Breast tumors Type: general – SubjectFull: Medical databases Type: general – SubjectFull: Risk assessment Type: general – SubjectFull: Health behavior Type: general – SubjectFull: Random forest algorithms Type: general – SubjectFull: Prediction models Type: general Titles: – TitleFull: Machine Learning–Based Prediction of Breast Cancer in Women: Insights From Feature Selection of Clinical and Lifestyle Data. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Allahqoli, Leila – PersonEntity: Name: NameFull: Behzadi, Mohammad Hassan – PersonEntity: Name: NameFull: Aghamohammadi, Seyedeh Zahra – PersonEntity: Name: NameFull: Hakimi, Sevil – PersonEntity: Name: NameFull: Salehiniya, Hamid – PersonEntity: Name: NameFull: Fallahi, Arezoo – PersonEntity: Name: NameFull: Rahmani, Azam – PersonEntity: Name: NameFull: Shahabinia, Zahra – PersonEntity: Name: NameFull: Mazidimoradi, Afrooz – PersonEntity: Name: NameFull: Momenimovahed, Zohre – PersonEntity: Name: NameFull: Ghiyasvand, Mohammadmatin – PersonEntity: Name: NameFull: Ribeiro, Ivana IsPartOfRelationships: – BibEntity: Dates: – D: 18 M: 05 Text: 5/18/2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 20903170 Numbering: – Type: volume Value: 2026 Titles: – TitleFull: International Journal of Breast Cancer Type: main |
| ResultId | 1 |