Academic Journal
Automated Multi-Label Adverse drug Reaction Extraction from Malaysia's National Pharmacovigilance Narratives.
| Τίτλος: | Automated Multi-Label Adverse drug Reaction Extraction from Malaysia's National Pharmacovigilance Narratives. |
|---|---|
| Συγγραφείς: | Lim SG; Department of Information Systems, Faculty of Computer Science and Information Technology, Universiti Malaya, Kuala Lumpur, 50603, Malaysia., Ismail MA; Department of Information Systems, Faculty of Computer Science and Information Technology, Universiti Malaya, Kuala Lumpur, 50603, Malaysia. maizatul@um.edu.my. |
| Πηγή: | Journal of medical systems [J Med Syst] 2026 Jul 31; Vol. 50 (1). Date of Electronic Publication: 2026 Jul 31. |
| Τύπος έκδοσης: | Journal Article |
| Γλώσσα: | English |
| Στοιχεία περιοδικού: | Publisher: Kluwer Academic/Plenum Publishers Country of Publication: United States NLM ID: 7806056 Publication Model: Electronic Cited Medium: Internet ISSN: 1573-689X (Electronic) Linking ISSN: 01485598 NLM ISO Abbreviation: J Med Syst Subsets: MEDLINE |
| Imprint Name(s): | Publication: 1999- : New York, NY : Kluwer Academic/Plenum Publishers Original Publication: New York, Plenum Press. |
| Ιατρικοί όροι (MeSH): | Adverse Drug Reaction Reporting Systems*/organization & administration , Drug-Related Side Effects and Adverse Reactions*/epidemiology , Data Mining*/methods , Pharmacovigilance* , Natural Language Processing*, Malaysia ; Humans ; Machine Learning |
| Περίληψη: | Extracting multiple adverse drug reaction (ADR) terms from unstructured narratives remains challenging, particularly under severe label imbalance that limits the detection of rare ADRs. This study aimed to develop and evaluate a multi-label natural language processing framework for automated ADR extraction within Malaysia's national pharmacovigilance reporting system. We evaluated classical machine learning and transformer-based models within a common framework, incorporating domain-specific preprocessing and imbalance-aware optimization strategies to improve the detection of rare ADRs. The framework was applied to 28,980 real-world ADR narratives annotated with 80 Medical Dictionary for Regulatory Activities (MedDRA) Preferred Terms (PTs) from the Skin and Subcutaneous Tissue Disorders System Organ Class (SOC). Performance was evaluated using micro-F1, macro-F1, precision, recall, and label coverage using a predefined 80/20 train-test evaluation strategy, complemented by pharmacist pilot review. Transformer-based models achieved the strongest performance, with the augmented model attaining a micro-F1 of 0.88, macro-F1 of 0.55, recall of 0.88, and the highest label coverage, correctly predicting 60 of 80 PTs (75%). However, a 70/10/20 sensitivity analysis on the improved transformer model yielded lower micro-F1 and precision, highlighting the impact of further data partitioning on model performance. Improved classical models also showed substantial gains over their baseline models, particularly in micro-F1, macro-F1, recall, and label coverage, consistent with improved detection of rare ADRs. During the pilot review, pharmacists accepted the model-preselected PTs without modification in 19 of 25 narratives (76%), while additional PTs were added in the remaining cases. Although the study was limited to a single SOC and one national pharmacovigilance database, the findings demonstrate that optimization strategies, domain-specific feature enrichment, and targeted augmentation can substantially improve multi-label ADR classification under severe imbalance. These findings support further evaluation of AI-assisted pharmacovigilance with pharmacist oversight in real-world practice. (© 2026. The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature.) |
| Competing Interests: | Declarations. Ethics Approval and Consent to Participate: The study was approved by the Medical Research and Ethics Committee, Ministry of Health Malaysia, approval number NMRR ID-25-01404-FFD. All narratives were fully anonymized before researcher access, with no patient identifiers retained. The analysis complied with NPRA confidentiality requirements, Malaysian regulatory standards, and institutional data-sharing agreements. Informed consent was not required as the study involved secondary analysis of fully anonymized PV reports. Competing Interests: The authors declare no competing interests. Consent for Publication: Not applicable. Clinical Trial Number: Not applicable. Generative AI and AI-Assisted Technologies: During the preparation of this work, the authors used GPT-4 via ChatGPT to generate synthetic adverse drug reaction narratives for training data expansion. All AI-generated content was reviewed by the authors for clinical plausibility, consistency with the original data, and suitability for model training and evaluation. Results involving synthetic data were interpreted in the context of known limitations of data augmentation. The authors take full responsibility for the content of the published article. |
| References: | World Health Organization. The importance of pharmacovigilance: safety monitoring of medicinal products. World Health Organization [Internet]. Geneva; 2002 [cited 2025 Sep 12]. Available from: https://www.who.int/publications/i/item/10665-42493. Van De Burgt BWM, Wasylewicz ATM, Dullemond B, Jessurun NT, Grouls RJE, Bouwman RA, et al. Development of a text mining algorithm for identifying adverse drug reactions in electronic health records. JAMIA Open. 2024;7(3). https://doi.org/10.1093/jamiaopen/ooae070. Zitu MM, Zhang S, Owen DH, Chiang C, Li L. Generalizability of machine learning methods in detecting adverse drug events from clinical narratives in electronic medical records. Front Pharmacol. 2023;14. doi: https://doi.org/10.3389/fphar.2023.1218679. Xia L. Historical profile will tell? A deep learning-based multi-level embedding framework for adverse drug event detection and extraction. Decis Support Syst. 2022;160. https://doi.org/10.1016/j.dss.2022.113832. Sankaranarayanapillai M, Wang S, Ji H, Song HY, Tao C. Lessons learned from annotation of VAERS reports on adverse events following influenza vaccination and related to Guillain-Barré syndrome. BMC Med Inform Decis Mak. 2024;23:298. https://doi.org/10.1186/s12911-023-02374-2. (PMID: 10.1186/s12911-023-02374-23818303410770878) Scaboro S, Portelli B, Chersoni E, Santus E, Serra G. Extensive evaluation of transformer-based architectures for adverse drug events extraction. Knowl Based Syst. 2023;275(110675). https://doi.org/10.1016/j.knosys.2023.110675. Ujiie S, Yada S, Wakamiya S, Aramaki E. Identification of adverse drug event–related japanese articles: Natural language processing analysis. JMIR Med Inform. 2020;8(11). https://doi.org/10.2196/22661. Fan B, Fan W, Smith C, Garner HS. Adverse drug event detection and extraction from open data: A deep learning approach. Inf Process Manag. 2020;57(1). https://doi.org/10.1016/j.ipm.2019.102131. Kalkhoran HA, Zwaveling J, Van Hunsel F, Kant A. An innovative method to strengthen evidence for potential drug safety signals using Electronic Health Records. J Med Syst. 2024;48(1):51. doi: https://doi.org/10.1007/s10916-024-02070-2. (PMID: 10.1007/s10916-024-02070-2) Hernández-Arango A, Arias MI, Pérez V, Chavarría LD, Jaimes F. Prediction of the Risk of Adverse Clinical Outcomes with Machine Learning Techniques in Patients with Noncommunicable Diseases. J Med Syst. 2025;49(1):19. doi: https://doi.org/10.1007/s10916-025-02140-z. (PMID: 10.1007/s10916-025-02140-z3990078411790785) Choo SM, Sartori D, Lee SC, Yang HC, Syed-Abdul S. Data-Driven Identification of Factors That Influence the Quality of Adverse Event Reports: 15-Year Interpretable Machine Learning and Time-Series Analyses of VigiBase and QUEST. JMIR Med Inform. 2024;12. doi: https://doi.org/10.2196/49643. National Pharmaceutical Regulatory Agency. National Centre for Adverse Drug Reaction Monitoring Annual Report [Internet]. [cited 2026 Feb 3]. Available from: https://npra.gov.my/index.php/en/informationen/annual-reports/national-centre-for-adverse-drug-reaction-monitoring-annual-report.html. Wallner V. Mapping medical expressions to MedDRA using Natural Language Processing [Master’s thesis] [Internet]. Uppsala University; 2020 [cited 2025 Nov 15]. Available from: http://urn.kb.se/resolve?urn=urn:nbn:se:uu:diva-426916. Nafea AA, Ibrahim MS, Mukhlif AA, Al-Ani MM, Omar N. An Ensemble Model for Detection of Adverse Drug Reactions. ARO-The Scientific Journal of Koya University. 2024;12(1):41–7. doi: https://doi.org/10.14500/aro.11403. (PMID: 10.14500/aro.11403) Kim S, Kang T, Chung TK, Choi Y, Hong YS, Jung K, et al. Automatic Extraction of Comprehensive Drug Safety Information from Adverse Drug Event Narratives in the Korea Adverse Event Reporting System Using Natural Language Processing Techniques. Drug Saf. 2023;46(8):781–95. doi: https://doi.org/10.1007/s40264-023-01323-2. (PMID: 10.1007/s40264-023-01323-23733041510344995) Chopard D, Treder MS, Corcoran P, Ahmed N, Johnson C, Busse M, et al. Text Mining of Adverse Events in Clinical Trials: Deep Learning Approach. JMIR Med Inform. 2021;9(12):e28632. doi: https://doi.org/10.2196/28632. (PMID: 10.2196/28632349516018742206) Mertes PM, Morgand C, Barach P, Jurkolow G, Assmann KE, Dufetelle E, et al. Validation of a natural language processing algorithm using national reporting data to improve identification of anesthesia-related ADVerse evENTs: The “ADVENTURE” study. Anaesth Crit Care Pain Med. 2024;43(4):101390. doi: https://doi.org/10.1016/j.accpm.2024.101390. (PMID: 10.1016/j.accpm.2024.10139038718923) Liu H, Chen Z, Yu M, Cui W, Zang L, Ren W, et al. Evaluation of the Relevance of Adverse Drug Reactions Based on ERNIE-DPCNN. In: 2023 IEEE 47th Annual Computers, Software, and Applications Conference (COMPSAC). IEEE; 2023;1525–9. https://doi.org/10.1109/COMPSAC57700.2023.00235. El-allaly E drissiya, Oubelkacem A, Bourray H. A Gated Multi-hop Attention Fusion Network for Extracting Inter-Sentential Adverse Drug Event Relation. J Healthc Inform Res. 2026;10:246–74. doi: https://doi.org/10.1007/s41666-025-00214-8. (PMID: 10.1007/s41666-025-00214-841658401) Chaichulee S, Promchai C, Kaewkomon T, Kongkamol C, Ingviya T, Sangsupawanich P. Multi-label classification of symptom terms from free-text bilingual adverse drug reaction reports using natural language processing. PLoS One. 2022;17(8):e0270595. doi: https://doi.org/10.1371/journal.pone.0270595. (PMID: 10.1371/journal.pone.0270595359259719352066) Tarekegn AN, Michalak K, Costa G, Ricceri F, Giacobini M. Predicting Multiple Outcomes Associated with Frailty based on Imbalanced Multi-label Classification. J Healthc Inform Res. 2024;8(4):594–618. doi: https://doi.org/10.1007/s41666-024-00173-6. (PMID: 10.1007/s41666-024-00173-63946385711499509) Montani I, Honnibal M, Boyd A, Van Landeghem S, Peters H. explosion/spaCy: v3.7.2: Fixes for APIs and requirements [Internet]. Zenodo; 2023 [cited 2026 Jan 12]. Available from: https://doi.org/10.5281/zenodo.1212303. Neumann M, King D, Beltagy I, Ammar W. ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing. In: Proceedings of the 18th BioNLP Workshop and Shared Task. Association for Computational Linguistics; 2019. p. 319–27. doi: https://doi.org/10.18653/v1/W19-5034. Wei T, Tu WW, Li YF, Yang GP. Towards Robust Prediction on Tail Labels. In: Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery; 2021. p. 1812–20. doi: https://doi.org/10.1145/3447548.3467223. Arslan M, Cruz C. Business text classification with imbalanced data and moderately large label spaces for digital transformation. Appl Netw Sci. 2024;9(1). doi: https://doi.org/10.1007/s41109-024-00623-5. OpenAI. GPT-4 research [Internet]. 2023 [cited 2025 Aug 15]. Available from: https://openai.com/index/gpt-4-research. Sakai H, Lam SS. Large Language Models for Healthcare Text Classification: A Systematic Review. JMIR AI. 2025;5(e79202). https://doi.org/10.2196/79202. Ge Y. Automated Text Classification in Electronic Health Record Narratives Using Machine Learning and Domain-Specific Transformers. In: Proceedings of 2025 6th International Conference on Computer Science and Management Technology, ICCSMT 2025. Association for Computing Machinery; 2026;25–31. https://doi.org/10.1145/3795154.3795158. Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT). Association for Computational Linguistics; 2019;4171–86. https://doi.org/10.18653/v1/N19-1423. Loshchilov I, Hutter F. Decoupled Weight Decay Regularization. In: International Conference on Learning Representations (ICLR) [Internet]. 2019. Available from: http://arxiv.org/abs/1711.05101. Gradio. Gradio documentation [Internet]. [cited 2025 Aug 15]. Available from: https://www.gradio.app. Zhao H, Zhong J, Liang X, Xie C, Wang S. Application of machine learning in drug side effect prediction: databases, methods, and challenges. Front Comput Sci. 2025;19(5):195902. doi: https://doi.org/10.1007/s11704-024-31063-0. (PMID: 10.1007/s11704-024-31063-0) ten Hoope S, Welvaars K, van Geijtenbeek K, Klok-Everaars M, van Schaik S, Karapinar-Çarkit F. Applying text-mining to clinical notes: the identification of patient characteristics from electronic health records (EHRs). BMC Med Inform Decis Mak. 2025;25(1):302. doi: https://doi.org/10.1186/s12911-025-03137-x. (PMID: 10.1186/s12911-025-03137-x4079731712344823) Van Dung H, Tan VM, Dieu NT, Van Linh P, Van Khai N, Ngan TT, et al. Development and external validation of a machine learning model for predicting drug-induced immune thrombocytopenia in a real-world hospital cohort. BMC Med Inform Decis Mak. 2025;25(1):265. doi: https://doi.org/10.1186/s12911-025-03107-3. (PMID: 10.1186/s12911-025-03107-34066530212261740) McMaster C, Chan J, Liew DFL, Su E, Frauman AG, Chapman WW, et al. Developing a deep learning natural language processing algorithm for automated reporting of adverse drug reactions. J Biomed Inform. 2023;137:104265. doi: https://doi.org/10.1016/j.jbi.2022.104265. (PMID: 10.1016/j.jbi.2022.10426536464227) Lauriola I, Lavelli A, Aiolli F. An introduction to Deep Learning in Natural Language Processing: Models, techniques, and tools. Neurocomputing. 2022;470:443–56. doi: https://doi.org/10.1016/j.neucom.2021.05.103. (PMID: 10.1016/j.neucom.2021.05.103) Oyebode O, Orji R. Identifying adverse drug reactions from patient reviews on social media using natural language processing. Health Informatics J. 2023;29(1):14604582221136712. doi: https://doi.org/10.1177/14604582221136712. (PMID: 10.1177/1460458222113671236857033) Song R, Chen X, Liu Z, An H, Zhang Z, Wang X, et al. Label prompt for multi-label text classification. Applied Intelligence. 2023;(53):8761–75. https://doi.org/10.1007/s10489-022-03896-4. Jasila EK, Saleena N, Nazeer KAA. The ScispaCy Clinical Term Extraction and Snomed CT Synonymy Elimination From Clinical Data For Clustering: A Novel Study. Communications in Mathematics and Applications. 2024;15(2):929–45. doi: https://doi.org/10.26713/cma.v15i2.2939. (PMID: 10.26713/cma.v15i2.2939) Guevara M, Chen S, Thomas S, Chaunzwa TL, Franco I, Kann BH, et al. Large language models to identify social determinants of health in electronic health records. NPJ Digit Med. 2024;7(1). https://doi.org/10.1038/s41746-023-00970-0. Yuan Y, Liu Y, Cheng L. A Multi-Faceted Evaluation Framework for Assessing Synthetic Data Generated by Large Language Models. arXiv preprint. 2025 Jul 24. https://doi.org/10.48550/arxiv.2404.14445v2. Chim J, Ive J, Liakata M. Evaluating Synthetic Data Generation from User Generated Text. Computational Linguistics. 2025;51(1):191–233. 10.1162/coli_a_00540 https://doi.org/10.1162/coli_a_00540. (PMID: 10.1162/coli_a_00540) Barucci A, Colcelli V, De Masi S, Falconi M, Leo MC, Marzola A, et al. Ethical and Regulatory Frameworks for Artificial Intelligence in Clinical Research: A European Perspective on the Artificial Intelligence Act for Ethics Committees and Researchers. European Cardiology Review. 2026;21:e01. doi: https://doi.org/10.15420/ecr.2025.59. (PMID: 10.15420/ecr.2025.594167636312888102) United Nations. Transforming our world: The 2030 agenda for sustainable development [Internet]. 2015 [cited 2025 Nov 15]. Available from: https://sdgs.un.org/goals. Economic Planning Unit PMD. MyDIGITAL: Malaysia Digital Economy Blueprint [Internet]. 2021 [cited 2025 Nov 15]. Available from: https://www.epu.gov.my/en/digital-economy-blueprint. |
| Contributed Indexing: | Keywords: Adverse Drug Reactions; MedDRA; Multi-label Classification; Natural Language Processing; Pharmacovigilance; Real-World Data |
| Entry Date(s): | Date Created: 20260731 Date Completed: 20260731 Latest Revision: 20260731 |
| Update Code: | 20260731 |
| DOI: | 10.1007/s10916-026-02448-4 |
| PMID: | 42536121 |
| Βάση Δεδομένων: | MEDLINE |
| ISSN: | 1573-689X |
|---|---|
| DOI: | 10.1007/s10916-026-02448-4 |