Academic Journal

Bi-directional Long Short-term Memory with Hybrid Optimiser for Image Captioning.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Bi-directional Long Short-term Memory with Hybrid Optimiser for Image Captioning.
Συγγραφείς: Hanumanth Raju, R. K., Vinay Kumar, S. B.
Πηγή: Journal of Information & Knowledge Management; Aug2026, Vol. 25 Issue 8, p1-37, 37p
Θεματικοί όροι: Long short-term memory, Photograph captions, Subroutines (Computer programs), Deep learning, Optimization algorithms, Mathematical optimization, Natural language processing, Metaheuristic algorithms
Περίληψη: Image captioning is a rapidly evolving Artificial Intelligence (AI) domain that integrates computer vision and natural language processing to generate descriptive text for visual content. However, existing image captioning models often struggle with limited contextual understanding and vanishing gradient issues. To address these challenges, this research develops a Deep Learning-based Image Captioning Framework (DLICF) designed to enhance the semantic and contextual coherence of generated captions. The main objective is to develop a model that effectively captures both visual and linguistic dependencies while overcoming the computational limitations of traditional architectures. The proposed framework introduces a Modified Activation Function-based Bidirectional Long Short-Term Memory (MAF-Bi-LSTM) network that integrates the hybrid APTx + MReLU activation function that effectively mitigates vanishing gradient and ensures stable and efficient learning. To further enhance optimisation and convergence, the model is fine-tuned using a hybrid Flower Pollination Assisted Black Widow Optimisation (FPABWO) algorithm, which combines biologically inspired mechanisms of Flower Pollination Optimisation and Black Widow Optimisation. Subsequently, the output from the MAF-Bi-LSTM is passed to a decoder that generates a coherent and contextually accurate textual description of the image. Experimental evaluation on benchmark datasets demonstrates that the MAF-Bi-LSTM + FPABWO model significantly outperforms conventional methods in terms of BLEU, CIDEr and ROUGE scores. The suggested MAF-Bi-LSTM + FPABWO method achieved a better BLEU-4 of 0.639, CIDER value of 0.751 and ROUGE score of 0.843 at 80% of training data. The findings confirm that the proposed framework achieves higher caption relevance, improved fluency and enhanced generalisation capability. This innovation provides a new theoretical and computational foundation for advancing image understanding systems. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Information & Knowledge Management is the property of World Scientific Publishing Company and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index