Academic Journal
Enhancing Tactile Paving Segmentation via Multidimensional Perception Networks.
| Τίτλος: | Enhancing Tactile Paving Segmentation via Multidimensional Perception Networks. |
|---|---|
| Συγγραφείς: | Zhang A; School of Computer and Software Engineering, Huaiyin Institute of Technology, Huai'an, Jiangsu, China., Zhang H; School of Computer and Software Engineering, Huaiyin Institute of Technology, Huai'an, Jiangsu, China., Zhu Y; School of Computer and Software Engineering, Huaiyin Institute of Technology, Huai'an, Jiangsu, China., Guo Y; School of Computer and Software Engineering, Huaiyin Institute of Technology, Huai'an, Jiangsu, China., Zhuang Z; School of Computer and Software Engineering, Huaiyin Institute of Technology, Huai'an, Jiangsu, China., Ma T; School of Computer and Software Engineering, Huaiyin Institute of Technology, Huai'an, Jiangsu, China., Cai Y; School of Computer and Software Engineering, Huaiyin Institute of Technology, Huai'an, Jiangsu, China., Liu L; School of Computer and Software Engineering, Huaiyin Institute of Technology, Huai'an, Jiangsu, China. |
| Πηγή: | Annals of the New York Academy of Sciences [Ann N Y Acad Sci] 2026 Jun; Vol. 1560 (1), pp. e70298. |
| Τύπος έκδοσης: | Journal Article |
| Γλώσσα: | English |
| Στοιχεία περιοδικού: | Publisher: New York Academy of Sciences Country of Publication: United States NLM ID: 7506858 Publication Model: Print Cited Medium: Internet ISSN: 1749-6632 (Electronic) Linking ISSN: 00778923 NLM ISO Abbreviation: Ann N Y Acad Sci Subsets: MEDLINE |
| Imprint Name(s): | Publication: 2006- : New York, NY : Malden, MA : New York Academy of Sciences ; Blackwell Original Publication: New York, The Academy. |
| Ιατρικοί όροι (MeSH): | Touch Perception*/physiology , Touch*/physiology , Classification Algorithms* , Construction Materials*, Attention/physiology ; Image Processing, Computer-Assisted/methods ; Humans ; Artificial Intelligence ; Convolutional Neural Networks ; Persons with Visual Disabilities |
| Περίληψη: | Tactile paving is essential for ensuring the safe and independent travel of visually impaired individuals in urban environments. However, existing segmentation models often fail to generalize well across diverse scenarios and rely heavily on color information, neglecting the physical structure of tactile patterns. This article introduces EGM-Unet, a novel segmentation method that incorporates an edge-aware multi-scale attentional fusion block (EMAF-Block) to enhance edge feature extraction, a recursive gated attention (RGA) mechanism to focus on critical regions, and a multidimensional cooperative aggregation attention (MCAA) module to refine texture features. Additionally, we utilize CLIPSeg with a correlative self-attention (CSA) mechanism to improve generalization. Our experiments demonstrate that EGM-Unet outperforms standard U-Net and other mainstream models, achieving a mean IoU of 93.73%, an IoU of 89.09%, and an accuracy of 98.56%. These results highlight the robustness of our method in segmenting tactile paving regions across various materials and lighting conditions, providing a solid foundation for high-precision tactile paving perception. (© 2026 The New York Academy of Sciences.) |
| References: | R. K. Katzschmann, B. Araki, and D. Rus, “Safe Local Navigation for Visually Impaired Users With a Time‐of‐Flight and Haptic Feedback Device,” IEEE Transactions on Neural Systems and Rehabilitation Engineering 26, no. 3 (2018): 583–593, https://doi.org/10.1109/TNSRE.2018.2800665. S. Zhang, X. Y. Liu, and J. Zeng, “Urban Safety Evacuation Strategy for the Visually Impaired,” Planners 34, no. 3 (2018): 103–107. Y. Q. Wang, X. H. Huang, and X. F. Shen, “Design of Crossing Guide and Road Obstacle Detection Warning System,” Measurement & Control Technology 37, no. 10 (2018): 114–118. S. Dong, W. Zhou, C. Xu, and W. Yan, “EGFNet: Edge‐Aware Guidance Fusion Network for RGB‐Thermal Urban Scene Parsing,” IEEE Transactions on Intelligent Transportation Systems 25, no. 1 (2024): 657–669, https://doi.org/10.1109/TITS.2023.3306368. W. Zhou, X. Lin, J. Lei, L. Yu, and J.‐N. Hwang, “MFFENet: Multiscale Feature Fusion and Enhancement Network for RGB–Thermal Urban Road Scene Parsing,” IEEE Transactions on Multimedia 24 (2022): 2526–2538, https://doi.org/10.1109/TMM.2021.3086618. W. Zhou, E. Yang, J. Lei, J. Wan, and L. Yu, “PGDENet: Progressive Guided Fusion and Depth Enhancement Network for RGB‐D Indoor Scene Parsing,” IEEE Transactions on Multimedia 25 (2023): 3483–3494, https://doi.org/10.1109/TMM.2022.3161852. W. Zhou, H. Zhang, W. Yan, and W. Lin, “MMSMCNet: Modal Memory Sharing and Morphological Complementary Networks for RGB‐T Urban Scene Semantic Segmentation,” IEEE Transactions on Circuits and Systems for Video Technology 33, no. 12 (2023): 7096–7108, https://doi.org/10.1109/TCSVT.2023.3275314. W. Zhou, H. Wu, and Q. Jiang, “MDNet: Mamba‐Effective Diffusion‐Distillation Network for RGB‐Thermal Urban Dense Prediction,” IEEE Transactions on Circuits and Systems for Video Technology 35, no. 4 (2025): 3222–3233, https://doi.org/10.1109/TCSVT.2024.3508058. W. Zhou, X. Fan, and W. Yan, “Graph Attention Guidance Network With Knowledge Distillation for Semantic Segmentation of Remote Sensing Images,” IEEE Transactions on Geoscience and Remote Sensing 61 (2023): 4506015, https://doi.org/10.1109/TGRS.2023.3311480. Y. Zhang and C. Liu, “Network for Robust and High‐Accuracy Pavement Crack Segmentation,” Automation in Construction 162 (2024): 105375, https://doi.org/10.1016/j.autcon.2024.105375. Y. Zhang and C. Liu, “Crack Segmentation Using Discrete Cosine Transform in Shadow Environments,” Automation in Construction 166 (2024): 105646, https://doi.org/10.1016/j.autcon.2024.105646. Y. Zhang and C. Liu, “Generative Adversarial Network Based on Domain Adaptation for Crack Segmentation in Shadow Environments,” Computer‐Aided Civil and Infrastructure Engineering 40, no. 24 (2025): 3997–4013, https://doi.org/10.1111/mice.13451. X. Wang, X. Wang, Y. Luo, et al., “Scene‐Aware Vectorized Memory Multi‐Agent Framework With Cross‐Modal Differentiated Quantization VLMs for Visually Impaired Assistance,” Expert Systems with Applications 303 (2026): 130662, https://doi.org/10.1016/j.eswa.2025.130662. X. Y. Wang, “RPIQ: Residual‐Projected Multi‐Collaboration Closed‐Loop and Single Instance Quantization for Visually Impaired Assistance,” arXiv (2026), https://arxiv.org/abs/2601.02888. A. Kirillov, E. Mintun, and N. Ravi, “Segment Anything,” in Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (IEEE, 2023), https://doi.org/10.48550/arXiv.2304.02643. X. Zou, J. Yang, and H. Zhang, “Segment Everything Everywhere All at Once,” arXiv (2023), https://doi.org/10.48550/arXiv.2304.06718. Y. Q. Peng, J. Xue, and Y. F. Guo, “Blind Road Recognition Algorithm Based on Color and Texture Information,” Journal of Computer Applications 34, no. 12 (2014): 3585–3588. L. Zhao, Z. W. Li, and X. L. Yang, “A Warning Blind Sidewalk Detection Method Based on Image Processing,” Computer Technology and Development 31, no. 2 (2021): 91–96. J. G. Ke, Q. F. Zhao, and P. F. Shi, “Blind Way Recognition Algorithm Based on Image Processing,” Computer Engineering 35, no. 1 (2009): 189–197, https://doi.org/10.3969/j.issn.1000‐3428.2009.01.065. T. Wei and Y. H. Zhou, “Blind Sidewalk Image Location Based on Machine Learning Recognition and Marked Watershed Segmentation,” Optics and Precision Engineering 27, no. 1 (2019): 201–210, https://doi.org/10.3788/OPE.20192701.0201. Z.‐Q. Zhao, P. Zheng, S.‐T. Xu, and X. Wu, “Object Detection With Deep Learning: A Review,” IEEE Transactions on Neural Networks and Learning Systems 30, no. 11 (2019): 3212–3232, https://doi.org/10.1109/TNNLS.2018.2876865. V. Athanasios, D. Nikolaos, and D. Anastasios, “Deep Learning for Computer Vision: A Brief Review,” Computational Intelligence and Neuroscience 2018, no. 1 (2018): 7068349, https://doi.org/10.1155/2018/7068349. Z. Chen, H. Wang, and M. Zhu, “Detection of Tactile Pavement and Crosswalk Based on YOLOv5 Algorithm,” Information Technology & Informatization 2022, no. 7 (2022): 10–14. G. Jocher. YOLOv5. (2022), https://github.com/ultralytics/YOLOv5. Y. Xia, Y. Li, Q. Ye, and J. Dong, “Image Segmentation for Blind Lanes Based on Improved SegNet Model,” Journal of Electronic Imaging 32, no. 1 (2023): 013038, https://doi.org/10.1117/1.JEI.32.1.013038. V. Badrinarayanan, A. Kendall, and R. Cipolla, “SegNet: A Deep Convolutional Encoder‐Decoder Architecture for Image Segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence 39, no. 12 (2017): 2481–2495, https://doi.org/10.1109/TPAMI.2016.2644615. J. Z. Chen and X. Z. Bai, “Atmospheric Transmission and Thermal Inertia Induced Blind Road Segmentation With a Large‐Scale Dataset TBRSD,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (IEEE, 2023), https://doi.org/10.1109/ICCV51070.2023.00103. O. Ronneberger, P. Fischer, and T. Brox, “U‐Net: Convolutional Networks for Biomedical Image Segmentation,” in International Conference on Medical Image Computing and Computer‐Assisted Intervention (Springer, 2015), https://doi.org/10.1007/978‐3‐319‐24574‐4_28. T. Lüddecke and A. Ecker, “Image Segmentation Using Text and Image Prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (IEEE, 2022). A. Radford, J. W. Kim, and C. Hallacy, “Learning Transferable Visual Models From Natural Language Supervision,” in International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, 2021), https://proceedings.mlr.press/v139/radford21a.html. Y. Zheng, J. H. Wu, and Y. Q. Qin, “Zero‐Shot Instance Segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2021), https://doi.org/10.1109/CVPR46437.2021.00262. Z. Wu and Z. Y. Xiao, “Few‐Shot Learning Based on Deep Learning: A Survey,” Mathematical Biosciences and Engineering 21, no. 1 (2024): 679–711, https://doi.org/10.3934/mbe.2024029. F. Wang, X. Mei, and Z. Zhang, “SCLIP: Rethinking Self‐Attention for Dense Vision‐Language Inference,” in Proceedings of the European Conference on Computer Vision (ECCV) (Springer, 2024), https://doi.org/10.1007/978‐3‐031‐72664‐4_18. B. Zhang, P. Zhang, and X. Dong, “Long‐CLIP: Unlocking the Long‐Text Capability of CLIP,” in Proceedings of the European Conference on Computer Vision (ECCV) (Springer, 2024), https://doi.org/10.1007/978‐3‐031‐72983‐6_18. X. Zhang, L. Liang, S. Zhao, and Z. Wang, “GRFB‐UNet: A New Multi‐Scale Attention Network With Group Receptive Field Block for Tactile Paving Segmentation,” Expert Systems With Applications 238 (2024): 122109, https://doi.org/10.1016/j.eswa.2023.122109. L.‐C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking Atrous Convolution for Semantic Image Segmentation,” arXiv (2017), https://arxiv.org/abs/1706.05587. J. Ma, F. Li, and B. Wang, “U‐Mamba: Enhancing Long‐Range Dependency for Biomedical Image Segmentation, ” arXiv (2024), https://arxiv.org/abs/2401.04722. J. Wang, K. Sun, T. Cheng, et al., “Deep High‐Resolution Representation Learning for Visual Recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence 43, no. 10 (2020): 3349–3364, https://doi.org/10.1109/TPAMI.2020.2983686. A. G. Howard, M. Zhu, and B. Chen, “Searching for MobileNetV3,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (IEEE, 2019), https://doi.org/10.1109/ICCV.2019.00140. J. Long, E. Shelhamer, and T. Darrell, “Fully Convolutional Networks for Semantic Segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (IEEE, 2015), https://doi.org/10.1109/CVPR.2015.7298965. X. Qin, Z. Zhang, C. Huang, M. Dehghan, O. R. Zaiane, and M. Jagersand, “U2Net: Going Deeper With Nested U‐Structure for Salient Object Detection,” Pattern Recognition 106 (2020): 107404, https://doi.org/10.1016/j.patcog.2020.107404. J. C. Ruan, J. C. Li, and S. C. Xiang, “VM‐UNet: Vision Mamba UNet for Medical Image Segmentation,” arXiv (2024), https://arxiv.org/abs/2402.02491. |
| Grant Information: | Z421A25056 Enterprise Collaboration Project |
| Contributed Indexing: | Keywords: U‐Net; attention mechanism; multi‐scale feature fusion; tactile paving segmentation |
| Entry Date(s): | Date Created: 20260602 Date Completed: 20260609 Latest Revision: 20260726 |
| Update Code: | 20260726 |
| PubMed Central ID: | PMC13390594 |
| DOI: | 10.1111/nyas.70298 |
| PMID: | 42229374 |
| Βάση Δεδομένων: | MEDLINE |
| ISSN: | 1749-6632 |
|---|---|
| DOI: | 10.1111/nyas.70298 |