Academic Journal

Improving intelligent perception and decision optimization of pedestrian crossing scenarios in autonomous driving environments through large visual language models.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Improving intelligent perception and decision optimization of pedestrian crossing scenarios in autonomous driving environments through large visual language models.
Συγγραφείς: Teng, Xiao, Huang, Lin, Shen, Zhenjiang, Li, Wankai
Πηγή: Scientific Reports; 8/25/2025, Vol. 15 Issue 1, p1-18, 18p
Θεματικοί όροι: Autonomous vehicles, Intelligent transportation systems, Decision support systems, Pedestrian crosswalks, Data structures, Traffic safety, Information processing, Dynamic models
Περίληψη: This study leverages large Visual Language Models (VLM) to develop an intelligent pedestrian crossing scenario system within autonomous driving environments. By establishing standardized checklists and prompts, the system minimizes the risks of misjudgment and omission through multimodal data processing. It offers data-driven decision-making support, presenting an innovative approach to integrating autonomous driving technology with intelligent transportation systems. The study begins by classifying pedestrian crossing scenarios based on international autonomous driving standards, distinguishing between pedestrian crossings and autonomous vehicle crossings, as well as dynamic and static entities. Next, standardized prompts derived from these standards are fed into the VLM, generating structured scenario checklists of dynamic and static entities, outputted in JSON format. This systematic identification and processing of entities—such as pedestrians, vehicles, and traffic facilities—enables the construction of structured data representations for complex traffic scenarios. Building on this foundation, the VLM analyzes scenario data to predict collision risks by modeling the behaviors of both pedestrians and vehicles, supporting real-time decision-making for autonomous vehicles and road users. Furthermore, the VLM processes scene data to anticipate potential conflicts and provide actionable safety recommendations, enhancing the overall security of all traffic participants. The system achieved a perception accuracy of 93.05%, with risk prediction consistency and decision-making rule consistency rates of 85.91% and 87.72% respectively. By constructing a VLM-based intelligent pedestrian crossing perception system, this study offers a novel technical framework for improving perception, prediction, and decision-making in autonomous driving. Unlike traditional rule-based and deep learning approaches, which struggle with complex pedestrian behaviors and dynamic environments, our method integrates visual perception with reasoning capabilities, enabling structured, standardized, and explainable decision-making in pedestrian crossing scenarios. [ABSTRACT FROM AUTHOR]
Copyright of Scientific Reports is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index
FullText Links:
  – Type: other
Text:
  Availability: 0
CustomLinks:
  – Url: https://dx.doi.org/doi:10.1038/s41598-025-14827-x
    Name: EDS - Springer Nature Journals (s7799221)
    Category: fullText
    Text: View record at Springer
Header DbId: edb
DbLabel: Complementary Index
An: 187534812
RelevancyScore: 1007
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 1007.33734130859
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Improving intelligent perception and decision optimization of pedestrian crossing scenarios in autonomous driving environments through large visual language models.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Teng%2C+Xiao%22">Teng, Xiao</searchLink><br /><searchLink fieldCode="AR" term="%22Huang%2C+Lin%22">Huang, Lin</searchLink><br /><searchLink fieldCode="AR" term="%22Shen%2C+Zhenjiang%22">Shen, Zhenjiang</searchLink><br /><searchLink fieldCode="AR" term="%22Li%2C+Wankai%22">Li, Wankai</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: Scientific Reports; 8/25/2025, Vol. 15 Issue 1, p1-18, 18p
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Autonomous+vehicles%22">Autonomous vehicles</searchLink><br /><searchLink fieldCode="DE" term="%22Intelligent+transportation+systems%22">Intelligent transportation systems</searchLink><br /><searchLink fieldCode="DE" term="%22Decision+support+systems%22">Decision support systems</searchLink><br /><searchLink fieldCode="DE" term="%22Pedestrian+crosswalks%22">Pedestrian crosswalks</searchLink><br /><searchLink fieldCode="DE" term="%22Data+structures%22">Data structures</searchLink><br /><searchLink fieldCode="DE" term="%22Traffic+safety%22">Traffic safety</searchLink><br /><searchLink fieldCode="DE" term="%22Information+processing%22">Information processing</searchLink><br /><searchLink fieldCode="DE" term="%22Dynamic+models%22">Dynamic models</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This study leverages large Visual Language Models (VLM) to develop an intelligent pedestrian crossing scenario system within autonomous driving environments. By establishing standardized checklists and prompts, the system minimizes the risks of misjudgment and omission through multimodal data processing. It offers data-driven decision-making support, presenting an innovative approach to integrating autonomous driving technology with intelligent transportation systems. The study begins by classifying pedestrian crossing scenarios based on international autonomous driving standards, distinguishing between pedestrian crossings and autonomous vehicle crossings, as well as dynamic and static entities. Next, standardized prompts derived from these standards are fed into the VLM, generating structured scenario checklists of dynamic and static entities, outputted in JSON format. This systematic identification and processing of entities—such as pedestrians, vehicles, and traffic facilities—enables the construction of structured data representations for complex traffic scenarios. Building on this foundation, the VLM analyzes scenario data to predict collision risks by modeling the behaviors of both pedestrians and vehicles, supporting real-time decision-making for autonomous vehicles and road users. Furthermore, the VLM processes scene data to anticipate potential conflicts and provide actionable safety recommendations, enhancing the overall security of all traffic participants. The system achieved a perception accuracy of 93.05%, with risk prediction consistency and decision-making rule consistency rates of 85.91% and 87.72% respectively. By constructing a VLM-based intelligent pedestrian crossing perception system, this study offers a novel technical framework for improving perception, prediction, and decision-making in autonomous driving. Unlike traditional rule-based and deep learning approaches, which struggle with complex pedestrian behaviors and dynamic environments, our method integrates visual perception with reasoning capabilities, enabling structured, standardized, and explainable decision-making in pedestrian crossing scenarios. [ABSTRACT FROM AUTHOR]
– Name: Abstract
  Label:
  Group: Ab
  Data: <i>Copyright of Scientific Reports is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=187534812
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1038/s41598-025-14827-x
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 18
        StartPage: 1
    Subjects:
      – SubjectFull: Autonomous vehicles
        Type: general
      – SubjectFull: Intelligent transportation systems
        Type: general
      – SubjectFull: Decision support systems
        Type: general
      – SubjectFull: Pedestrian crosswalks
        Type: general
      – SubjectFull: Data structures
        Type: general
      – SubjectFull: Traffic safety
        Type: general
      – SubjectFull: Information processing
        Type: general
      – SubjectFull: Dynamic models
        Type: general
    Titles:
      – TitleFull: Improving intelligent perception and decision optimization of pedestrian crossing scenarios in autonomous driving environments through large visual language models.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Teng, Xiao
      – PersonEntity:
          Name:
            NameFull: Huang, Lin
      – PersonEntity:
          Name:
            NameFull: Shen, Zhenjiang
      – PersonEntity:
          Name:
            NameFull: Li, Wankai
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 25
              M: 08
              Text: 8/25/2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 20452322
          Numbering:
            – Type: volume
              Value: 15
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Scientific Reports
              Type: main
ResultId 1