Academic Journal
GeoPyEval: An Automated Evaluation Framework for Python Code Generation Capabilities of Large Language Models in Geospatial Domains.
| Title: | GeoPyEval: An Automated Evaluation Framework for Python Code Generation Capabilities of Large Language Models in Geospatial Domains. |
|---|---|
| Authors: | Wu, Shaowen1 (AUTHOR), Xie, Lutong1 (AUTHOR), Hou, Shuyang1 (AUTHOR) whuhsy@whu.edu.cn, Chen, Guanyu2 (AUTHOR), Jiao, Haoyue2 (AUTHOR), Liu, Ziqi1 (AUTHOR), Guan, Xuefeng1 (AUTHOR), Wu, Huayi1 (AUTHOR) |
| Source: | Transactions in GIS. Jun2026, Vol. 30 Issue 4, p1-28. 28p. |
| Subject Terms: | *Benchmarking (Management), Python programming language, Geoinformatics, Language models |
| Abstract: | Despite the increasing adoption of large language models (LLMs) for Python geospatial code generation, no automated evaluation framework exists for this domain. We propose GeoPyEval, the first function‐level framework in this field, featuring an expert‐in‐the‐loop, automated execution design with three core modules: (a) GeoPyEval‐Bench, constructed from expert‐verified LLM‐generated content, covering 18 mainstream geospatial libraries with 1074 unit test tasks; (b) an expert‐configurable submission program for invoking LLMs to perform code generation and execution; and (c) a judging program for correctness evaluation. We systematically assessed 12 mainstream LLMs available as of November 2025 regarding accuracy, resource consumption, operational efficiency, and error type logs. Results show that GPT‐5 achieved the highest accuracy with a Pass@5 of 64.28%. GeoPyEval extends existing LLM evaluation frameworks to Python geospatial code generation, establishing a novel benchmarking standard for this domain. [ABSTRACT FROM AUTHOR] |
| Copyright of Transactions in GIS is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Business Source Index |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://resolver.ebsco.com/c/fiv2js/result?sid=EBSCO:bsx&genre=article&issn=13611682&ISBN=&volume=30&issue=4&date=20260601&spage=1&pages=1-28&title=Transactions in GIS&atitle=GeoPyEval%3A%20An%20Automated%20Evaluation%20Framework%20for%20Python%20Code%20Generation%20Capabilities%20of%20Large%20Language%20Models%20in%20Geospatial%20Domains.&aulast=Wu%2C%20Shaowen&id=DOI:10.1111/tgis.70324 Name: Full Text Finder (for New FTF UI) (ns324271) Category: fullText Text: Full Text Finder MouseOverText: Full Text Finder |
|---|---|
| Header | DbId: bsx DbLabel: Business Source Index An: 194919340 RelevancyScore: 1435 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 1435 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: GeoPyEval: An Automated Evaluation Framework for Python Code Generation Capabilities of Large Language Models in Geospatial Domains. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Wu%2C+Shaowen%22">Wu, Shaowen</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Xie%2C+Lutong%22">Xie, Lutong</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Hou%2C+Shuyang%22">Hou, Shuyang</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> whuhsy@whu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Chen%2C+Guanyu%22">Chen, Guanyu</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Jiao%2C+Haoyue%22">Jiao, Haoyue</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Liu%2C+Ziqi%22">Liu, Ziqi</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Guan%2C+Xuefeng%22">Guan, Xuefeng</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Wu%2C+Huayi%22">Wu, Huayi</searchLink><relatesTo>1</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Transactions+in+GIS%22">Transactions in GIS</searchLink>. Jun2026, Vol. 30 Issue 4, p1-28. 28p. – Name: Subject Label: Subject Terms Group: Su Data: *<searchLink fieldCode="DE" term="%22Benchmarking+%28Management%29%22">Benchmarking (Management)</searchLink><br /><searchLink fieldCode="DE" term="%22Python+programming+language%22">Python programming language</searchLink><br /><searchLink fieldCode="DE" term="%22Geoinformatics%22">Geoinformatics</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Despite the increasing adoption of large language models (LLMs) for Python geospatial code generation, no automated evaluation framework exists for this domain. We propose GeoPyEval, the first function‐level framework in this field, featuring an expert‐in‐the‐loop, automated execution design with three core modules: (a) GeoPyEval‐Bench, constructed from expert‐verified LLM‐generated content, covering 18 mainstream geospatial libraries with 1074 unit test tasks; (b) an expert‐configurable submission program for invoking LLMs to perform code generation and execution; and (c) a judging program for correctness evaluation. We systematically assessed 12 mainstream LLMs available as of November 2025 regarding accuracy, resource consumption, operational efficiency, and error type logs. Results show that GPT‐5 achieved the highest accuracy with a Pass@5 of 64.28%. GeoPyEval extends existing LLM evaluation frameworks to Python geospatial code generation, establishing a novel benchmarking standard for this domain. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Transactions in GIS is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=bsx&AN=194919340 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/tgis.70324 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 28 StartPage: 1 Subjects: – SubjectFull: Benchmarking (Management) Type: general – SubjectFull: Python programming language Type: general – SubjectFull: Geoinformatics Type: general – SubjectFull: Language models Type: general Titles: – TitleFull: GeoPyEval: An Automated Evaluation Framework for Python Code Generation Capabilities of Large Language Models in Geospatial Domains. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Wu, Shaowen – PersonEntity: Name: NameFull: Xie, Lutong – PersonEntity: Name: NameFull: Hou, Shuyang – PersonEntity: Name: NameFull: Chen, Guanyu – PersonEntity: Name: NameFull: Jiao, Haoyue – PersonEntity: Name: NameFull: Liu, Ziqi – PersonEntity: Name: NameFull: Guan, Xuefeng – PersonEntity: Name: NameFull: Wu, Huayi IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 06 Text: Jun2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 13611682 Numbering: – Type: volume Value: 30 – Type: issue Value: 4 Titles: – TitleFull: Transactions in GIS Type: main |
| ResultId | 1 |