Academic Journal
Describing Land Cover Changes via Multi-Temporal Remote Sensing Image Captioning Using LLM, ViT, and LoRA
| Τίτλος: | Describing Land Cover Changes via Multi-Temporal Remote Sensing Image Captioning Using LLM, ViT, and LoRA |
|---|---|
| Συγγραφείς: | Lamar-Leon, Javier, Nogueira, Vitor, Salgueiro, Pedro, Quaresma, Paulo |
| Συνεισφορές: | Pan, Jiayi, Li, Xinghua |
| Στοιχεία εκδότη: | Remote Sensing MDPI |
| Έτος έκδοσης: | 2026 |
| Συλλογή: | Repositório Científico da Universidade de Évora |
| Θεματικοί όροι: | Image Captioning, Remote Sensing, LLM, LoRA |
| Περιγραφή: | Describing land cover changes from multi-temporal remote sensing imagery requires capturing both visual transformations and their semantic meaning in natural language. Existing methods often struggle to balance visual accuracy with descriptive coherence. We propose MVLT-LoRA-CC (Multi-modal Vision Language Transformer with Low-Rank Adaptation for Change Captioning), a framework that integrates a Vision Transformer (ViT), a Large Language Model (LLM), and Low-Rank Adaptation (LoRA) for efficient multi-modal learning. The model processes paired temporal images through patch embeddings and transformer blocks, aligning visual and textual representations via a multi-modal adapter. To improve efficiency and avoid unnecessary parameter growth, LoRA modules are selectively inserted only into the attention projection layers and cross-modal adapter blocks rather than being uniformly applied to all linear layers. This targeted design preserves general linguistic knowledge while enabling effective adaptation to remote sensing change description. To assess performance, we introduce the Complementary Consistency Score (CCS) framework, which evaluates both descriptive fidelity for change instances and classification accuracy for no change cases. Experiments on the LEVIR-CC test set demonstrate that MVLT-LoRA-CC generates semantically accurate captions, surpassing prior methods in both descriptive richness and temporal change recognition. The approach establishes a scalable solution for multi-modal land cover change description in remote sensing applications. |
| Τύπος εγγράφου: | article in journal/newspaper |
| Γλώσσα: | Portuguese |
| Relation: | https://hdl.handle.net/10174/41416; jlamarleon@uevora.pt; vbn@uevora.pt; pds@uevora.pt; pq@uevora.pt; 283; https://doi.org/10.3390/rs18010166 |
| DOI: | 10.3390/rs18010166 |
| Διαθεσιμότητα: | https://hdl.handle.net/10174/41416 https://doi.org/10.3390/rs18010166 |
| Rights: | openAccess |
| Αριθμός Καταχώρησης: | edsbas.EFEEF817 |
| Βάση Δεδομένων: | BASE |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://hdl.handle.net/10174/41416# Name: EDS - BASE (ns324271) Category: fullText Text: View record from BASE |
|---|---|
| Header | DbId: edsbas DbLabel: BASE An: edsbas.EFEEF817 RelevancyScore: 1025 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 1024.80883789063 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Describing Land Cover Changes via Multi-Temporal Remote Sensing Image Captioning Using LLM, ViT, and LoRA – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Lamar-Leon%2C+Javier%22">Lamar-Leon, Javier</searchLink><br /><searchLink fieldCode="AR" term="%22Nogueira%2C+Vitor%22">Nogueira, Vitor</searchLink><br /><searchLink fieldCode="AR" term="%22Salgueiro%2C+Pedro%22">Salgueiro, Pedro</searchLink><br /><searchLink fieldCode="AR" term="%22Quaresma%2C+Paulo%22">Quaresma, Paulo</searchLink> – Name: Author Label: Contributors Group: Au Data: Pan, Jiayi<br />Li, Xinghua – Name: Publisher Label: Publisher Information Group: PubInfo Data: Remote Sensing MDPI – Name: DatePubCY Label: Publication Year Group: Date Data: 2026 – Name: Subset Label: Collection Group: HoldingsInfo Data: Repositório Científico da Universidade de Évora – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Image+Captioning%22">Image Captioning</searchLink><br /><searchLink fieldCode="DE" term="%22Remote+Sensing%22">Remote Sensing</searchLink><br /><searchLink fieldCode="DE" term="%22LLM%22">LLM</searchLink><br /><searchLink fieldCode="DE" term="%22LoRA%22">LoRA</searchLink> – Name: Abstract Label: Description Group: Ab Data: Describing land cover changes from multi-temporal remote sensing imagery requires capturing both visual transformations and their semantic meaning in natural language. Existing methods often struggle to balance visual accuracy with descriptive coherence. We propose MVLT-LoRA-CC (Multi-modal Vision Language Transformer with Low-Rank Adaptation for Change Captioning), a framework that integrates a Vision Transformer (ViT), a Large Language Model (LLM), and Low-Rank Adaptation (LoRA) for efficient multi-modal learning. The model processes paired temporal images through patch embeddings and transformer blocks, aligning visual and textual representations via a multi-modal adapter. To improve efficiency and avoid unnecessary parameter growth, LoRA modules are selectively inserted only into the attention projection layers and cross-modal adapter blocks rather than being uniformly applied to all linear layers. This targeted design preserves general linguistic knowledge while enabling effective adaptation to remote sensing change description. To assess performance, we introduce the Complementary Consistency Score (CCS) framework, which evaluates both descriptive fidelity for change instances and classification accuracy for no change cases. Experiments on the LEVIR-CC test set demonstrate that MVLT-LoRA-CC generates semantically accurate captions, surpassing prior methods in both descriptive richness and temporal change recognition. The approach establishes a scalable solution for multi-modal land cover change description in remote sensing applications. – Name: TypeDocument Label: Document Type Group: TypDoc Data: article in journal/newspaper – Name: Language Label: Language Group: Lang Data: Portuguese – Name: NoteTitleSource Label: Relation Group: SrcInfo Data: https://hdl.handle.net/10174/41416; jlamarleon@uevora.pt; vbn@uevora.pt; pds@uevora.pt; pq@uevora.pt; 283; https://doi.org/10.3390/rs18010166 – Name: DOI Label: DOI Group: ID Data: 10.3390/rs18010166 – Name: URL Label: Availability Group: URL Data: https://hdl.handle.net/10174/41416<br />https://doi.org/10.3390/rs18010166 – Name: Copyright Label: Rights Group: Cpyrght Data: openAccess – Name: AN Label: Accession Number Group: ID Data: edsbas.EFEEF817 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.EFEEF817 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.3390/rs18010166 Languages: – Text: Portuguese Subjects: – SubjectFull: Image Captioning Type: general – SubjectFull: Remote Sensing Type: general – SubjectFull: LLM Type: general – SubjectFull: LoRA Type: general Titles: – TitleFull: Describing Land Cover Changes via Multi-Temporal Remote Sensing Image Captioning Using LLM, ViT, and LoRA Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Lamar-Leon, Javier – PersonEntity: Name: NameFull: Nogueira, Vitor – PersonEntity: Name: NameFull: Salgueiro, Pedro – PersonEntity: Name: NameFull: Quaresma, Paulo – PersonEntity: Name: NameFull: Pan, Jiayi – PersonEntity: Name: NameFull: Li, Xinghua IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2026 Identifiers: – Type: issn-locals Value: edsbas – Type: issn-locals Value: edsbas.oa |
| ResultId | 1 |