Academic Journal

CytoCLIP: Learning Cytoarchitectural Characteristics in Developing Human Brain Using Contrastive Language Image Pre-Training.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: CytoCLIP: Learning Cytoarchitectural Characteristics in Developing Human Brain Using Contrastive Language Image Pre-Training.
Συγγραφείς: Ta P; Sudha Gopalakrishnan Brain Centre, Indian Institute of Technology Madras, Chennai, India. pralaypati@htic.iitm.ac.in.; Department of Electrical Engineering, Indian Institute of Technology Madras, Chennai, India. pralaypati@htic.iitm.ac.in., Venkatesaperumal S; Sudha Gopalakrishnan Brain Centre, Indian Institute of Technology Madras, Chennai, India., Ram K; Sudha Gopalakrishnan Brain Centre, Indian Institute of Technology Madras, Chennai, India., Sivaprakasam M; Sudha Gopalakrishnan Brain Centre, Indian Institute of Technology Madras, Chennai, India.; Department of Electrical Engineering, Indian Institute of Technology Madras, Chennai, India.
Πηγή: Neuroinformatics [Neuroinformatics] 2026 Jun 26; Vol. 24 (3). Date of Electronic Publication: 2026 Jun 26.
Τύπος έκδοσης: Journal Article
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: Humana Press, Inc Country of Publication: United States NLM ID: 101142069 Publication Model: Electronic Cited Medium: Internet ISSN: 1559-0089 (Electronic) Linking ISSN: 15392791 NLM ISO Abbreviation: Neuroinformatics Subsets: MEDLINE
Imprint Name(s): Original Publication: Totowa, NJ : Humana Press, Inc., c2003-
Ιατρικοί όροι (MeSH): Brain*/cytology , Brain*/growth & development , Brain*/diagnostic imaging , Brain*/embryology , Image Processing, Computer-Assisted*/methods , Language*, Humans
Περίληψη: The functions of different regions of the human brain are closely linked to their distinct cytoarchitecture, which is defined by the spatial arrangement and morphology of the cells. Identifying brain regions by their cytoarchitecture enables various scientific analyses of the brain. However, delineating these areas manually in brain histological sections is time-consuming and requires specialized knowledge. An automated approach is necessary to minimize the effort needed from human experts. To address this, we propose CytoCLIP, a suite of vision-language models derived from pre-trained Contrastive Language-Image Pre-Training (CLIP) frameworks to learn joint visual-text representations of brain cytoarchitecture. CytoCLIP comprises two model variants: one is trained using low-resolution whole-region images to understand the overall cytoarchitectural pattern of an area, and the other is trained on high-resolution image tiles for detailed cellular-level representation. The training dataset is created from NISSL-stained histological sections of developing fetal brains of different gestational weeks. It includes 86 distinct regions for low-resolution images and 379 brain regions for high-resolution tiles. We evaluate the model's understanding of the cytoarchitecture and generalization ability using region classification and cross-modal retrieval tasks. Multiple experiments are performed under various data setups, including data from samples of different ages and sectioning planes. Experimental results demonstrate that CytoCLIP outperforms existing methods. It achieves a weighted F1 score of 0.87 for whole-region classification and 0.91 for high-resolution image tile classification.
(© 2026. The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature.)
Competing Interests: Declarations. Competing interests: The authors declare no competing interests.
References: Amunts, K., & Zilles, K. (2015). Architectonic mapping of the human brain beyond brodmann. Neuron, 88(6), 1086–1107. https://doi.org/10.1016/j.neuron.2015.12.001. (PMID: 10.1016/j.neuron.2015.12.00126687219)
Brancati, N., Anniciello, A. M., Pati, P., Riccio, D., Scognamiglio, G., Jaume, G., De Pietro, G., Di Bonito, M., Foncubierta, A., Botti, G., Gabrani, M., Feroce, F., & Frucci, M. (2022). BRACS: A dataset for breast carcinoma subtyping in H&E histology images. Database: The Journal of Biological Databases and Curation. https://doi.org/10.1093/database/baac093. (PMID: 10.1093/database/baac093362517769575967)
Chen, R. J., Chen, C., Li, Y., Chen, T. Y., Trister, A. D., Krishnan, R. G., & Mahmood, F. (2022). Scaling vision transformers to gigapixel images via hierarchical self-supervised learning. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
Chen, R. J., Ding, T., Lu, M. Y., Williamson, D. F. K., Jaume, G., Song, A. H., Chen, B., Zhang, A., Shao, D., Shaban, M., Williams, M., Oldenburg, L., Weishaupt, L. L., Wang, J. J., Vaidya, A., Le, L. P., Gerber, G., Sahai, S., Williams, W., & Mahmood, F. (2024). Towards a general-purpose foundation model for computational pathology. Nature Medicine, 30(3), 850–862. https://doi.org/10.1038/s41591-024-02857-3. (PMID: 10.1038/s41591-024-02857-33850401811403354)
Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A simple framework for contrastive learning of visual representations. Proceedings of the 37th International Conference on Machine Learning, 1597–1607.
Ciga, O., Xu, T., & Martel, A. L. (2022). Self supervised contrastive learning for digital histopathology. Machine Learning With Applications, 7(100198), Article 100198. https://doi.org/10.1016/j.mlwa.2021.100198. (PMID: 10.1016/j.mlwa.2021.100198)
Das, A., Chaudhuri, M., Bhat, K., Ram, K., Bota, M., & Sivaprakasam, M. (2025). PosDiffAE: Position-aware diffusion auto-encoder for high-resolution brain tissue classification incorporating artifact restoration. IEEE Journal of Biomedical and Health Informatics. https://doi.org/10.1109/JBHI.2025.3567708. (PMID: 10.1109/JBHI.2025.356770841359720)
Ding, S.-L., Royall, J. J., Lesnar, P., Facer, B. A. C., Smith, K. A., Wei, Y., Brouner, K., Dalley, R. A., Dee, N., Dolbeare, T. A., Ebbert, A., Glass, I. A., Keller, N. H., Lee, F., Lemon, T. A., Nyhus, J., Pendergraft, J., Reid, R., Sarreal, M., … Lein, E. S. (2022). Cellular resolution anatomical and molecular atlases for prenatal human brains. Journal of Comparative Neurology, 530(1), 6–503. https://doi.org/10.1002/cne.25243. (PMID: 10.1002/cne.2524334525221)
Ding, S.-L., Royall, J. J., Sunkin, S. M., Ng, L., Facer, B. A. C., Lesnar, P., Guillozet-Bongaarts, A., McMurray, B., Szafer, A., Dolbeare, T. A., Stevens, A., Tirrell, L., Benner, T., Caldejon, S., Dalley, R. A., Dee, N., Lau, C., Nyhus, J., Reding, M., … Lein, E. S. (2016). Comprehensive cellular-resolution atlas of the adult human brain. Journal of Comparative Neurology, 524(16), 3127–3481. https://doi.org/10.1002/cne.24080. (PMID: 10.1002/cne.24080274182735054943)
Eslami, S., Meinel, C., & de Melo, G. (2023). PubMedCLIP: How much does CLIP benefit visual question answering in the medical domain? Findings of the Association for Computational Linguistics: EACL 2023, 1181–1193.
Fashi, P. A., Hemati, S., Babaie, M., Gonzalez, R., & Tizhoosh, H. R. (2022). A self-supervised contrastive learning approach for whole slide image representation in digital pathology. Journal of Pathology Informatics, 13(100133), Article 100133. https://doi.org/10.1016/j.jpi.2022.100133. (PMID: 10.1016/j.jpi.2022.100133366051149808093)
Hu, W., Li, X., Li, C., Li, R., Jiang, T., Sun, H., Huang, X., Grzegorzek, M., & Li, X. (2023). A state-of-the-art survey of artificial neural networks for whole-slide image analysis: From popular convolutional neural networks to potential visual transformers. Computers in Biology and Medicine, 161(107034), Article 107034. https://doi.org/10.1016/j.compbiomed.2023.107034. (PMID: 10.1016/j.compbiomed.2023.10703437230019)
Huang, Z., Bianchi, F., Yuksekgonul, M., Montine, T. J., & Zou, J. (2023). A visual-language foundation model for pathology image analysis using medical Twitter. Nature Medicine, 29(9), 2307–2316. https://doi.org/10.1038/s41591-023-02504-3. (PMID: 10.1038/s41591-023-02504-337592105)
Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., & Schmidt, L. (2025). OpenCLIP. Zenodo.
Kakkar, M., Shanbhag, D., Aladahalli, C., & M, G. R. (2024). Language augmentation in CLIP for improved anatomy detection on multi-modal medical images. Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference, 2024, 1–4. https://doi.org/10.1109/EMBC53108.2024.10781689.
Komura, D., & Ishikawa, S. (2018). Machine learning methods for histopathological image analysis. Computational and Structural Biotechnology Journal, 16, 34–42. https://doi.org/10.1016/j.csbj.2018.01.001. (PMID: 10.1016/j.csbj.2018.01.001302759366158771)
Kosaraju, S., Park, J., Lee, H., Yang, J. W., & Kang, M. (2022). Deep learning-based framework for slide-based histopathological image analysis. Scientific Reports, 12(1), 19075. https://doi.org/10.1038/s41598-022-23166-0. (PMID: 10.1038/s41598-022-23166-0363519979646838)
Loshchilov, I., & Hutter, F. (2019). Decoupled Weight Decay Regularization. 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6–9, 2019. https://openreview.net/forum?id=Bkg6RiCqY7.
Lu, M. Y., Chen, B., Williamson, D. F. K., Chen, R. J., Liang, I., Ding, T., Jaume, G., Odintsov, I., Le, L. P., Gerber, G., Parwani, A. V., Zhang, A., & Mahmood, F. (2024). A visual-language foundation model for computational pathology. Nature Medicine, 30(3), 863–874. https://doi.org/10.1038/s41591-024-02856-4. (PMID: 10.1038/s41591-024-02856-43850401711384335)
Lu, M. Y., Chen, B., Zhang, A., Williamson, D. F. K., Chen, R. J., Ding, T., Le, L. P., Chuang, Y. S., & Mahmood, F. (2023). Visual language pretrained multiple instance zero-shot transfer for histopathology images. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
Madabhushi, A., & Lee, G. (2016). Image analysis and machine learning in digital pathology: Challenges and opportunities. Medical Image Analysis, 33, 170–175. https://doi.org/10.1016/j.media.2016.06.037. (PMID: 10.1016/j.media.2016.06.037274234095556681)
Patel, T., El-Sayed, H., & Sarker, M. K. (2024). Evaluating Vision-Language Models for hematology image Classification: Performance Analysis of CLIP and its Biomedical AI Variants. 2024 36th Conference of Open Innovations Association (FRUCT), 578–584.
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning Transferable Visual Models From Natural Language Supervision. In M. Meila & T. Zhang (Eds.), Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 8748–8763). PMLR.
Rahaman, M. M., Millar, E. K. A., & Meijering, E. (2025). Generalized deep learning for histopathology image classification using supervised contrastive learning. Journal of Advanced Research, 75, 389–404. https://doi.org/10.1016/j.jare.2024.11.013. (PMID: 10.1016/j.jare.2024.11.01339551131)
Rottschy, C., Eickhoff, S. B., Schleicher, A., Mohlberg, H., Kujovic, M., Zilles, K., & Amunts, K. (2007). Ventral visual cortex in humans: Cytoarchitectonic mapping of two extrastriate areas. Human Brain Mapping, 28(10), 1045–1059. https://doi.org/10.1002/hbm.20348. (PMID: 10.1002/hbm.20348172661066871378)
Schiffer, C., Amunts, K., Harmeling, S., & Dickscheid, T. (2021a). Contrastive representation learning for whole brain cytoarchitectonic mapping in histological human brain Sect. 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI).
Schiffer, C., Spitzer, H., Kiwitz, K., Unger, N., Wagstyl, K., Evans, A. C., Harmeling, S., Amunts, K., & Dickscheid, T. (2021b). Convolutional neural networks for cytoarchitectonic brain mapping at large scale. NeuroImage, 240(118327), Article 118327. https://doi.org/10.1016/j.neuroimage.2021.118327. (PMID: 10.1016/j.neuroimage.2021.11832734224853)
SGBC-IITM (2025). DHARANI Dataset. Sudha Gopalakrishnan Brain Centre IIT Madras. https://brainportal.humanbrain.in/publicview/index.html.
Spitzer, H., Amunts, K., Harmeling, S., & Dickscheid, T. (2017). Parcellation of visual cortex on high-resolution histological brain sections using convolutional neural networks. 2017 IEEE 14th International Symposium on Biomedical Imaging (ISBI 2017).
Spitzer, H., Kiwitz, K., Amunts, K., Harmeling, S., & Dickscheid, T. (2018). Improving cytoarchitectonic segmentation of human brain areas with self-supervised Siamese networks. Medical Image Computing and Computer Assisted Intervention – MICCAI 2018 (pp. 663–671). Springer International Publishing.
Tan, J. W., & Jeong, W. K. (2023). Histopathology image classification using deep manifold contrastive learning. Lecture Notes in Computer Science (pp. 683–692). Springer Nature Switzerland.
Verma, R., Bota, M., Ram, K., Jayakumar, J., Folkerth, R., Pandurangan, K., Ramesh, J. J., Majumder, M., Raveendran, R., Nanda, R., K, S., S, A. D., Karthik, S., Kumarasami, R., S, S., Lata, S., Kumar, E. H., Rangasami, R., Srinivasan, C., & Sivaprakasam, M. (2025). DHARANI: A 3D developing human-brain atlas resource to advance neuroscience internationally integrated multimodal imaging and high-resolution histology of the second trimester. Journal of Comparative Neurology, 533(2), Article e70006. https://doi.org/10.1002/cne.70006.
Wang, Z., Wu, Z., Agarwal, D., & Sun, J. (2022). MedCLIP: Contrastive learning from unpaired medical images and text. Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing, 2022, 3876–3887. https://doi.org/10.18653/v1/2022.emnlp-main.256.
Wei, X., Kurtz, C., & Cloppet, F. (2025). Enhancing vision–language contrastive representation learning using domain knowledge. Computer Vision and Image Understanding, 259(104403), Article 104403. https://doi.org/10.1016/j.cviu.2025.104403. (PMID: 10.1016/j.cviu.2025.104403)
Weiner, K. S., Barnett, M. A., Lorenz, S., Caspers, J., Stigliani, A., Amunts, K., Zilles, K., Fischl, B., & Grill-Spector, K. (2017). The cytoarchitecture of domain-specific regions in human high-level visual cortex. Cerebral Cortex, 27(1), 146–161. https://doi.org/10.1093/cercor/bhw361. (PMID: 10.1093/cercor/bhw361279090035939223)
Yang, P., Hong, Z., Yin, X., Zhu, C., & Jiang, R. (2021). Self-supervised visual representation learning for histopathological images. Lecture Notes in Computer Science (pp. 47–57). Springer International Publishing.
Zhang, S., Xu, Y., Usuyama, N., Xu, H., Bagga, J., Tinn, R., Preston, S., Rao, R., Wei, M., Valluri, N., Wong, C., Tupini, A., Wang, Y., Mazzola, M., Shukla, S., Liden, L., Gao, J., Crabtree, A., Piening, B., … Poon, H. (2025). A multimodal biomedical foundation model trained from fifteen million image–text pairs. NEJM AI. https://doi.org/10.1056/aioa2400640. (PMID: 10.1056/aioa2400640)
Zhao, F., Wang, Z., Du, H., He, X., & Cao, X. (2023). Self-supervised triplet contrastive learning for classifying endometrial histopathological images. IEEE Journal of Biomedical and Health Informatics, 27(12), 5970–5981. https://doi.org/10.1109/JBHI.2023.3314663. (PMID: 10.1109/JBHI.2023.331466337698968)
Zhao, Z., Liu, Y., Wu, H., Wang, M., Li, Y., Wang, S., Teng, L., Liu, D., Cui, Z., Wang, Q., & Shen, D. (2025). CLIP in medical imaging: A survey. Medical Image Analysis, 102(103551), Article 103551. https://doi.org/10.1016/j.media.2025.103551. (PMID: 10.1016/j.media.2025.10355140127590)
Zilles, K., & Amunts, K. (2010). Centenary of Brodmann’s map–conception and fate. Nature Reviews. Neuroscience, 11(2), 139–145. https://doi.org/10.1038/nrn2776. (PMID: 10.1038/nrn277620046193)
Contributed Indexing: Keywords: CLIP; Contrastive learning; Cytoarchitecture; Histological Image processing
Entry Date(s): Date Created: 20260626 Date Completed: 20260626 Latest Revision: 20260626
Update Code: 20260627
DOI: 10.1007/s12021-026-09793-2
PMID: 42360543
Βάση Δεδομένων: MEDLINE
Περιγραφή
ISSN:1559-0089
DOI:10.1007/s12021-026-09793-2