Academic Journal

Combined evidence from artificial neural networks and human brain-lesion models reveals that language modulates vision in human perception.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Combined evidence from artificial neural networks and human brain-lesion models reveals that language modulates vision in human perception.
Συγγραφείς: Chen H; School of Psychological and Cognitive Sciences and Beijing Key Laboratory of Behavior and Mental Health, Peking University, Beijing, China.; State Key Laboratory of Cognitive Neuroscience and Learning & IDG/McGovern Institute for Brain Research, Beijing Normal University, Beijing, China., Liu B; Department of Radiology, First Hospital of Shanxi Medical University, Taiyuan, China.; College of Medical Imaging, Shanxi Medical University, Taiyuan, China.; Shanxi Key Laboratory of Intelligent Imaging, First Hospital of Shanxi Medical University, Taiyuan, China., Wang S; State Key Laboratory of Cognitive Neuroscience and Learning & IDG/McGovern Institute for Brain Research, Beijing Normal University, Beijing, China., Wang X; State Key Laboratory of Cognitive Neuroscience and Learning & IDG/McGovern Institute for Brain Research, Beijing Normal University, Beijing, China., Han W; School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China., Wang X; Department of Radiology, First Hospital of Shanxi Medical University, Taiyuan, China. wangxiaochun@sydyy.com.; College of Medical Imaging, Shanxi Medical University, Taiyuan, China. wangxiaochun@sydyy.com.; Shanxi Key Laboratory of Intelligent Imaging, First Hospital of Shanxi Medical University, Taiyuan, China. wangxiaochun@sydyy.com., Zhu Y; School of Psychological and Cognitive Sciences and Beijing Key Laboratory of Behavior and Mental Health, Peking University, Beijing, China. yixin.zhu@pku.edu.cn.; Institute for Artificial Intelligence, Peking University, Beijing, China. yixin.zhu@pku.edu.cn.; State Key Laboratory of General Artificial Intelligence, Peking University, Beijing, China. yixin.zhu@pku.edu.cn.; Key Laboratory of Machine Perception (Ministry of Education), Peking University, Beijing, China. yixin.zhu@pku.edu.cn., Bi Y; School of Psychological and Cognitive Sciences and Beijing Key Laboratory of Behavior and Mental Health, Peking University, Beijing, China. ybi@pku.edu.cn.; State Key Laboratory of Cognitive Neuroscience and Learning & IDG/McGovern Institute for Brain Research, Beijing Normal University, Beijing, China. ybi@pku.edu.cn.; Institute for Artificial Intelligence, Peking University, Beijing, China. ybi@pku.edu.cn.; Key Laboratory of Machine Perception (Ministry of Education), Peking University, Beijing, China. ybi@pku.edu.cn.; IDG/McGovern Institute for Brain Research, Peking University, Beijing, China. ybi@pku.edu.cn.
Πηγή: Nature human behaviour [Nat Hum Behav] 2026 Mar; Vol. 10 (3), pp. 615-631. Date of Electronic Publication: 2025 Dec 15.
Τύπος έκδοσης: Journal Article; Research Support, Non-U.S. Gov't
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: Springer Nature Publishing Country of Publication: England NLM ID: 101697750 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 2397-3374 (Electronic) Linking ISSN: 23973374 NLM ISO Abbreviation: Nat Hum Behav Subsets: MEDLINE
Imprint Name(s): Original Publication: [London] : Springer Nature Publishing, [2017]-
Ιατρικοί όροι (MeSH): Visual Perception*/physiology , Brain*/physiology , Neural Networks, Computer* , Language*, Temporal Lobe/physiology ; Stroke/physiopathology ; Humans ; Male ; Female ; Adult ; Magnetic Resonance Imaging ; Models, Neurological ; Middle Aged
Περίληψη: Comparing information structures in between deep neural networks (DNNs) and the human brain has become a key method for exploring their similarities and differences. Recent research has shown better alignment of vision-language DNN models, such as contrastive language-image pretraining (CLIP), with the activity of the human ventral occipitotemporal cortex (VOTC) than earlier vision models, supporting the idea that language modulates human visual perception. However, interpreting the results from such comparisons is inherently limited owing to the 'black box' nature of DNNs. Here we combine model-brain fitness analyses with human brain lesion data to examine how disrupting the communication pathway between the visual and language systems causally affects the ability of vision-language DNNs to explain the activity of the VOTC to address this. Across four diverse datasets, CLIP consistently captured unique variance in VOTC neural representations, relative to both label-supervised (ResNet) and unsupervised (MoCo) models. This advantage tended to be left-lateralized at the group level, aligning with the human language network. Analyses of 33 patients who experienced a stroke revealed that reduced white matter integrity between the VOTC and the language region in the left angular gyrus was correlated with decreased CLIP-brain correspondence and increased MoCo-brain correspondence, indicating a dynamic influence of language processing on the activity of the VOTC. These findings support the integration of language modulation in neurocognitive models of human vision, reinforcing concepts from vision-language DNN models. The sensitivity of model-brain similarity to specific brain lesions demonstrates that leveraging the manipulation of the human brain is a promising framework for evaluating and developing brain-like computer models.
(© 2025. The Author(s), under exclusive licence to Springer Nature Limited.)
Competing Interests: Competing interests: The authors declare no competing interests.
References: Bao, P., She, L., McGill, M. & Tsao, D. Y. A map of object space in primate inferotemporal cortex. Nature 583, 103–108 (2020). (PMID: 32494012808838810.1038/s41586-020-2350-5)
Schrimpf, M. et al. The neural architecture of language: Integrative modeling converges on predictive processing. Proc. Natl Acad. Sci. USA 118, e2105646118 (2021). (PMID: 34737231869405210.1073/pnas.2105646118)
Schrimpf, M. et al. Brain-score: which artificial neural network for object recognition is most brain-like? Preprint at bioRxiv https://doi.org/10.1101/407007 (2018).
Yamins, D. L. et al. Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proc. Natl Acad. Sci. USA 111, 8619–8624 (2014). (PMID: 24812127406070710.1073/pnas.1403112111)
Krizhevsky, A., Sutskever, I. & Hinton, G. E. Imagenet classification with deep convolutional neural networks. Adv. Neural Inf. Process. Syst. 25, 1097–1105 (2012).
He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. In Proc. IEEE Conference on Computer Vision and Pattern Recognition 770–778 (IEEE, 2016).
Kriegeskorte, N. et al. Matching categorical object representations in inferior temporal cortex of man and monkey. Neuron 60, 1126–1141 (2008). (PMID: 19109916314357410.1016/j.neuron.2008.10.043)
Ungerleider, L. G. & Haxby, J. V. ‘What’ and ‘where’in the human brain. Curr. Opin. Neurobiol. 4, 157–165 (1994). (PMID: 803857110.1016/0959-4388(94)90066-3)
Dobs, K., Martinez, J., Kell, A. J. E. & Kanwisher, N. Brain-like functional specialization emerges spontaneously in deep neural networks. Sci. Adv. 8, eabl8913 (2022). (PMID: 35294241892634710.1126/sciadv.abl8913)
Konkle, T. & Alvarez, G. A. A self-supervised domain-general learning framework for human ventral stream representation. Nat. Commun. 13, 491 (2022). (PMID: 35078981878981710.1038/s41467-022-28091-4)
Vinken, K., Prince, J. S., Konkle, T. & Livingstone, M. S. The neural code for ‘face cells’ is not face-specific. Sci. Adv. 9, eadg1736 (2023). (PMID: 376474001046812310.1126/sciadv.adg1736)
Prince, J. S., Alvarez, G. A. & Konkle, T. Contrastive learning explains the emergence and function of visual category-selective regions. Sci. Adv. 10, eadl1776 (2024). (PMID: 393213041142389610.1126/sciadv.adl1776)
Wang, A. Y., Kay, K., Naselaris, T., Tarr, M. J. & Wehbe, L. Better models of human high-level visual cortex emerge from natural language supervision with a large and diverse dataset. Nat. Mach. Intell. 5, 1415–1426 (2023). (PMID: 10.1038/s42256-023-00753-y)
Zhou, Q., Du, C., Wang, S. & He, H. CLIP-MUSED: CLIP-guided multi-subject visual neural information semantic decoding. In Proc. 12th International Conference on Learning Representations (eds Kim, B. et al.) https://openreview.net/pdf?id=lKxL5zkssv (ICLR, 2024).
Doerig, A. et al. High-level visual representations in the human brain are aligned with large language models. Nat. Mach.Intell. 7, 1220–1234 (2025). (PMID: 408424851236471010.1038/s42256-025-01072-0)
Conwell, C., Prince, J. S., Hamblin, C. J. & Alvarez, G. A. Controlled assessment of CLIP-style language-aligned vision models in prediction of brain and behavioral data. In ICLR 2023 Workshop on Mathematical and Empirical Understanding of Foundation Models (ME-FoMo, 2023).
Luo, A. F., Henderson, M. M., Wehbe, L. & Tarr, M. J. Brain diffusion for visual exploration: cortical discovery using large-scale generative models. In Proc. 37th International Conference on Neural Information Processing Systems (eds Oh, A. et al.) 75740–75781 (Curran Associates, 2023).
Luo, A. F., Henderson, M. M., Tarr, M. J. & Wehbe, L. BrainSCUBA: fine-grained natural language captions of visual cortex selectivity. In Proc. 12th International Conference on Learning Representations (eds Kim, B. et al.) https://openreview.net/pdf?id=mQYHXUUTkU (ICLR, 2024).
Lupyan, G. The centrality of language in human cognition. Lang. Learn. 66, 516–553 (2016). (PMID: 10.1111/lang.12155)
Thierry, G. Neurolinguistic relativity: how language flexes human perception and cognition. Lang. Learn. 66, 690–713 (2016). (PMID: 27642191500688210.1111/lang.12186)
Gilbert, A. L., Regier, T., Kay, P. & Ivry, R. B. Whorf hypothesis is supported in the right visual field but not the left. Proc. Natl Acad. Sci. USA 103, 489–494 (2006). (PMID: 1638784810.1073/pnas.0509868103)
Drivonikou, G. V. et al. Further evidence that Whorfian effects are stronger in the right visual field than the left. Proc. Natl Acad. Sci. USA 104, 1097–1102 (2007). (PMID: 17213312178337010.1073/pnas.0610132104)
Winawer, J. et al. Russian blues reveal effects of language on color discrimination. Proc. Natl Acad. Sci. USA 104, 7780–7785 (2007). (PMID: 17470790187652410.1073/pnas.0701644104)
Ting Siok, W. et al. Language regions of brain are operative in color perception. Proc. Natl Acad. Sci. USA 106, 8140–8145 (2009). (PMID: 19416812268888810.1073/pnas.0903627106)
Martinovic, J., Paramei, G. V. & MacInnes, W. J. Russian blues reveal the limits of language influencing colour discrimination. Cognition 201, 104281 (2020). (PMID: 3227623610.1016/j.cognition.2020.104281)
Fedorenko, E., Piantadosi, S. T. & Gibson, E. A. Language is primarily a tool for communication rather than thought. Nature 630, 575–586 (2024). (PMID: 3889829610.1038/s41586-024-07522-w)
Maier, M. & Abdel Rahman, R. No matter how: top-down effects of verbal and semantic category knowledge on early visual perception. Cogn. Affect. Behav. Neurosci. 19, 859–876 (2019). (PMID: 3060783110.3758/s13415-018-00679-8)
Conwell, C., Prince, J. S., Kay, K. N., Alvarez, G. A. & Konkle, T. A large-scale examination of inductive biases shaping high-level visual representation in brains and machines. Nat. Commun. 15, 9383 (2024). (PMID: 394779231152613810.1038/s41467-024-53147-y)
Allen, E. J. et al. A massive 7T fMRI dataset to bridge cognitive neuroscience and artificial intelligence. Nat. Neurosci. 25, 116–126 (2022). (PMID: 3491665910.1038/s41593-021-00962-x)
Conwell, C. et al. Monkey See, model knew: large language models accurately predict visual brain responses in humans and non-human primates. Preprint at bioRxiv https://doi.org/10.1101/2025.03.05.641284 (2025).
Kriegeskorte, N., Mur, M. & Bandettini, P. A. Representational similarity analysis-connecting the branches of systems neuroscience. Front. Syst. Neurosci. 2, 249 (2008).
Fu, Z. et al. Different computational relations in language are captured by distinct brain systems. Cereb. Cortex. 33, 997–1013 (2023). (PMID: 3533291410.1093/cercor/bhac117)
Liu, B. et al. Object knowledge representation in the human visual cortex requires a connection with the language system. PLoS Biol. 23, e3003161 (2025). (PMID: 403928021209177010.1371/journal.pbio.3003161)
Hebart, M. N. et al. THINGS-data, a multimodal collection of large-scale datasets for investigating object representations in human brain and behavior. Elife 12, e82580 (2023). (PMID: 368473391003866210.7554/eLife.82580)
Güntürkün, O., Ströckens, F. & Ocklenburg, S. Brain lateralization: a comparative perspective. Physiol. Rev. 100, 1019–1063 (2020). (PMID: 3223391210.1152/physrev.00006.2019)
Wilke, M. & Lidzba, K. LI-tool: a new toolbox to assess lateralization in functional MR-data. J. Neurosci. Methods 163, 128–136 (2007). (PMID: 1738694510.1016/j.jneumeth.2007.01.026)
Seghier, M. L. Laterality index in functional MRI: methodological issues. Magn. Reson. Imaging 26, 594–601 (2008). (PMID: 1815822410.1016/j.mri.2007.10.010)
Fedorenko, E., Hsieh, P. J., Nieto-Castañón, A., Whitfield-Gabrieli, S. & Kanwisher, N. New method for fMRI investigations of language: defining ROIs functionally in individual subjects. J. Neurophysiol. 104, 1177–1194 (2010). (PMID: 20410363293492310.1152/jn.00032.2010)
Oliva, A. & Torralba, A. Modeling the shape of the scene: a holistic representation of the spatial envelope. Int. J. Comput. Vis. 42, 145–175 (2001). (PMID: 10.1023/A:1011139631724)
Hua, K. et al. Tract probability maps in stereotaxic spaces: analyses of white matter anatomy and tract-specific quantification. Neuroimage 39, 336–347 (2008). (PMID: 1793189010.1016/j.neuroimage.2007.07.053)
Brown, T. B. et al. Language models are few-shot learners. In Proc. 34th International Conference on Neural Information Processing Systems: Advances in Neural Information Processing Systems Vol. 33 (eds Larochelle, H. et al.) 1877–1901 (Curran Associates, Inc., 2020).
Kaplan, J. et al. Scaling laws for neural language models. Preprint at https://arxiv.org/abs/2001.08361 (2020).
Mu, N., Kirillov, A., Wagner, D. & Xie, S. SLIP: self-supervision meets language-image pre-training. In European Conference on Computer Vision (eds Avidan, S. et al.) 529–544 (Springer, 2022).
Radford, A. et al. Learning transferable visual models from natural language supervision. In Proc. 38th International Conference on Machine Learning: Proc. Machine Learning Research Vol. 139 (eds Meila, M. & Zhang, T.) 8748–8763 (PMLR, 2021).
Gelman, S. A. & Roberts, S. O. How language shapes the cultural inheritance of categories. Proc. Natl Acad. Sci. USA 114, 7900–7907 (2017). (PMID: 28739931554427810.1073/pnas.1621073114)
Unger, L. & Fisher, A. V. The emergence of richly organized semantic knowledge from simple statistics: a synthetic review. Dev. Rev. 60, 100949 (2021). (PMID: 33840880802614410.1016/j.dr.2021.100949)
Xu, Y., He, Y. & Bi, Y. A tri-network model of human semantic processing. Front. Psychol. 8, 1538 (2017). (PMID: 28955266560090510.3389/fpsyg.2017.01538)
Seghier, M. L. The angular gyrus: multiple functions and multiple subdivisions. Neuroscientist 19, 43–61 (2013). (PMID: 2254753010.1177/1073858412440596)
Xu, Y. et al. Doctor, teacher, and stethoscope: neural representation of different types of semantic relations. J. Neurosci. 38, 3303–3317 (2018). (PMID: 29476016659606010.1523/JNEUROSCI.2562-17.2018)
Schwartz, M. F. et al. Neuroanatomical dissociation for taxonomic and thematic knowledge in the human brain. Proc. Natl Acad. Sci. USA 108, 8520–8524 (2011). (PMID: 21540329310092810.1073/pnas.1014935108)
Zhang, W., Xiang, M. & Wang, S. The role of left angular gyrus in the representation of linguistic composition relations. Hum. Brain Mapp. 43, 2204–2217 (2022). (PMID: 35064707899636210.1002/hbm.25781)
Price, A. R., Bonner, M. F., Peelle, J. E. & Grossman, M. Converging evidence for the neuroanatomic basis of combinatorial semantics in the angular gyrus. J. Neurosci. 35, 3276–3284 (2015). (PMID: 25698762433163910.1523/JNEUROSCI.3446-14.2015)
Lupyan, G., Rahman, R. A., Boroditsky, L. & Clark, A. Effects of language on visual perception. Trends Cogn. Sci. 24, 930–944 (2020). (PMID: 3301268710.1016/j.tics.2020.08.005)
Mattioni, S. et al. Categorical representation from sound and sight in the ventral occipito-temporal cortex of sighted and blind. Elife 9, e50732 (2020). (PMID: 32108572710886610.7554/eLife.50732)
van den Hurk, J., Van Baelen, M. & Op de Beeck, H. P. Development of visual category selectivity in ventral visual cortex does not require visual experience. Proc. Natl Acad. Sci. USA 114, E4501–E4510 (2017). (PMID: 285071275465914)
Wang, X. et al. How visual is the visual cortex? Comparing connectional and functional fingerprints between congenitally blind and sighted individuals. J. Neurosci. 35, 12545–12559 (2015). (PMID: 26354920660540510.1523/JNEUROSCI.3914-14.2015)
Ricciardi, E., Bonino, D., Pellegrini, S. & Pietrini, P. Mind the blind brain to understand the sighted one! Is there a supramodal cortical functional architecture?. Neurosci. Biobehav. Rev. 41, 64–77 (2014). (PMID: 2415772610.1016/j.neubiorev.2013.10.006)
Bi, Y., Wang, X. & Caramazza, A. Object domain and modality in the ventral visual pathway. Trends Cogn. Sci. 20, 282–290 (2016). (PMID: 2694421910.1016/j.tics.2016.02.002)
Peelen, M. V. & Downing, P. E. Category selectivity in human visual cortex: beyond visual object recognition. Neuropsychologia 105, 177–183 (2017). (PMID: 2837716110.1016/j.neuropsychologia.2017.03.033)
Mahon, B. Z. et al. Action-related properties shape object representations in the ventral stream. Neuron 55, 507–520 (2007). (PMID: 17678861200082410.1016/j.neuron.2007.07.011)
Striem-Amit, E. et al. Functional connectivity of visual cortex in the blind follows retinotopic organization principles. Brain 138, 1679–1695 (2015). (PMID: 25869851461414210.1093/brain/awv083)
Burton, H., Snyder, A. Z. & Raichle, M. E. Resting state functional connectivity in early blind humans. Front. Syst. Neurosci. 8, 51 (2014). (PMID: 24778608398501910.3389/fnsys.2014.00051)
Ashburner, J. & Friston, K. J. Unified segmentation. Neuroimage 26, 839–851 (2005). (PMID: 1595549410.1016/j.neuroimage.2005.02.018)
Chen, X., Xie, S. & He, K. An empirical study of training self-supervised vision transformers. In Proc. IEEE/CVF International Conference on Computer Vision 9640–9649 (IEEE, 2021).
Kriegeskorte, N., Goebel, R. & Bandettini, P. Information-based functional brain mapping. Proc. Natl Acad. Sci. USA 103, 3863–3868 (2006). (PMID: 16537458138365110.1073/pnas.0600244103)
Vallat, R. Pingouin: statistics in Python. J. Open Source Softw. 3, 1026 (2018). (PMID: 10.21105/joss.01026)
Fonov, V. et al. Unbiased average age-appropriate atlases for pediatric studies. Neuroimage 54, 313–327 (2011). (PMID: 2065603610.1016/j.neuroimage.2010.07.033)
Xia, M., Wang, J. & He, Y. BrainNet Viewer: a network visualization tool for human brain connectomics. PLoS ONE 8, e68910 (2013). (PMID: 23861951370168310.1371/journal.pone.0068910)
Yan, C. G., Wang, X. D., Zuo, X. N. & Zang, Y. F. DPABI: data processing and analysis for (resting-state) brain imaging. Neuroinformatics 14, 339–351 (2016). (PMID: 2707585010.1007/s12021-016-9299-4)
Chen, H. Language modulates vision: evidence from neural networks and human brain-lesion models. figshare https://doi.org/10.6084/m9.figshare.29531288.v3 (2025).
Stoinski, L. M., Perkuhn, J. & Hebart, M. N. THINGSplus: new norms and metadata for the THINGS database of 1854 object concepts and 26,107 natural object images. Behav. Res. 56, 1583–1603 (2024). (PMID: 10.3758/s13428-023-02110-8)
Grant Information: 32171052 National Natural Science Foundation of China (National Science Foundation of China); 62406020 National Natural Science Foundation of China (National Science Foundation of China); 62376009 National Natural Science Foundation of China (National Science Foundation of China); 31925020; 82021004 National Natural Science Foundation of China (National Science Foundation of China)
Entry Date(s): Date Created: 20251216 Date Completed: 20260326 Latest Revision: 20260724
Update Code: 20260725
DOI: 10.1038/s41562-025-02357-5
PMID: 41398465
Βάση Δεδομένων: MEDLINE
Περιγραφή
ISSN:2397-3374
DOI:10.1038/s41562-025-02357-5