Abstract
Large language models (LLMs) produce seemingly meaningful outputs, yet they are trained on text alone without direct interaction with the world. This leads to a modern variant of the classical symbol grounding problem in AI: can LLMs' internal states and outputs be about extra-linguistic reality, independently of the meaning human interpreters project onto them? We argue that they can. We first distinguish referential grounding—the connection between a representation and its worldly referent—from other forms of grounding and argue it is the only kind essential to solving the problem. We contend that referential grounding is achieved when a system's internal states satisfy two conditions derived from teleosemantic theories of representation: (1) they stand in appropriate causal-informational relations to the world, and (2) they have a history of selection that has endowed them with the function of carrying this information. We argue that LLMs can meet both conditions, even without multimodality or embodiment.
References
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., … Zeng, A. (2022). Do as I can, not as I say: Grounding language in robotic affordances (No. arXiv:2204.01691). arXiv. https://doi.org/10.48550/arXiv.2204.01691
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T., Clark, J., … Kaplan, J. (2021). A general language assistant as a laboratory for alignment (No. arXiv:2112.00861). arXiv. https://doi.org/10.48550/arXiv.2112.00861
Barsalou, L. W. (1999). Perceptual symbol systems. Behavioral and Brain Sciences, 22, 577–660. https://doi.org/10.1017/s0140525x99002149
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021, March). On the dangers of stochastic parrots. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. https://doi.org/10.1145/3442188.3445922
Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.463
Block, N. (1986). Advertisement for a semantics for psychology. Midwest Studies in Philosophy, 10, 615–678. https://doi.org/10.1111/j.1475-4975.1987.tb00558.x
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., … Liang, P. (2022). On the opportunities and risks of foundation models (No. arXiv:2108.07258). arXiv. https://doi.org/10.48550/arXiv.2108.07258
Bordes, F., Pang, R. Y., Ajay, A., Li, A. C., Bardes, A., Petryk, S., Mañas, O., Lin, Z., Mahmoud, A., Jayaraman, B., Ibrahim, M., Hall, M., Xiong, Y., Lebensold, J., Ross, C., Jayakumar, S., Guo, C., Bouchacourt, D., Al-Tahan, H., … Chandra, V. (2024). An introduction to vision-language modeling (No. arXiv:2405.17247). arXiv. https://doi.org/10.48550/arXiv.2405.17247
Borg, E. (2025). LLMs, turing tests and Chinese rooms: The prospects for meaning in large language models. Inquiry, 1–31. https://doi.org/10.1080/0020174X.2024.2446241
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Driessche, G. B. V. D., Lespiau, J.-B., Damoc, B., Clark, A., Casas, D. D. L., Guy, A., Menick, J., Ring, R., Hennigan, T., Huang, S., Maggiore, L., Jones, C., Cassirer, A., … Sifre, L. (2022). Improving language models by retrieving from trillions of tokens. Proceedings of the 39th International Conference on Machine Learning, 2206–2240.
Brandom, R. (1994). Making it explicit: Reasoning, representing, and discursive commitment. Harvard University Press.
Brennan, S. E. (1998). The grounding problem in conversations with and through computers. In S. R. Fussell & R. J. Kreuz (Eds.), Social and cognitive approaches to interpersonal communication (pp. 201–225). Lawrence Erlbaum.
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., & Zhang, Y. (2023). Sparks of artificial general intelligence: Early experiments with GPT-4 (No. arXiv:2303.12712). arXiv. https://doi.org/10.48550/arXiv.2303.12712
Butlin, P. (2021). Sharing our concepts with machines. Erkenntnis, 88(7), 3079–3095. https://doi.org/10.1007/s10670-021-00491-w
Cappelen, H., & Dever, J. (2021). Making AI intelligible: Philosophical foundations. Oxford University Press.
Chalmers, D. J. (2023). Does thought require sensory grounding? From pure thinkers to large language models. Proceedings and Addresses of the American Philosophical Association, 97, 22–45. https://doi.org/10.1098/rsta.2022.0041
Chalmers, D. J. (2025). Propositional interpretability in artificial intelligence (No. arXiv:2501.15740). ar22iv. https://doi.org/10.48550/arXiv.2501.15740
Chandu, K. R., Bisk, Y., & Black, A. W. (2021). Grounding “grounding” in NLP. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 4283–4305. https://doi.org/10.18653/v1/2021.findings-acl.375
Chomsky, N. (1957). Syntactic structures. Mouton.
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30.
Clark, H. H., & Brennan, S. E. (1991). Grounding in communication. In Perspectives on socially shared cognition (pp. 127–149). American Psychological Association. https://doi.org/10.1037/10096-006
Coelho Mollo, D. (2015). Being clear on content. Philosophia, 43(3), 687–699. https://doi.org/10.1007/s11406-015-9622-6
Coelho Mollo, D. (2022). Deflationary realism: Representation and idealisation in cognitive science. Mind & Language, 37(5), 1048–1066. https://doi.org/10.1111/mila.12364
de Saussure, F. (1916). Cours de linguistique générale. Payot.
Delétang, G., Ruoss, A., Duquenne, P.-A., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L. K., Aitchison, M., Orseau, L., Hutter, M., & Veness, J. (2024). Language modeling is compression (No. arXiv:2309.10668). arXiv. https://doi.org/10.48550/arXiv.2309.10668
Di Maro, M. (2021). Computational grounding: An overview of common ground applications in conversational agents. IJCoL. Italian Journal of Computational Linguistics, 7(1 | 2), 133–156. https://doi.org/10.4000/ijcol.890
Firth, J. R. (1957). Papers in linguistics, 1934-1951. Oxford University Press.
Fodor, J. A. (1990). A theory of content and other essays. MIT Press.
Garson, J. (2019). What biological functions are and why they matter. Cambridge University Press. https://doi.org/10.1017/9781108560764
Godfrey-Smith, P. (1994). A modern history theory of functions. Noûs, 28(3), 344–362.
Goldstein, S., & Levinstein, B. A. (2024). Does ChatGPT have a mind? (No. arXiv:2407.11015). arXiv. https://doi.org/10.48550/arXiv.2407.11015
Grindrod, J. (2024). Large language models and linguistic intentionality. Synthese, 204(2), 71. https://doi.org/10.1007/s11229-024-04723-8
Grzankowski, A. (2024). Real sparks of artificial intelligence and the importance of inner interpretability. Inquiry, 0(0), 1–27. https://doi.org/10.1080/0020174X.2023.2296468
Haiman, J. (1985). Iconicity in syntax (Vol. 6). John Benjamins Publishing.
Harding, J. (2023). Operationalising representation in natural language processing (No. 2306.08193). arXiv. https://doi.org/10.1086/728685
Harnad, S. (1990). The symbol-grounding problem. Physica D, 42, 335–346. https://doi.org/10.1016/0167-2789(90)90087-6
Harris, Z. S. (1954). Distributional structure. Word, 10, 146–162. https://doi.org/10.1080/00437956.1954.11659520
Heyes, C. M. (2018). Cognitive gadgets: The cultural evolution of thinking. The Belknap Press of Harvard University Press.
jylin04, JackS, Karvonen, A., & Can. (2024). OthelloGPT learned a bag of heuristics. https://www.lesswrong.com/posts/gcpNuEZnxAPayaKBY/othellogpt-learned-a-bag-of-heuristics-1
Kim, M. J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., Vuong, Q., Kollar, T., Burchfiel, B., Tedrake, R., Sadigh, D., Levine, S., Liang, P., & Finn, C. (2024). OpenVLA: An open-source vision-language-action model (No. arXiv:2406.09246). arXiv. https://doi.org/10.48550/arXiv.2406.09246
LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444. https://doi.org/10.1038/nature14539
Lederman, H., & Mahowald, K. (2024). Are language models more like libraries or like librarians? Bibliotechnism, the novel reference problem, and the attitudes of LLMs. https://arxiv.org/abs/2401.04854v1
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., & Wattenberg, M. (2023). Emergent world representations: Exploring a sequence model trained on a synthetic task (No. arXiv:2210.13382). arXiv. https://doi.org/10.48550/arXiv.2210.13382
Mandelkern, M., & Linzen, T. (2024). Do language models’ words refer? Computational Linguistics, 1–10. https://doi.org/10.1162/coli_a_00522
Mcclelland, J. L., Rumelhart, D. E., & Group, P. R. (1986). Parallel distributed processing, explorations in the microstructure of cognition. MIT Press.
McCoy, R. T., Yao, S., Friedman, D., Hardy, M. D., & Griffiths, T. L. (2024). Embers of autoregression show how large language models are shaped by the problem they are trained to solve. Proceedings of the National Academy of Sciences, 121(41). https://doi.org/10.1073/pnas.2322420121
Merrill, W., Wu, Z., Naka, N., Kim, Y., & Linzen, T. (2024). Can you learn semantics through next-word prediction? The case of entailment. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024 (pp. 2752–2773). Association for Computational Linguistics. https://doi.org/10.18653/V1/2024.FINDINGS-ACL.161
Millière, R. (forthcoming). Language models as models of language. In R. Nefdt, G. Dupre, & K. Stanton (Eds.), The Oxford Handbook of the Philosophy of Linguistics. Oxford University Press.
Millière, R., & Buckner, C. (2024). A philosophical introduction to language models – part I: Continuity with classic debates (No. arXiv:2401.03910). arXiv. https://doi.org/10.48550/arXiv.2401.03910
Millière, R., & Buckner, C. (2025). Interventionist methods for interpreting deep neural networks. In G. Piccinini (Ed.), Neurocognitive Foundations of Mind. Routledge.
Millikan, R. G. (1989). In defense of proper functions. Philosophy of Science, 56(2), 288–302. https://doi.org/10.1086/289488
Millikan, R. G. (2017). Beyond concepts: Unicepts, language, and natural information. Oxford University Press.
Miracchi Titus, L. (2024). Does ChatGPT have semantic understanding? A problem with the statistics-of-occurrence strategy. Cognitive Systems Research, 83, 101174. https://doi.org/10.1016/j.cogsys.2023.101174
Montague, R. (1970). Universal grammar. Theoria, 36(3), 373–398. https://doi.org/10.1111/j.1755-2567.1970.tb00434.x
Nanda, N., Lee, A., & Wattenberg, M. (2023). Emergent linear representations in world models of self-supervised sequence models. In Y. Belinkov, S. Hao, J. Jumelet, N. Kim, A. McCarthy, & H. Mohebbi (Eds.), Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLPEMNLP 2023, Singapore, December 7, 2023 (pp. 16–30). Association for Computational Linguistics. https://doi.org/10.18653/V1/2023.BLACKBOXNLP-1.2
Neander, K. (1991). Functions as selected effects: The conceptual analyst’s defence. Philosophy of Science, 58(2), 168–184. https://doi.org/10.1086/289610
Neander, K. (2017). A mark of the mental: In defense of informational semantics. The MIT Press.
O’Neill, A., Rehman, A., Gupta, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., Jain, A., Tung, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Gupta, A., Wang, A., … Lin, Z. (2024). Open X-embodiment: Robotic learning datasets and RT-X models (No. arXiv:2310.08864). arXiv. https://doi.org/10.48550/arXiv.2310.08864
OpenAI. (2023). GPT-4 technical report (No. arXiv:2303.08774). arXiv. https://doi.org/10.48550/arXiv.2303.08774
Oswald, J. von, Schlegel, M., Meulemans, A., Kobayashi, S., Niklasson, E., Zucchet, N., Scherrer, N., Miller, N., Sandler, M., Arcas, B. A. y, Vladymyrov, M., Pascanu, R., & Sacramento, J. (2024). Uncovering mesa-optimization algorithms in transformers (No. arXiv:2309.05858). arXiv. https://doi.org/10.48550/arXiv.2309.05858
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.
Partee, B. H. (1981). Montague grammar, mental representations, and reality. In S. Kanger & S. Ōhman (Eds.), Philosophy and Grammar: Papers on the Occasion of the Quincentennial of Uppsala University (pp. 59–78). Springer Netherlands. https://doi.org/10.1007/978-94-009-9012-8_5
Pavlick, E. (2022). Semantic structure in deep learning. Annual Review of Linguistics, 8(1), 447–471. https://doi.org/10.1146/annurev-linguistics-031120-122924
Pepp, J. (2025). Reference without intentions in large language models. Inquiry, 1–19. https://doi.org/10.1080/0020174X.2024.2448482
Perniss, P., & Vigliocco, G. (2014). The bridge of iconicity: From a world of experience to the experience of language. Philosophical Transactions of the Royal Society B: Biological Sciences, 369(1651), 20130300. https://doi.org/10.1098/rstb.2013.0300
Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., & Riedel, S. (2019). Language models as knowledge bases? (No. arXiv:1909.01066). arXiv. https://doi.org/10.48550/arXiv.1909.01066
Piantadosi, S. T., & Hill, F. (2022). Meaning without reference in large language models. arXiv. https://doi.org/10.48550/ARXIV.2208.02957
Poon, H. (2013). Grounded unsupervised semantic parsing. Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 933–943.
Pulvermüller, F. (1999). Words in the brain’s language. Behavioural and Brain Sciences, 22, 253–336.
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. arXiv:2103.00020 [Cs]. https://arxiv.org/abs/2103.00020
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2024). Direct preference optimization: Your language model is secretly a reward model (No. arXiv:2305.18290). arXiv. https://doi.org/10.48550/arXiv.2305.18290
Searle, J. R. (1980). Minds, brains and programs. Behavioral and Brain Sciences, 3(3), 417–457. https://doi.org/10.1017/s0140525x00005781
Shea, N. (2012). Millikan’s isomorphism requirement. In D. Ryder, J. Kingsbury, & K. Williford (Eds.), Millikan and her critics (pp. 63–86). Wiley.
Shea, N. (2018). Representation in cognitive science. Oxford University Press.
Shortliffe, E. H. (1977). Mycin: A knowledge-based computer program applied to infectious diseases. Proceedings of the Annual Symposium on Computer Application in Medical Care, 66–69.
Simon, H. A., & Newell, A. (1971). Human problem solving: The state of the theory in 1970. American Psychologist, 26(2), 145–159. https://doi.org/10.1037/h0030806
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Power, A., Ray, A., Warstadt, A., Kocurek, A. W., Safaya, A., Tazarv, A., … Wu, Z. (2023). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research. https://openreview.net/forum?id=uyTL5Bvosj
Stalnaker, R. (2002). Common ground. Linguistics and Philosophy, 25(5-6), 701–721. https://doi.org/10.1023/a:1020867916902
Sterelny, K. (2014). The evolved apprentice: How evolution made humans unique (1st paperback edition). The MIT Press.
Stoljar, D., & Zhang, Z. V. (2025). Why ChatGPT doesn’t think: An argument from rationality. Inquiry, 1–29. https://doi.org/10.1080/0020174X.2024.2427061
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., Li, Y., Lee, H., Zheng, H. S., Ghafouri, A., Menegali, M., Huang, Y., Krikun, M., Lepikhin, D., Qin, J., … Le, Q. (2022). LaMDA: Language models for dialog applications (No. arXiv:2201.08239). arXiv. https://doi.org/10.48550/arXiv.2201.08239
Tomasello, M. (2009). Cultural origins of human cognition. Harvard University Press.
Traum, D. R. (1994). A computational theory of grounding in natural language conversation [Ph.D. Thesis]. University of Rochester.
Tsai, C.-T., & Roth, D. (2016). Concept grounding to multiple knowledge bases via indirect supervision. Transactions of the Association for Computational Linguistics, 4(0), 141–154. https://doi.org/10.1162/tacl_a_00089
Van Langendonck, W. (2010). Iconicity. In D. Geeraerts & H. Cuyckens (Eds.), The Oxford Handbook of Cognitive Linguistics (p. 0). Oxford University Press. https://doi.org/10.1093/oxfordhb/9780199738632.013.0016
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. arXiv. https://doi.org/10.48550/ARXIV.1706.03762
Williams, I. (2025). Can structural correspondences ground real world representational content in large language models? https://arxiv.org/abs/2506.16370
Wright, L. (1973). Functions. The Philosophical Review, 82(2), 139–168.
Wu, X., & Varshney, L. R. (2024). A meta-learning perspective on transformers for causal language modeling. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024 (pp. 15612–15622). Association for Computational Linguistics. https://doi.org/10.18653/V1/2024.FINDINGS-ACL.922
Yildirim, I., & Paul, L. A. (2023). From task structures to world models: What do LLMs know? Trends in Cognitive Sciences. https://doi.org/10.48550/arXiv.2310.04276
Zheng, C., Huang, W., Wang, R., Wu, G., Zhu, J., & Li, C. (2024). On mesa-optimization in autoregressively trained transformers: Emergence and capability. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, & C. Zhang (Eds.), Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024.
Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., Wahid, A., Vuong, Q., Vanhoucke, V., Tran, H. T., Soricut, R., Singh, A., Singh, J., Sermanet, P., Sanketi, P. R., Salazar, G., … Han, K. (2023). RT-2: Vision-language-action models transfer web knowledge to robotic control. In J. Tan, M. Toussaint, & K. Darvish (Eds.), Conference on Robot Learning, CoRL 2023, 6-9 November 2023, Atlanta, GA, USA (Vol. 229, pp. 2165–2183). PMLR.

This work is licensed under a Creative Commons Attribution 4.0 International License.
Copyright (c) 2026 Dimitri Coelho Mollo, Raphaël Millière
