Does ChatGPT think like humans?

Reviewers: Two anonymous reviewers

Editorial Assistant: Maren Giersiepen

AIs like ChatGPT are trained only on texts and do not actually experience the world. It’s been argued that for this reason, AI’s cannot really ‘understand’. This is called the symbol grounding problem: symbols have to be connected to perception and action. To address the question whether AIs can think like humans, this article looks at research on cognition and language.

Have you ever considered whether ChatGPT has thoughts? If it did, what would they be made of? Now what about your own thoughts? What are they made of? By answering this second question, we can come to learn much about the first and about the potential future of AI.

First, it is important to establish what “AI” even entails. When we talk about a system being an artificial intelligence (AI), it is typically because it can answer questions about seemingly any topic and respond like a knowledgeable human. For comparison, a calculator, despite performing arithmetic faster and more exactly than any human, is not considered AI. That’s because despite being able to multiply 28x354^2, it cannot tell you whether that number is more or less than the fingers on a human hand. For this reason, a classic test for identifying AI has been the so-called Turing Test: if someone were unaware that they were talking to an artificial system, could they confuse it for a human? ChatGPT, but not a calculator, can often manage this feat [1].

How do large language models work?

Currently-available AIs, like ChatGPT, are usually large language models. These belong to a group of methods called neural networks, statistical models inspired by the way neurons work in the brain. These statistical models have many nodes (roughly, statistical units inspired by neurons) organized into layers. In each layer, many nodes transform information, receiving incoming input and sending a modified output to the next layer. How each node manipulates information is extensively refined through a form of trial-and-error training. To reach the incredible ability that we know from these models, this training is vast. To illustrate just how vast, a conservative estimate is that an early model of ChatGPT (GPT-4, released in 2023) is thought to be trained on some 9.75 trillion words, around 2000 times the amount of the whole of English Wikipedia [2]. With such extensive training, these many nodes become supremely refined and come to respond in human-like fashion.

Picture 1: What AIs are made of.
Image 1. What AIs are made of

Intuitively, one could be forgiven for wondering whether ‘what’s under the hood’ is not similar to human intelligence. These models speak like humans, even share their ‘thoughts’ or ‘beliefs,’ and even the statistical model behind it is based on the human brain. The only difference is the scale. Human intelligence is the product of over a hundred billion neurons and over a hundred trillion connections between them, while language models’ neural networks are only made of at most a few billion nodes. All that is missing is to keep adding more of the same.

Probably not. Even when glossing over many important technical details, many experts in human and artificial cognition have argued that there is a foundational attribute that is decisive to human cognition and which large-language models, as they currently exist, do not have [3]. Specifically, it is argued that AIs’ “thoughts” are not grounded.

Grounding cognition

To understand what it means for thoughts to be grounded and why this is important, let’s take a step back and consider what a thought even is. In research, all the processes related to thinking are called cognition. One, quite obvious, but important insight about cognition is that it must be about something. In other words, to think about something requires knowing what one is thinking about (imagine if someone told you they had invented a new number and asked you to do arithmetic with it. You would be unable to do so until you were told what magnitude this number refers to). For this reason, a considerable point of debate in cognitive science is about how information is represented in cognition, called a representation’s format.

The same information can be formatted in different ways. For example, “2” or “••” represent the same information, the concept of two in two different formats. The first is a symbolic representation: digits are a form of symbol because they share no obvious features with the thing they represent. They are arbitrary. It is because we have learned that the digit “2” describes the concept, that this symbol comes to represent it, but in a parallel world, it would be possible that the symbol “3” could represent the concept of two. On the other hand, “••” is an iconic representation. It represents two because it is literally two things. This is called iconic because the representation in some way looks like the thing it is representing. Here, it would not be possible for “•••” to represent two, because there are three objects. One example of iconic representations most people are familiar with are emojis which (for the most part) represent things by looking like the thing they represent [4].

In a way, in symbolic representations, the actual information is not inside the symbol but in the person looking at the symbol. This is shown by looking at a type of symbol we are very familiar with: words. The word “two” written as such would carry no information for a non-English-speaking Chinese person, and the Chinese symbol 双 carries no information for a non-Chinese speaking person. This is not the case for an iconic representation, where the information is inherently part of the representation itself. Speakers of any language recognize which number “••” would refer to. This will become more important later.

Picture 2: Both symbolic and iconic representations of the concept two can be found on playing cards.
Image 2. Both symbolic and iconic representations of the concept two can be found on playing cards

Cognitive scientists have long discussed what the format of human cognition is. One set of theories, especially popular in the late 1960s to 1980s, argued that human cognition has the symbolic format [5]. Despite being very impactful and widely accepted at the time, two very influential papers by philosophers John Searle [6] and Steven Harnad [7] pointed out that such symbol-based theories have a significant issue, now called the symbol grounding problem. It states that cognition could hardly consist only of symbols because these are not grounded (the same problem ChatGPT, trained only on words, encounters today).

The Chinese room and dictionary

Grounding and its relevance to human and artificial cognition is best understood by looking to two thought experiments presented by Searle and Harnad. Searle proposed to imagine a hypothetical person placed into a small room. This small room has no windows and only two small slits on opposite sides. Through one of these slits, from the outside, people pass a single page with Chinese writing. Some moments later, a response in Chinese emerges from the other slit. To the outside world, it seems the person inside the room would write and understand Chinese. Yet, this task was done without them actually knowing what they responded. They are using an if-then book: “if the original page has this letter, then respond with this letter.” Using this book, they manage to respond intelligently, without actually understanding their own responses. This portrays how intelligent responses can be produced without understanding, given the response (if-then) rules are specific enough. Harnad expanded on Searle’s Chinese room experiment to a Chinese dictionary. Imagine you are trying to navigate through a Chinese airport, but all you have as a translational aid is a Chinese/Chinese dictionary. When you attempt to translate a sign, the dictionary, despite being able to define a word for you, does so in Chinese. You can, of course, look up each of the Chinese words in the definition, but these are again defined only in Chinese. It is impossible for any of these Chinese symbols to bear any meaning for you, because at no point are they attached to an actual experience. In other words, these symbols are not grounded.

Picture 3: A dictionary, each word refers to other entries. A Chinese dictionary without translations leaves the user in a loop of Chinese definitions.
Image 3. A dictionary, each word refers to other entries. A Chinese dictionary without translations leaves the user in a loop of Chinese definitions.

The symbol grounding problem, therefore, says that a system that truly understands things cannot rely purely on symbolic representations. Our thoughts certainly do have meaning attached to them, and we do not feel lost in a Chinese dictionary. This issue needed to be addressed, and for theories of human cognition, ‘embodied’ theories came to the rescue. These theories, such as the Perceptual Symbol Systems theory by Larry Barsalou [8], propose iconic mental representations. Thinking about an apple, for example, consists of the reactivation of perceptions that we experienced when we interacted with apples in the past. It is the visual image, the sweet taste, its heaviness when holding it in the hand, and so on. Of course, we do not see and taste an apple every time we think of one. These are covert (i.e., not consciously aware) reactivations of these sensory signals [9]. Such theories are called embodied because they put the body center stage, arguing that thoughts have meaning because they consist of reactivations of bodily sensations. Of course, that does not imply that language is unimportant for cognition. It certainly is. It is just that our thoughts are not made of symbols. Language and other symbols refer to concepts that are represented in actual experiences [10].

With these thought experiments and cognitive science’s embodied solution in hand, we can return to looking at large language models: How could ChatGPT surmount the symbol grounding problem? It is obvious that ChatGPT will never have seen or bitten into an apple, as we have. It cannot relate the word apple to what the taste of an apple actually is. Just like translating Chinese signs with a Chinese/Chinese dictionary, each word used to train the model is in turn made up of other words, and the model never makes actual contact with the world. In short, its symbols are not grounded, and therefore its “thoughts” do not actually have information in them. Just like a man locked in a Chinese room with an if-then book, ChatGPT is locked in silicon with an exquisitely refined neural network [11].

The conclusion one may draw is that ChatGPT’s very intelligent answers are not much more than facades. Facades that merely seem smart to us because we, in our own brains, imbue the symbols with meaning. So it is impossible that large language models could ever approach human-like understanding because all they have are symbols in the form of words. Right? Currently, this is most likely the case, but some emerging findings suggest this may not demonstrate the full capability.

Picture 4: How to represent apple: re-activating past perceptions of interactions with apples.
Image 4. How to represent apple: re-activating past perceptions of interactions with apples

The world hidden in language

We may have underestimated the amount of information hidden in the patterns of language  [12]. If the way words co-occur and the structure in their co-occurrence is enough to reverse-engineer a representation of the world [13], it could be argued that large language models acquire something akin to ‘grounding’, based solely on language [14]. To illustrate, assessing the degree to which a large language model can reverse-engineer a representation of space, computer scientist Wes Gurnee and physicist Max Tegmark [15] extracted a model’s representation (in the form of numerical weights) of various locations around the world. When organizing these locations on a two-dimensional plane, they form a decently accurate map of the world. They found the same for history, where events and figures were similarly accurately represented. Indeed, they even found specific nodes that encode spatial or temporal location.

Picture 5: The world map in a large-language model. Colored dots are extractions of place representations in the model overlaid over a map over the world.
Image 5. The world map in a large-language model. Colored dots are extractions of place representations in the model overlaid over a map of the world.

Astonishingly, these patterns in language may even more closely reflect the world as humans see it than the actual physical world itself. To understand this better, a group of researchers from Pavia and Berlin extracted the relations and co-occurrences of body-part words to create a ‘body map.’ This map of the body as represented in language was then used in experiments where participants were tasked with locating parts of the body in relation to others. The language-derived body map predicted participants’ responses better than the factually accurate actual body [16]. Language therefore carries in its structure not only a model of the world but perhaps even one that reflects the world as seen through human eyes.

Does that mean AIs can think like humans? Not necessarily. For one, because what we consider as “thinking” does not exist as such in these models. Still, there are processes that produce sophisticated responses, which may be enough for some to consider these “thoughts.” More importantly though, the research in this field is ever-changing and large language models often show unexpected behavior, which makes the future of this research notoriously difficult to predict. What is clear is that human cognition is an extraordinary miracle. One that we have not begun to understand in the slightest. It may be tempting at times to assume that these large language models must be like human cognition because they pass the Turing test. Yet, as the Chinese room shows us, the actual knowledge and understanding we attribute is coming more from our own head than the model’s. It is perhaps too early to write them off completely though, lest one ignore the fantastic wealth of knowledge hidden in our language.

Bibliography

[1]    C. R. Jones and B. K. Bergen, “Large Language Models Pass the Turing Test,” Mar. 31, 2025, arXiv: arXiv:2503.23674. doi: 10.48550/arXiv.2503.23674.
[2]    V. Samborska, “Scaling up: how increasing inputs has made artificial intelligence more capable,” Our World in Data, Jan. 2025, Accessed: Jan. 19, 2026. [Online]. Available: https://ourworldindata.org/scaling-up-ai
[3]    G. Pezzulo, T. Parr, P. Cisek, A. Clark, and K. J. Friston, “Generating meaning: active inference and the scope and limits of passive AI,” Trends in Cognitive Sciences, vol. 28, no. 2, pp. 97–112, Feb. 2024, doi: 10.1016/j.tics.2023.10.002.
[4]    L. R. A. Wilde, “The Elephant in the Room of Emoji Research: Or, Pictoriality, to what Extent?,” in Emoticons, Kaomoji, and Emoji, Routledge, 2019.
[5]    J. A. Fodor, The language of thought, Digit. repr. in The language and thought series. Cambridge, Mass: Harvard Univ. Pr, 1975.
[6]    J. R. Searle, “Minds, brains, and programs,” Behavioral and Brain Sciences, vol. 3, no. 3, pp. 417–424, Sep. 1980, doi: 10.1017/S0140525X00005756.
[7]    S. Harnad, “The Symbol Grounding Problem,” Physica D: Nonlinear Phenomena, vol. 42, no. 1–3, pp. 335–346, Jun. 1990, doi: 10.1016/0167-2789(90)90087-6.
[8]    L. W. Barsalou, “Perceptual Symbol Systems,” The Behavioral and brain sciences, vol. 22, no. 4, pp. 577–660, 1999.
[9]    J. Friedrich, M. H. Fischer, and M. Raab, “Issues in Grounded Cognition and How to Solve Them – the Minimalist Account,” Journal of Cognition, vol. 8, no. 1, Apr. 2025, doi: 10.5334/joc.444.
[10]    A. M. Borghi, L. Barca, F. Binkofski, C. Castelfranchi, G. Pezzulo, and L. Tummolini, “Words as social tools: Language, sociality and inner grounding in abstract concepts,” Physics of Life Reviews, vol. 29, pp. 120–153, Jul. 2019, doi: 10.1016/j.plrev.2018.12.001.
[11]    E. Borg, “LLMs, Turing tests and Chinese rooms: the prospects for meaning in large language models,” Inquiry, pp. 1–31, Jan. 2025, doi: 10.1080/0020174X.2024.2446241.
[12]    J. Firth, “A Synopsis of Linguistic Theory, 1930-1955,” 1957. Accessed: Jan. 23, 2026. [Online]. Available: https://www.semanticscholar.org/paper/A-Synopsis-of-Linguistic-Theory%2…
[13]    E. Pavlick, “Symbols and grounding in large language models,” Phil. Trans. R. Soc. A., vol. 381, no. 2251, p. 20220041, Jul. 2023, doi: 10.1098/rsta.2022.0041.
[14]    M. M. Louwerse, “Symbol Interdependency in Symbolic and Embodied Cognition,” Topics in Cognitive Science, vol. 3, no. 2, pp. 273–302, Apr. 2011, doi: 10.1111/j.1756-8765.2010.01106.x.
[15]    W. Gurnee and M. Tegmark, “Language Models Represent Space and Time,” Oct. 03, 2023, arXiv: arXiv:2310.02207. doi: 10.48550/arXiv.2310.02207.
[16]    D. Gatti, F. Günther, and L. Rinaldi, “A Body Map Beyond Perceptual Experience,” J Cogn, vol. 7, no. 1, p. 22, 2024, doi: 10.5334/joc.347.

Image sources

Image 1–4: Unsplash
Image 5: Wes Gurnee, used with permission.