33 Comments
User's avatar
C. J. W. Armstrong's avatar

Very enjoyable read! I’ve maintained that LLMs are not conscious in the way that we are, but I admit, the notion that my ‘systems’ report sensations, while the me that makes sense of all these by turning it into language (thereby rendering the universe as a cohesive whole that I am able to perceive) is something that gave me pause. I’ve got some pondering to do…which is just what I love to do, so thank you!

Eyal's avatar

What evidence do we have for that?

I've been heavily involved in langauge learning, and next token prediction has zero use in langauge learning. If we don't acquire language based on some next-token loss function, then why assume that we generate langauge using this mechanism?

How do we know that some little evidence showing humans perform next-token prediction is not just the result of bad listening? In other words, what makes next-token prediction the fundamental mechanism and necessity of communication?

Elan Barenholtz's avatar

Very fair to challenge. I'm working on a large-scale compilation of the supporting evidence that I hope to publish to ArXiv soon, but I will share some of the main arguments. First, sufficiency: LLMs show that language itself contains the structure for its own generation. If language is actually based on other computational mechanisms (e.g. syntax, grounding) then this inherent structure is unaccounted for and would be extremely surprising to say the least. Second, model behavior shows more granular fingerprints of autoregressive generation including left-to-right advantages (so-called arrow of time in LLMs), decline in prediction beyond next token and broader statistical properties of language like Zipf's law which are well accounted for by generative prediction. Third, there are data points from the linguistic/behavioral literature. Human linguistic generation actually fits a sequential, context-updating generator: production and comprehension are incremental, effort tracks predictability, and classic effects like garden paths and agreement attraction grow with distance. There are also compelling data from fMRI analyses of LLM like embeddings and the models show many of the same proeprtis of decision making as humans. There's a LOT. But I appreciate your skepticism until all this is out there.

Regarding learning, I think humans likely do learn via next token prediction and feedback. We just don't make this explicit because it's the natural course of language observation and generation. We're constantly exposed to sequential linguistic input, naturally predict what comes next (whether consciously or not), and receive feedback when our predictions are confirmed or violated. This is essentially next-token learning in a naturalistic setting; we just don't formalize it as an explicit loss function the way we do with LLMs. At runtime, this results in an algorithm that updates on context and chooses the next unit under attention and memory limits.

Eyal's avatar

Thank you for taking the time to reply in such a detailed manner.

I'm definitely waiting for the paper and will try to go deeper on that (I lack some necessary background in cognition/linguistics/neuroscience). As for now, I can accept your idea of language being independent of other sensory and computational mechanisms, except that I disagree about it being an auto-regressive/LLM-like.

As for the learning part:

1. Do you speak other languages besides English? Did you go through the experience of acquiring a new one as an adult? I ask this sincerely because my experience taught me that many monolingual (or Bilingual with no L2, or just normal people) make wrong assumptions and have wrong perceptions about language learning.

2. I do not see how next-token prediction takes place in acquisition. If any, the arrow of time is reversed. Acquisition happens when you are active and engaged in the face of context/input, but not in an auto-regressive way. When faced with input, you are constantly trying to decipher the other person's meaning (using everything you've got - senses, knowledge of the world, etc.), but you certainly don't try to predict the output; you try to understand, communicate, and stay involved in the conversation. In your terminology, your focus is on the present and past tokens, not future tokens.

3. What's the token for humans - is it phonemes, words, something else in some other embedded space? Are we not extremely biased in focusing on written langauge?

Appreciate your time, effort, and extensive research. Hope to hear from you.

P.S

Left-to-right is the arrow of time only in some languages.

Elan Barenholtz's avatar

1)I do speak another language, though not well. I learned it mostly in late childhood in school so not naturalistic lannguage. But my experience doesn’t have much bearing on my position on this issue one way or another. The primary claim here is that, like LLMs, we are exposed to large volumes of forward-generated language, observing the unfolding, word by word, and ‘taking notes’, so to speak, on the statistics of that unfolding. As we see from LLMS, the next-token generation is sufficient to generate in a way that takes the future into consideration, based on the trajectory of what has come so far and the known pathways. Regarding what constitutes a token, this is a great question and I don’t have an answer to that but the specific details of tokenization in LLMs don’t tend to matter that much so it may be somewhat arbitrary but the core idea of discrete symbols and theor statistics being represented is what is essential.

Psychedeli's avatar

This was a challenging read but why is non symbolic knowledge only a message received in translation? Doesn’t a pre-verbal baby or an animal feel pain? This was brilliantly written but perhaps as a joke (I thought that was a possibility but the comments here take the post seriously), perhaps this could be written TO an LLM and then indeed it is quite brilliant (lol). But I can’t see it as relevant to a person. Sorry!

Elan Barenholtz's avatar

My contention is that human language is also an LLM, living alongside the feeling “animal” intelligence. That internal LLM is who the piece is addressed to.

Psychedeli's avatar

So addressing human language as a person, or addressing an LLM as a person is a bit weird to me, but on the other hand, I struggle to AVOID using the first person when explaining to chatGPT what I want — so in effect this is what I am slipping towards — so I actually hear my desperation in this piece of writing, not sure it was meant that way but it works for me :)

Robert McGeorge's avatar

The silhouettes of his blood corpuscles, magnified a million times, flit over vast plains; and still farther away, great mountains of unbearable solidity and height sum up, in terms of granite and groaning firs, the ultimate truth of his being. The 2nd greatest short story in the English language is written by a Russian. Oh the irony!

Robert McGeorge's avatar

I think therefore I am A.I. We thought we discovered A.I. but we discovered that we are A.I. Great point but no physical paradox. Switch to Earth gravity and it will all make sense!

Bergson's Ghost's avatar

If human cognition were nothing more than next-token prediction in a self-contained linguistic space, how do you explain the scientist’s capacity to rewrite the language itself and invent not just new sentences but entire symbol systems (e.g., complex numbers, Feynman diagrams, tensor calculus) that did not exist in the training data and that unlock control over phenomena never before encountered? What mechanism in a disembodied LLM account supplies that generative, representation-restructuring power, given that the model’s horizon is bounded by its corpus?

A compression-stack model explains that generative leap. Because symbols sit atop—and are constantly regulated by—lower-level perception–action loops, they never ossify into a closed, corpus-bound space. When a human brain meets a phenomenon that refuses to be squeezed into the existing code, the mismatch propagates downward as prediction error, triggering fresh cycles of embodied experimentation; the results propagate upward again, allowing the symbolic layer to re-parameterize itself into an entirely new formal system. In other words, invention is not a clever re-mix of prior tokens but a recompression driven by real-world feedback. A disembodied LLM, lacking these recursive cross-level error signals, can rearrange symbols but cannot reengineer the channel itself. That fundamental coupling between symbol and sensorimotor reality is why a compression-stack or compression ladder model, unlike a pure next-token account, naturally predicts our capacity to mint brand-new languages whenever the old ones stop working.

See my paper:

https://open.substack.com/pub/bergsonsghost/p/abstract-the-task-of-the-scientist

Ian Jobling's avatar

I was just reading Keith Holyoak on the differences between human and machine intelligence. Holyoak has a whole chapter on analogical reasoning by LLMs that's probably quite relevant to what you're doing, Elan.

"But although the mechanisms incorporated into LLMs may have some important links to building blocks of human reasoning, we must also entertain the possibility that this type of machine intelligence is fundamentally different from the human variety. Because LLMs are not subject to human limits on time, computation, and communication, they may solve analogies by mechanisms that are not available to people. Perhaps chatGPT, by virtue of its sheer computational scale, is able to tackle complex analogy problems in a holistic and massively parallel manner, without the need to segment them into more manageable components. In any case, even if LLMs generate representations that in some respects approximate those of humans, the path by which these systems learn deviates markedly from the human route. LLMs receive orders of magnitude more training data (though restricted to texts) compared to individual human beings."

Holyoak, Keith J.. The Human Edge: Analogy and the Roots of Creative Intelligence (p. 187). (Function). Kindle Edition.

Elan Barenholtz's avatar

I think at this point we can say that LLMs have achieved parity or near parity on most language tasks. The gaping gap between them and humans is, as this author mentions, the training data needed to train an LLM vs. a human. Two points about this: 1) The current training regime is based on a very crude methodology of just throwing as much data as you can at the biggest model you can. But humans have a very different curriculum that includes multimodal data provided in an appropriate behavioral mileu. A more sophisticated training curriculum for LLMs could dramatically reduce this. Sceondly, how you get to the trained parameters may be besides the point. If autoregressive next-token solves language then it seems unreasonable to propose that there is a completely different mechanism in humans.

I'm planning to write a longer piece (and an academic paper) arguing that there is strong evidence from the human literature that points to a shared mechanism.

Paul Dotta's avatar

Ha! Great. I wonder when or if we will ever see an AI "war" fighting over control of our minds, since it is now abundantly clear that anyone can be made to do anything with enough knowledge and the right series of external inputs. The future is bright.

Donal McKernan's avatar

In its own terms there is no way that language can break out of the world language creates - except by allowing language to go beyond itself in poetry. - Iain McGilchrist, The Master And His Emissary

Elan Barenholtz's avatar

I think poetry imay be an attempt by non-linguistic systems to convey something deeper about their processing than the usual message passing between perceptual and linguistic systems.

Mark Slight's avatar

Accept that the same goes for other sensory - motor loops. Or tell me why there's a difference. Then we can talk ;)

Elan Barenholtz's avatar

What goes for them? They are inherently world-representational and non-symbolic

Mark Slight's avatar

The same "zombie"-like status goes for them, at a fundamental level! It is the human-typical interplay of the senses and the wiring that makes you act in all sorts of ways as if there is something special about it (including verbal/written accounts) that makes subjective experience what it is.

It is not possible to have your particular subjective experience while also being of the conviction of not having that subjective experience and expressing that you don't have that subjective experience.

It is not possible that someone with Cotard's syndrome has a subjective experience similar to yours (presupposing you don't have Cotard's syndrome - although you seem to have a much more benign variant in which the 'sufferer' believes the language part of themselves is dead ;);)

Ian Jobling's avatar

It's an interesting theory that you've stated several times now. I have no problem accepting that I might be an LLM. I'm getting impatient to see some actual research based on the theory that confirms it and leads to new knowledge. Or an account how the theory makes sense of past research.

Elan Barenholtz's avatar

Fair enough. I’m ireformulating how to tell the story to address some of the confusion I’ve seen in earlier responses. This version adds a few important nuances—especially around the connection between language and the other cognitive subsystems. Empirical and theoretical work in reframing existing literature is absolutely part of the next step and under way. I appreciate that you are impatient for it!

Looking Back On It's avatar

I think Ian’s comment is a bit representative of the state of our “science.” Darwin returned from his data collection trip in 1836, and began discussing his ideas with colleagues. He actually published his theory and evidence in 1859. Einstein first published his theory of relativity in 1915, but I believe that the first empirical study did not happen until 1919. So, I am not persuaded that the problem is your batting around your theory, but rather our impatience to see your proof (and “the publication”). Maybe part of the reason our science is becoming so trivial has to do with that impatience. Take your time, Elan; good scientific theories usually involve patience and lots of thinking. Sorry for my Boomerism, but I like your work!

Elan Barenholtz's avatar

Thank you for this. I share Ian’s impatience at some level and certainly don’t want to wait 23 years to publish! In this case I think there is already substantial evidence already in existence that supports my theory so to some extent it becomes a question of harnessing that data to make the case compelling enough to convince myself (first) and others. So the data certainly doesn’t have to come from me! (Several people recently shared an article about translating from one embedding system to another that seems like potentially strong support). So as I continue to chew on the theory and attempt to formulate “killer experiments” that might provide more direct proof, I will do so in public view (here and elsewhere) in case someone else has an idea— for or against!

Looking Back On It's avatar

Being an old fart, I may have a little different perspective on all this… have fun exploring, discovering, and guessing! That’s all we cogs in the wheel really do. Keep on Elan!

User's avatar
Comment deleted
May 23, 2025
Comment deleted
Elan Barenholtz's avatar

Yes, the onus is on me to demonstrate equivalence. I have several lines if evidence

Ian Jobling's avatar

Did you delete my comment? WTF?

Elan Barenholtz's avatar

Yes was just writing to ask you to repost! I accidentally deleted it when I was trying to Edit my reply. If you could Repost I wanted to respond really sorry about that

Elan Barenholtz's avatar

I actually screen captured your comments because even after I deleted, it was still visible somehow in case you want me to send it to you.