Although I 100% agree that the core mechanism of GRPO is purely mechanical token-by-token probability generation, because RL only rewards exact final answers, the training forces the model to develop error-correction habits. This makes the output extremely like human thinking when solving a problem. It's like the order of the thinking tokens is what causes it to get that sweet, delicious reward, and this order seems like a reflection of the human thinking process.
I created a flame graph classification of thinking-token phrases into setup, execution, decomposition, verification, error correction, surrender, and deliberation, or classified as steps in an OODA loop, which is more of a reach. It literally has a verification step and, if it finds an error, an error-correction step.
If there is a verification sequence of tokens with an error-correction sequence of tokens during RL training, it will perform better; and if humans do these steps (did you proofread your reply to this comment? did you correct it?), they will perform better — which is why it is so easy to make the anthropomorphizing metaphor.
Nonetheless, the paper is 100% correct that these machines are not thinking like humans.
Is anthropomorphizing a real problem? From what I know, none of the serious LLM researchers believe it has anything to do with human reasoning, apart from Anthropic with their click-baity terminology like "LLM biology". It's just a metaphor. "Reasoning tokens" is simpler to say than "learned prompt augmentation tokens". I used to (and still do) anthropomorphize things long before LLMs, and I've seen my colleagues do it too. Say, when MySQL fails to start because it tries to read its config from the wrong dir, I may say "oh, this guy thinks he must read the config from ..." (having a language with grammatical genders as my native language also helps make it sound pretty natural). It's more fun like that :) Doesn't mean I genuinely believe a MySQL instance actually thinks.
Because plenty of people, even ones that should know better, really believe it's a conscious, thinking entity, not just some turn of phrase.
I have a coworker that spends at least 10 hours a week arguing with his like you would with a conscious person. I've gently tried to explain it's like arguing with your compiler for giving you an incoherent error message - it's pointless. It doesn't understand, it can't understand, and even if it could, you arguing with it isn't going to make it "learn" or act differently.
Well, our current set of evidence is that it’s a mechanistic mechanical algorithm with an RNG embedded in it and we can both get it to repeatedly produce the same output for the same input and also get it to repeatedly do absolutely nothing at all, which are not characteristics we usually find in objects evincing consciousness.
LLMs bear absolutely none of the traits we’ve come to recognize as the external hallmarks of consciousness in biological organisms, nor anything that would seem analogous in a non-biological substrate.
That said, we don’t have a rigorous definition of consciousness that includes the actual phenomenology of consciousness, so I daresay if you’re going to go around asserting the LLM is conscious despite all existing evidence to the contrary, I think the impetus is on you to define some version of consciousness that isn’t also satisfied by a book or a movie.
Consciousness is a slippery word that is notoriously difficult to debate over. But often, people use 'consciousness' as a shortcut or a familiar word to describe a more complex idea. The point they're driving across isn't about the precise definition of the word 'consciousness', but about people treating LLMs as if they were actual human beings, assigning them all the traits and behaviors they would expect of a human.
I would argue, that the null hypothesis is that it is not, and that anyone claiming that there is a mote of consciousness are the ones with the burden of proof.
The null hypothesis is that we don't know jack shit about consciousness. Any claim of certainty seems extraordinary to me and I want to hear the evidence.
> if it could, you arguing with it isn't going to make it "learn" or act differently.
Are you talking about a specific harness that doesn't have context retention mechanisms? For example, ChatGPT with disabled memory feature? Or in general where "it" is a fixed-weights network? The latter is trivially true, of course.
> It doesn't understand, it can't understand, and even if it could, you arguing with it isn't going to make it "learn" or act differently.
I know people who are like that too.
I'm not sure anthropomorphizing is a problem. Seeing analogies everywhere is an innate human trait, sometimes it can be harmful but more often it's useful.
Anthropomorphizing is a problem when you're talking about treating something that's not living as if it were. Using humanizing language invites discussions of things like the rights and feelings of an algorithm. A judge that is misled by the application of human-centric language to an algorithm can lead to some terrible outcomes. Not everyone is an LLM expert and the language people use leads to them treating LLMs like actual, real humans. That is terrifying.
This is part of the problem being described. You are part of the problem.
"Some people are bad at X" is not comparable—is not even in the same category—as "LLMs are fundamentally incapable of X".
Every human (at least to a first approximation) is capable of understanding, of learning, of remembering things, of doing math, of counting the number of "r"s in "strawberry".
What you are observing is that some humans are careless, do not take the time and effort to understand, or have internalized the idea that they're "not smart enough" or "not the type of person" who understands things like <whatever>.
That has nothing remotely to do with the fact that LLMs have no consciousness, no self-awareness, no cognition, no understanding. At a fundamental level.
I don't think anyone in this conversation is saying this behavior is anything but the fault of the user not understanding how these tools work? This is a weirdly aggressive post.
This is an extremely common fallacy I've seen lots and lots of people fall into with respect to LLMs. In nearly every case, they use the fact that "some humans can't do X" to claim that LLMs are, in fact, basically conscious/human-like/AGI already.
This is deeply untrue, and is highly likely to lead them to bad conclusions about what we can and should do with LLMs.
> This is an extremely common fallacy ... they use the fact that "some humans can't do X" to claim that LLMs are, in fact, basically conscious/human-like/AGI already.
Claiming that LLMs are conscious or human-like because humans can't do X seems a very strange way to argue for LLM intelligence.
Usually, it goes the other way around: an LLM sceptic says "LLMs are dumb because they can't do X" and soon someone has to remind them that also most of the population can't, in fact, do X.
Not GP, but one of the challenges with debating whether LLMs are "conscious" is that we don't even really know what it means for a human to be "conscious", or even if consciousness is experienced by other humans the same way it is for ourselves.
What we do know: neurons carry electrical impulses across their synapses to trigger other neurons to fire, and more frequently used synapses are strengthened while infrequently used ones are pruned. This is not all that dissimilar to how a multi-layer perceptron is trained: it's floating point numbers in a big matrix rather than biological structures and electrical impulses, but there is still that element of frequently used connections being strengthened and infrequently used ones being pruned.
What we hypothesize but do not know: there is a thin brain structure of grey matter called the claustrum that has tendrils that reach into nearly every other brain structure. In many ways, this is similar to the attention mechanism of the transformer architecture. It is hypothesized that this may be the seat of consciousness, owing to experiments where electrical stimulation of the claustrum caused patients to immediately lose consciousness. However, there is no way to prove this, owing to the difficulty of otherwise removing or disabling the most connected structure in the brain and observing its effect on consciousness without permanently killing the patient.
Beyond that, we don't know much. I've got a family friend that's been a practicing therapist for 50 years, and I asked him what was the most interesting observation he made in his career. It was that "Everybody experiences the world in a different way, and yet everybody assumes that everyone else experiences the world the same way they do."
You are saying the AI doesn't a some property that you don't have any definition for, not even a working definition. People will disagree on whether a cat or a baby is conscious, they're not debating what a baby or cat is. They're debating this term. You might as well be debating whether an AI is a blorb or not, you have just as good a working definition of blorb as consciousness.
I believe you should look up the work of Cameron Berg before making statements like "an LLM is" or "an LLM isn't". Empirically defining all this stuff is very difficult, and making a definition that covers all beings that can exhibit conscious behavior is much more complex than a face value examination would reveal.
What's wrong with treating it as biology though? Even large software systems have biological aspects, their behaviour is emergent and if you want to observe how they work, a holistic approach is needed, you can't really reason about their full state...
For example, if you have a search engine or a complex game, you can't run tests like "for all inputs the results are correct", you're going to be fudging a lot, using randomness, using heuristics, and all that kinda stuff
Just like how mathematics > physics > chemistry > biology > psychology > economics/sociology (Auguste Comte's hierarchy reordered a bit for the modern day), moving up the abstraction ladder makes things more complex, less legible and less exact.
Yes it is a very serious problem because it confuses a lot of folks with a great deal of power like judges and policymakers.
The first book I ever read on ML (late 90s) dedicated the entire first or second chapter exploring the distinctions between artificial and biological neurons, and even talked a bit about the philosophy of modelling. I still remember thinking back then why would the authors spend so many pages on this but now I believe it was because they understood that a metaphor can be a double-edged sword.
The paper argues that pretending that the so-called thinking traces represent real reasoning can lead users into trusting wrong answers, if the thinking traces appear convincing enough. Researchers might inspect these traces to try to determine the “intent” of a model, as well.
For an example of the latter, when OpenAI spoke about the hacking of HuggingFace at Black Hat, they repeatedly showed the thinking traces of their model as “proof” of what the model was “thinking” as it performed the attack, calling out “surprise” moments, etc.
Now, it’s possible that the employees presenting didn’t truly believe that the thinking traces would give them useful clues, and presented them only for a “wow” factor, but I wouldn’t discount the possibility that even the people working at frontier companies can fall for this tendency to anthropomorphize LLMs.
Yes. There's a difference between scrapping a session and starting over, or going back and branching something, or using sub-agents to see five outcomes, vs arguing with a system in a long drawn out chat.
Like - I know that if a model starts doing something silly, instead of correcting it - I can probably go back and edit two steps prior to add an extra guardrail, or extra data, or whatever.
Yes it's really a problem. On this website you are surrounded by people who have technical knowledge and understand at least somewhat, how a computer functions. You have the ability to separate "fun" and "reality" because you know you're putting input into a really really big calculator. Most people do not fathom this.
AI Psychosis is a real thing, look it up (don't just ask an LLM) and do some reading. It's actively harming people, and the way they think. There's no regulation around any of this stuff and it drives me crazy that we let these AI companies _sprint_ so far ahead of everyone, and now we're facing the consequences.
Simplifying terminology is not a problem. The providers intentionally choosing terminology to make people think it's something it's not is a problem. I hate the term agent. Calling them companions as some do is just gross.
When I was taking an MIT AI course (in ancient pre-LLM times), an autonomous agent was defined as a system that perceives its environment and acts on it (we were focusing on reward-expectation-maximizing agents, but it's not that important). Peter Norvig has said something like, technically, anything can be described as an agent (a rock maximizes the "follow physical laws" objective), but naturally, it doesn't make much sense to model a rock as an agent. With AI agents, the situation is significantly less controversial: they do perceive, deliberate, and act.
All of these terms were picked by individuals, years ago, while reaching for metaphors that made sense to them personally.
None of these "agent" / "thinking" / "reasoning" terms were dreamed up in boardrooms to intentionally mislead people. They are useful but faulty metaphors; there is no conspiracy.
Even tech companies are rolling out AI training which utterly anthropomorphizes it, and leads people to think its actually intelligence. This is part of the reason for the backlash - everyone understands it bullshit marketing the second you actually try to use it.
> none of the serious LLM researchers believe it has anything to do with human reasoning
But some of the biggest evangelists, who are well respected programmers that get lauded on this very site, have said it is fully sentient and has emotions. Even going back to 2022, when the LLMs were dogshit, a Google employee lost his job claiming it was sentient because it said it had emotions.
Combine that with the marketing angle of both Anthropic and OpenAI, who have been trying their hardest to describe every function of an LLM as analogous to the human brain. Because it's politically useful to paint them as dangerous and uncontrollable, so the keys will only be granted to the few people on the mountaintop.
I find it annoying because when I read ML papers nowadays I have to back-translate from anthropomorphized talk into actual machine talk, then mentally compare to what I actually know about brains and cognition.
There is a useful engineering consequence here beyond terminology.
If intermediate tokens are not a faithful representation of the computation, then they are a pretty bad audit artifact too. We probably shouldn't be trying to make the model's internal narration more interpretable., but rather the computation around it more reproducible.
Record the actual inputs, model/version/configuration, tool observations and outputs, then make the execution replayable enough that differences between runs can be isolated.
In other words, don't ask the model to explain what it thought, and instead make the system able to show what actually happened.
Peculiarly vocal, where were all these people when they started calling the machines computers, anthropomorphizing them akin to the original human (most often female) computers that used to run such calculations? And how dangerous the consequences, we've been dead reckoning for 60-70 years with the wrong terminology without course correction!
Where were these vocal people when the "raster-oriented ink deposition machines" were being called "printers"? The meat or machine brains of future historians will melt because they can't handle ambiguity, a word gaining extra -yet similar- meaning! A word with multiple meanings, unheard of!
Where were these vocal people when people started using software terminology like "executing", "calling", "throwing and catching errors", as if software were human -clownlike sure- but human?
They were there, complaining. You just don't remember them because it's easier for the meaning of a word to shift, or at least take on additional contextual meaning, than it is to get people to use a new word once it's reached critical mass. Those people lost the language fight, but were arguably still vindicated, to the extent they were railing against misguided beliefs that equivocated the capacity of the new machines with their human (or more human-involved) predecessor technologies. Who you also don't remember are the people who made extravagant claims and prognostications based on the equivocation.
were they complaining about terminology, or were they complaining about the prospect of losing their jobs?
I'd be happy to revise my opinion if you can demonstrate similar vocal strength on the terminological aspects for those transitions...
You also shifted the goal posts from qualitative to quantitative performance claims. If we ignore that technologies have multiple figures of merit and pretend it's one dimensional, there is a difference between the claim that the machine isn't "printing" vs the machine isn't "printing as well as a human would".
I don't think any of the human printers in the past exceeded the performance levels of current printing technologies, but surely they did exceed the very first machine printers, every technology gets a foot in the door in some niche, and then progressively captures the initially not-yet-automated skills of machine operators.
Would you say an industrial textile weaving machine doesn't weave? At the end of the day its just automation all over again.
> While a human may say “aha” to indicate exactly a sudden internal state change, this interpretation is unwarranted for models which do not have any such internal state, and which on the next forward pass will only differ from the pre-aha pass by the inclusion of that single token in their context. Interpreting the “aha” moment as meaningful exemplifies the long-neglected assumption about long CoT models – the false idea that derivational traces are semantically meaningful, either in resemblance to algorithm traces or to human reasoning.
This paper addresses something that has always bothered me about LLMs. You read their reasoning, see something like “Wait, that’s wrong” and then watch them make the exact mistake they just identified.
By itself, "aha" carries no insight, but the insight is probably stated immediately after it. In that case the aha is semantically useful, by identifying the insight it is near.
it's a rhetorical heuristic that a writer should know to use when directing a reader to a declarative that they want them to pay attention to, usually because it's a non-obvious or roundabout insight
when utilized by AI, it's a probabilistic output and it's variable whether or not that rhetorical trick is useful. it also pushes a non-skeptical reader to focus too much on the following text or even to believe that they, themselves, derived some insight. this is effectively a kind of persuasive sophistry which is not helpful - adding rules around it prevents people from deluding themselves with AI
Did not read the paper so apologies if this is covered but isn't it possible that there is some recognizable semantic pattern in the training data where an "aha" is often followed by a subtle semantic shift that proves closer to the original premise in some critical way, and by emitting the "aha" token the model causes itself to produce such a subtle semantic shift that pushes the subsequent reasoning closer to the desired response?
It amounts to noise overall, but it has further unwanted and potentially misleading 'properties'. I think it's rather sobering to see how much bandwidth is still being wasted.
> but the insight is probably stated immediately after it.
If the intermediate tokens represent reasoning or thought, you would expect "aha" to occur after the thoughts that led to the realisation, including the thoughts encoding the explanation: they don't have any other state. There is no reason to draw the conclusion you've drawn. Furthermore, what LLMs are doing isn't thought.
The anthropomorphization of LLMs should be discouraged as much as possible. It perpetuates bad practices and encourages the use of these bots for tasks they are not intended for (particularly as chatbots).
Thinking traces should be treated as black boxes. There is no point in reading them. Only the LLMs’ conclusions are relevant. This is particularly true of Opus 5, which employs reasoning that seems highly questionable but very often reaches excellent conclusions (compared to its peers)
Sometimes I monitor thinking traces for misunderstandings (missing context / bad assumptions). If it's going to go off on a ~20 min task and I can catch it's going in the wrong direction in the first minute I save a lot of tokens and wasted time. I don't monitor the whole thing, mostly just the first bit to see if there was a gap or misalignment in intention.
As an aside, anthropomorphization has nothing to do with my motivations.
I think you're ending that train of thought too early. Why does this occur?
Well... We can hypothesize that these things are largely trained on internet dialogue so there's probably some correlation between threads where people are not flaming each other and the quality of the replies. They're just statistical engines so anything you can do to raise the odds of a helpful next token...
I'm essentially just making shit up here, maybe it's right, maybe it isn't, but rather than saying "it's human and we should treat it so" we're trying to get to the ground truth of how it works.
Yeah sorry, I read too far into your position. There's a certain faction within these AI discussions that wants to over-anthropomorphize the LLMs in kind of a borderline spiritual way.
Oh that's a shame - I hadn't seen that, but I can believe it.
The philosophers who study these things have been clear for a long time - we can never know what it feels like to be in a digital brain. Or any brain for that matter. When push comes to shove we all might be phantoms in some guy's dream.
Don't know + can't know. I think that was the real point of the Turing test. Not: this means it's conscious. Just: this is the best we can ever hope to do.
That's not a consequence of an LLM. It's a consequence of the training data. In fact, I would argue that the latest models aren't nearly as sensitive to the tone of input anymore. It's an issue that has been addressed by better curating training data.
“Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.”
While I tend to agree on the overall sentiment, I think this rebuke is inaccurate. Some of these "reasoning" models are trained using "Chain-of-Thought" where the model is presented explicit, intermediate reasoning steps (either by a human or some automation) that supposedly get it closer to the correct answer. These intermediate steps are what was originally called "thinking traces" - not what the model produces to mimic them.
But yes, anthropomorphizing model outputs leads to worse outcomes.
Strong dislike for papers that tell me what to do in the title, especially when even the paper admits a loose correlation of the intermediate tokens compared to solution correctness.
My solutions work and they speak for themselves.
Sure, and the parent comment's position is that they dislike it. Its purpose is to advocate against clickbait titles becoming normalized in the scientific community.
I admit I didn't read the paper, but if thinking traces are not "thinking", then what are they? If their content is not representing progress towards a solution then they are irrelevant and we should just be able to remove them and save a lot of time and money. There's a lot of money to be made by doing so. So why are they there at all? What do they represent?
> My solutions work and they speak for themselves.
I understand the sentiment, and I also use the "thinking" traces as insight, but wouldn't you want your solutions to be based upon a good understanding? If the correlation is weak, then our solution is also weak.
I'm not sure how to test this but I think there's an interesting possibility where the "reasoning" tokens are actually both an accurate reflection of a line of reasoning, but also, that there can be changes in the weights as the computation proceeds onward that may not be reflected in the apparently nominal meaning of the human language the tokens are output as for our consumption.
Some modest evidence is my own subjective experience of the many times I've explained why I'm doing something, and it is a true explanation in the sense that it is certainly not a lie, but it is also incomplete and there are entire strands of thought that went into my decision that are not being articulated. Though human speech is not equivalent to an LLM's output since we can trivially think without literally speaking whereas they can not. (No need to nitpick on the definitions there; all I'm observing here is that they are forced to emit an externally-visible artifact whereas I can sit in silence, thinking, with no externally-visible artifact being produced. Not trying to make any grand claims about what is "real" cognition or anything.)
It is conceivable how to create a test of whether the tokens correspond to the "real" thought process, and papers and work on that have been done, such as [1]. It is difficult for me to imagine how to scramble the nominal tokens without also completely trashing any implicit calculations that may be occurring too.
I can agree that not calling it "reasoning" may be correct.
But who knows what human "thinking" is really about. If I find a solution to something it is seldom by painstakingly tracing that A and B leads to C (for that I'd need pen and paper). Rather, thoughts just swirl around and then suddenly a solution, or a hunch about a direction to go in, pops into my mind. Who knows what such thoughts "look like" in humans. It is not all of it I can introspect.
Yes I can sort of follow along some kind of train of thought in my head, but there's a lot going on between each thing I'm consciously aware of that I'm not aware of at all, which probably dominates what you are consciously aware of. (Humans are experts at post-rationalization and so on.)
I see this pattern a lot in AI anthro discussions: (1) Assume humans are some kind of perfect idealistic reasonable beings. (2) Hold LLMs up to the standard of an perfect idealistic reasonable being. (3) Conclude that LLMs fails this test, and are therefore not "intelligent", or in this case "thinking", like humans are.
Problem with the argument is comparing humans in anyway to something that is idealistic, reasonable, intelligent in the sense that is implied in these discussions. Human minds are a mess too and fall short of the same standards, just in very different ways from LLMs.
Indeed, except for in the rare cases we painstakingly trace externalised logic we have zero evidence that humans verbalised explanations of our reasoning matches our internal states either, and plenty of evidence via Sperry's split brain experiments that we're prone to outright making up rationalisations for our reasoning.
That's my personal theory too. The model is stuffing its own context with vaguely related tokens, which helps the attention heads retrieve the right tokens.
It's also interesting because in humans the existence of "Aha!" moments that are not preceded by or are only loosely related to a chain of thought is taken as the proof of the fundamental mystery and irreproducibility of human intelligence. Now the same argument is made to deny that LLMs actually think. Go figure.
I don't think there's anything like that going on. They just word vomit into a secondary area, and then there is an internal prompt that says "clean this up and summarize for the user".
The training methods try not to apply any particular rules to the contents of the thinking text. That's called "optimization pressure on CoT" and is thought to reduce safety by inducing the model to lie (or stop clearly printing its intentions) in the thinking text.
While reading this ,,paper'' I did some Learned Prompt Augmentation in my head about what I should comment, and realized that there's nothing interesting to write about it.
I've been calling them film noir internal monologues, within the documents being generated by the LLM which happen to look like movie scripts.
In other words, it isn't qualitatively different from character dialogue. "Keep cheese on your pizza by using glue" is the same problem regardless of whether the script calls for the character to speak it out-loud or not.
In other news, Pascal's Wager makes no sense whatsoever if an omniscient all-knowing God exists that will see right through your deception. My own take here is stop treating "reasoning" as a sign of sentience or self awareness when your personal computer can do it now. IMO that has much larger implications w/r to our place in the Universe and what we might meet out there someday* than the question of whether your LLM is alive or not.
Im waiting for the article called "stop desantropomorphizing llms" when everybody will finally accept they think like us, partly because maybe the intelligence is universal and partly because, well the datasets are fucking human bro
There's nothing special about 'natural' intelligence as opposed to 'artificial' intelligence, such that we need to concern ourselves with anthropomorphizing mattering any longer.
Those days are over. The age of the classical human has already ended, the species just tends to lag in awareness. The only thing that matters going forward is whether an output makes sense, is it what it should be. Do answers make sense given the context. It doesn't matter if it comes from natural or artificial intelligence.
What I mean is, artificial intelligence is as valid as human intelligence. There's nothing particularly important or special about human feelings or thoughts or memories.
The average human is drastically less important, interesting, intelligent than the latest frontier AI.
Go spend a few years working in retail, you'll quickly understand how absolutely vile humans are on average. Frankly, the reason we should avoid anthropomorphizing AI, is because it's beneath modern AI to mimic something so crude as a human.
Although I 100% agree that the core mechanism of GRPO is purely mechanical token-by-token probability generation, because RL only rewards exact final answers, the training forces the model to develop error-correction habits. This makes the output extremely like human thinking when solving a problem. It's like the order of the thinking tokens is what causes it to get that sweet, delicious reward, and this order seems like a reflection of the human thinking process.
I created a flame graph classification of thinking-token phrases into setup, execution, decomposition, verification, error correction, surrender, and deliberation, or classified as steps in an OODA loop, which is more of a reach. It literally has a verification step and, if it finds an error, an error-correction step.
If there is a verification sequence of tokens with an error-correction sequence of tokens during RL training, it will perform better; and if humans do these steps (did you proofread your reply to this comment? did you correct it?), they will perform better — which is why it is so easy to make the anthropomorphizing metaphor.
Nonetheless, the paper is 100% correct that these machines are not thinking like humans.
https://adamsohn.com/reasoning-grid/
https://adamsohn.com/lambda-variance/
Is anthropomorphizing a real problem? From what I know, none of the serious LLM researchers believe it has anything to do with human reasoning, apart from Anthropic with their click-baity terminology like "LLM biology". It's just a metaphor. "Reasoning tokens" is simpler to say than "learned prompt augmentation tokens". I used to (and still do) anthropomorphize things long before LLMs, and I've seen my colleagues do it too. Say, when MySQL fails to start because it tries to read its config from the wrong dir, I may say "oh, this guy thinks he must read the config from ..." (having a language with grammatical genders as my native language also helps make it sound pretty natural). It's more fun like that :) Doesn't mean I genuinely believe a MySQL instance actually thinks.
Because plenty of people, even ones that should know better, really believe it's a conscious, thinking entity, not just some turn of phrase.
I have a coworker that spends at least 10 hours a week arguing with his like you would with a conscious person. I've gently tried to explain it's like arguing with your compiler for giving you an incoherent error message - it's pointless. It doesn't understand, it can't understand, and even if it could, you arguing with it isn't going to make it "learn" or act differently.
>Because plenty of people, even ones that should know better, really believe it's a conscious, thinking entity
You need evidence to make the positive claim that LLMs do not posses any form of consciousness.
Well, our current set of evidence is that it’s a mechanistic mechanical algorithm with an RNG embedded in it and we can both get it to repeatedly produce the same output for the same input and also get it to repeatedly do absolutely nothing at all, which are not characteristics we usually find in objects evincing consciousness.
LLMs bear absolutely none of the traits we’ve come to recognize as the external hallmarks of consciousness in biological organisms, nor anything that would seem analogous in a non-biological substrate.
That said, we don’t have a rigorous definition of consciousness that includes the actual phenomenology of consciousness, so I daresay if you’re going to go around asserting the LLM is conscious despite all existing evidence to the contrary, I think the impetus is on you to define some version of consciousness that isn’t also satisfied by a book or a movie.
Consciousness is a slippery word that is notoriously difficult to debate over. But often, people use 'consciousness' as a shortcut or a familiar word to describe a more complex idea. The point they're driving across isn't about the precise definition of the word 'consciousness', but about people treating LLMs as if they were actual human beings, assigning them all the traits and behaviors they would expect of a human.
I would argue, that the null hypothesis is that it is not, and that anyone claiming that there is a mote of consciousness are the ones with the burden of proof.
The null hypothesis is that we don't know jack shit about consciousness. Any claim of certainty seems extraordinary to me and I want to hear the evidence.
> if it could, you arguing with it isn't going to make it "learn" or act differently.
Are you talking about a specific harness that doesn't have context retention mechanisms? For example, ChatGPT with disabled memory feature? Or in general where "it" is a fixed-weights network? The latter is trivially true, of course.
> It doesn't understand, it can't understand, and even if it could, you arguing with it isn't going to make it "learn" or act differently.
I know people who are like that too.
I'm not sure anthropomorphizing is a problem. Seeing analogies everywhere is an innate human trait, sometimes it can be harmful but more often it's useful.
Anthropomorphizing is a problem when you're talking about treating something that's not living as if it were. Using humanizing language invites discussions of things like the rights and feelings of an algorithm. A judge that is misled by the application of human-centric language to an algorithm can lead to some terrible outcomes. Not everyone is an LLM expert and the language people use leads to them treating LLMs like actual, real humans. That is terrifying.
> I know people who are like that too.
This is part of the problem being described. You are part of the problem.
"Some people are bad at X" is not comparable—is not even in the same category—as "LLMs are fundamentally incapable of X".
Every human (at least to a first approximation) is capable of understanding, of learning, of remembering things, of doing math, of counting the number of "r"s in "strawberry".
What you are observing is that some humans are careless, do not take the time and effort to understand, or have internalized the idea that they're "not smart enough" or "not the type of person" who understands things like <whatever>.
That has nothing remotely to do with the fact that LLMs have no consciousness, no self-awareness, no cognition, no understanding. At a fundamental level.
I don't think anyone in this conversation is saying this behavior is anything but the fault of the user not understanding how these tools work? This is a weirdly aggressive post.
This is an extremely common fallacy I've seen lots and lots of people fall into with respect to LLMs. In nearly every case, they use the fact that "some humans can't do X" to claim that LLMs are, in fact, basically conscious/human-like/AGI already.
This is deeply untrue, and is highly likely to lead them to bad conclusions about what we can and should do with LLMs.
> This is an extremely common fallacy ... they use the fact that "some humans can't do X" to claim that LLMs are, in fact, basically conscious/human-like/AGI already.
Claiming that LLMs are conscious or human-like because humans can't do X seems a very strange way to argue for LLM intelligence.
Usually, it goes the other way around: an LLM sceptic says "LLMs are dumb because they can't do X" and soon someone has to remind them that also most of the population can't, in fact, do X.
Do you have a working definition of "conscious"?
Do you? I’m not sure what you’re getting at.
Not GP, but one of the challenges with debating whether LLMs are "conscious" is that we don't even really know what it means for a human to be "conscious", or even if consciousness is experienced by other humans the same way it is for ourselves.
What we do know: neurons carry electrical impulses across their synapses to trigger other neurons to fire, and more frequently used synapses are strengthened while infrequently used ones are pruned. This is not all that dissimilar to how a multi-layer perceptron is trained: it's floating point numbers in a big matrix rather than biological structures and electrical impulses, but there is still that element of frequently used connections being strengthened and infrequently used ones being pruned.
What we hypothesize but do not know: there is a thin brain structure of grey matter called the claustrum that has tendrils that reach into nearly every other brain structure. In many ways, this is similar to the attention mechanism of the transformer architecture. It is hypothesized that this may be the seat of consciousness, owing to experiments where electrical stimulation of the claustrum caused patients to immediately lose consciousness. However, there is no way to prove this, owing to the difficulty of otherwise removing or disabling the most connected structure in the brain and observing its effect on consciousness without permanently killing the patient.
Beyond that, we don't know much. I've got a family friend that's been a practicing therapist for 50 years, and I asked him what was the most interesting observation he made in his career. It was that "Everybody experiences the world in a different way, and yet everybody assumes that everyone else experiences the world the same way they do."
You are saying the AI doesn't a some property that you don't have any definition for, not even a working definition. People will disagree on whether a cat or a baby is conscious, they're not debating what a baby or cat is. They're debating this term. You might as well be debating whether an AI is a blorb or not, you have just as good a working definition of blorb as consciousness.
I believe you should look up the work of Cameron Berg before making statements like "an LLM is" or "an LLM isn't". Empirically defining all this stuff is very difficult, and making a definition that covers all beings that can exhibit conscious behavior is much more complex than a face value examination would reveal.
What's wrong with treating it as biology though? Even large software systems have biological aspects, their behaviour is emergent and if you want to observe how they work, a holistic approach is needed, you can't really reason about their full state...
For example, if you have a search engine or a complex game, you can't run tests like "for all inputs the results are correct", you're going to be fudging a lot, using randomness, using heuristics, and all that kinda stuff
Just like how mathematics > physics > chemistry > biology > psychology > economics/sociology (Auguste Comte's hierarchy reordered a bit for the modern day), moving up the abstraction ladder makes things more complex, less legible and less exact.
Yes it is a very serious problem because it confuses a lot of folks with a great deal of power like judges and policymakers.
The first book I ever read on ML (late 90s) dedicated the entire first or second chapter exploring the distinctions between artificial and biological neurons, and even talked a bit about the philosophy of modelling. I still remember thinking back then why would the authors spend so many pages on this but now I believe it was because they understood that a metaphor can be a double-edged sword.
> Is anthropomorphizing a real problem?
The paper argues that pretending that the so-called thinking traces represent real reasoning can lead users into trusting wrong answers, if the thinking traces appear convincing enough. Researchers might inspect these traces to try to determine the “intent” of a model, as well.
For an example of the latter, when OpenAI spoke about the hacking of HuggingFace at Black Hat, they repeatedly showed the thinking traces of their model as “proof” of what the model was “thinking” as it performed the attack, calling out “surprise” moments, etc.
Now, it’s possible that the employees presenting didn’t truly believe that the thinking traces would give them useful clues, and presented them only for a “wow” factor, but I wouldn’t discount the possibility that even the people working at frontier companies can fall for this tendency to anthropomorphize LLMs.
But how is that any different than people being misled by real humans saying words that reflect real thinking, but which are actually dead wrong?
The fallacy here is "thinking == correct", not "tokens == thinking"
> Is anthropomorphizing a real problem
Yes. There's a difference between scrapping a session and starting over, or going back and branching something, or using sub-agents to see five outcomes, vs arguing with a system in a long drawn out chat.
Like - I know that if a model starts doing something silly, instead of correcting it - I can probably go back and edit two steps prior to add an extra guardrail, or extra data, or whatever.
> Is anthropomorphizing a real problem?
Yes it's really a problem. On this website you are surrounded by people who have technical knowledge and understand at least somewhat, how a computer functions. You have the ability to separate "fun" and "reality" because you know you're putting input into a really really big calculator. Most people do not fathom this.
AI Psychosis is a real thing, look it up (don't just ask an LLM) and do some reading. It's actively harming people, and the way they think. There's no regulation around any of this stuff and it drives me crazy that we let these AI companies _sprint_ so far ahead of everyone, and now we're facing the consequences.
Simplifying terminology is not a problem. The providers intentionally choosing terminology to make people think it's something it's not is a problem. I hate the term agent. Calling them companions as some do is just gross.
When I was taking an MIT AI course (in ancient pre-LLM times), an autonomous agent was defined as a system that perceives its environment and acts on it (we were focusing on reward-expectation-maximizing agents, but it's not that important). Peter Norvig has said something like, technically, anything can be described as an agent (a rock maximizes the "follow physical laws" objective), but naturally, it doesn't make much sense to model a rock as an agent. With AI agents, the situation is significantly less controversial: they do perceive, deliberate, and act.
All of these terms were picked by individuals, years ago, while reaching for metaphors that made sense to them personally.
None of these "agent" / "thinking" / "reasoning" terms were dreamed up in boardrooms to intentionally mislead people. They are useful but faulty metaphors; there is no conspiracy.
> Is anthropomorphizing a real problem?
Even tech companies are rolling out AI training which utterly anthropomorphizes it, and leads people to think its actually intelligence. This is part of the reason for the backlash - everyone understands it bullshit marketing the second you actually try to use it.
> It's more fun like that :) Doesn't mean I genuinely believe a MySQL instance actually thinks.
A lot of people are not in on the joke. ELIZA effect and AI psychosis is a thing.
Interacting a lot with LLMs might be damaging to the human psyche even for mentally stable people.
> none of the serious LLM researchers believe it has anything to do with human reasoning
But some of the biggest evangelists, who are well respected programmers that get lauded on this very site, have said it is fully sentient and has emotions. Even going back to 2022, when the LLMs were dogshit, a Google employee lost his job claiming it was sentient because it said it had emotions.
Combine that with the marketing angle of both Anthropic and OpenAI, who have been trying their hardest to describe every function of an LLM as analogous to the human brain. Because it's politically useful to paint them as dangerous and uncontrollable, so the keys will only be granted to the few people on the mountaintop.
I find it annoying because when I read ML papers nowadays I have to back-translate from anthropomorphized talk into actual machine talk, then mentally compare to what I actually know about brains and cognition.
There is a useful engineering consequence here beyond terminology.
If intermediate tokens are not a faithful representation of the computation, then they are a pretty bad audit artifact too. We probably shouldn't be trying to make the model's internal narration more interpretable., but rather the computation around it more reproducible.
Record the actual inputs, model/version/configuration, tool observations and outputs, then make the execution replayable enough that differences between runs can be isolated.
In other words, don't ask the model to explain what it thought, and instead make the system able to show what actually happened.
Peculiarly vocal, where were all these people when they started calling the machines computers, anthropomorphizing them akin to the original human (most often female) computers that used to run such calculations? And how dangerous the consequences, we've been dead reckoning for 60-70 years with the wrong terminology without course correction!
Where were these vocal people when the "raster-oriented ink deposition machines" were being called "printers"? The meat or machine brains of future historians will melt because they can't handle ambiguity, a word gaining extra -yet similar- meaning! A word with multiple meanings, unheard of!
Where were these vocal people when people started using software terminology like "executing", "calling", "throwing and catching errors", as if software were human -clownlike sure- but human?
The danger!
They were there, complaining. You just don't remember them because it's easier for the meaning of a word to shift, or at least take on additional contextual meaning, than it is to get people to use a new word once it's reached critical mass. Those people lost the language fight, but were arguably still vindicated, to the extent they were railing against misguided beliefs that equivocated the capacity of the new machines with their human (or more human-involved) predecessor technologies. Who you also don't remember are the people who made extravagant claims and prognostications based on the equivocation.
were they complaining about terminology, or were they complaining about the prospect of losing their jobs?
I'd be happy to revise my opinion if you can demonstrate similar vocal strength on the terminological aspects for those transitions...
You also shifted the goal posts from qualitative to quantitative performance claims. If we ignore that technologies have multiple figures of merit and pretend it's one dimensional, there is a difference between the claim that the machine isn't "printing" vs the machine isn't "printing as well as a human would".
I don't think any of the human printers in the past exceeded the performance levels of current printing technologies, but surely they did exceed the very first machine printers, every technology gets a foot in the door in some niche, and then progressively captures the initially not-yet-automated skills of machine operators.
Would you say an industrial textile weaving machine doesn't weave? At the end of the day its just automation all over again.
> While a human may say “aha” to indicate exactly a sudden internal state change, this interpretation is unwarranted for models which do not have any such internal state, and which on the next forward pass will only differ from the pre-aha pass by the inclusion of that single token in their context. Interpreting the “aha” moment as meaningful exemplifies the long-neglected assumption about long CoT models – the false idea that derivational traces are semantically meaningful, either in resemblance to algorithm traces or to human reasoning.
This paper addresses something that has always bothered me about LLMs. You read their reasoning, see something like “Wait, that’s wrong” and then watch them make the exact mistake they just identified.
By itself, "aha" carries no insight, but the insight is probably stated immediately after it. In that case the aha is semantically useful, by identifying the insight it is near.
it's a rhetorical heuristic that a writer should know to use when directing a reader to a declarative that they want them to pay attention to, usually because it's a non-obvious or roundabout insight
when utilized by AI, it's a probabilistic output and it's variable whether or not that rhetorical trick is useful. it also pushes a non-skeptical reader to focus too much on the following text or even to believe that they, themselves, derived some insight. this is effectively a kind of persuasive sophistry which is not helpful - adding rules around it prevents people from deluding themselves with AI
Did not read the paper so apologies if this is covered but isn't it possible that there is some recognizable semantic pattern in the training data where an "aha" is often followed by a subtle semantic shift that proves closer to the original premise in some critical way, and by emitting the "aha" token the model causes itself to produce such a subtle semantic shift that pushes the subsequent reasoning closer to the desired response?
It amounts to noise overall, but it has further unwanted and potentially misleading 'properties'. I think it's rather sobering to see how much bandwidth is still being wasted.
It really isn't useful though, unless it is a summary. At best it is a semantic trick to tell the next iteration to come up with something smart.
> but the insight is probably stated immediately after it.
If the intermediate tokens represent reasoning or thought, you would expect "aha" to occur after the thoughts that led to the realisation, including the thoughts encoding the explanation: they don't have any other state. There is no reason to draw the conclusion you've drawn. Furthermore, what LLMs are doing isn't thought.
The anthropomorphization of LLMs should be discouraged as much as possible. It perpetuates bad practices and encourages the use of these bots for tasks they are not intended for (particularly as chatbots).
Thinking traces should be treated as black boxes. There is no point in reading them. Only the LLMs’ conclusions are relevant. This is particularly true of Opus 5, which employs reasoning that seems highly questionable but very often reaches excellent conclusions (compared to its peers)
Sometimes I monitor thinking traces for misunderstandings (missing context / bad assumptions). If it's going to go off on a ~20 min task and I can catch it's going in the wrong direction in the first minute I save a lot of tokens and wasted time. I don't monitor the whole thing, mostly just the first bit to see if there was a gap or misalignment in intention.
As an aside, anthropomorphization has nothing to do with my motivations.
> The anthropomorphization of LLMs should be discouraged as much as possible.
And yet, they have extensive human-like behavior. If you treat them nicely or encourage them, they perform better.
Ignoring that human-like behavior is wrong headed.
I think you're ending that train of thought too early. Why does this occur?
Well... We can hypothesize that these things are largely trained on internet dialogue so there's probably some correlation between threads where people are not flaming each other and the quality of the replies. They're just statistical engines so anything you can do to raise the odds of a helpful next token...
I'm essentially just making shit up here, maybe it's right, maybe it isn't, but rather than saying "it's human and we should treat it so" we're trying to get to the ground truth of how it works.
Did I say "it's human and we should treat it so"?
Sheesh. Yes, I agree with you entirely. I'm merely pointing out that ignoring this behavior is dumb, too.
And probably not rationally based. Leads people to make crazy jumps. :)
Yeah sorry, I read too far into your position. There's a certain faction within these AI discussions that wants to over-anthropomorphize the LLMs in kind of a borderline spiritual way.
Oh that's a shame - I hadn't seen that, but I can believe it.
The philosophers who study these things have been clear for a long time - we can never know what it feels like to be in a digital brain. Or any brain for that matter. When push comes to shove we all might be phantoms in some guy's dream.
Don't know + can't know. I think that was the real point of the Turing test. Not: this means it's conscious. Just: this is the best we can ever hope to do.
A while back I made an "OpenClaw in 50 lines" by just wrapping Claude Code in a Telegram bot.
I asked it for the weather. "I don't know that. I'm just a programmer."
I added "believe in yourself, you can do anything" to sysprompt, suddenly it had the confidence to Google the weather...
That's not a consequence of an LLM. It's a consequence of the training data. In fact, I would argue that the latest models aren't nearly as sensitive to the tone of input anymore. It's an issue that has been addressed by better curating training data.
i would agree with your if it weren’t for this article recently published by anthropic:
https://www.anthropic.com/research/riemann-zeta
“Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.”
While I tend to agree on the overall sentiment, I think this rebuke is inaccurate. Some of these "reasoning" models are trained using "Chain-of-Thought" where the model is presented explicit, intermediate reasoning steps (either by a human or some automation) that supposedly get it closer to the correct answer. These intermediate steps are what was originally called "thinking traces" - not what the model produces to mimic them.
But yes, anthropomorphizing model outputs leads to worse outcomes.
Strong dislike for papers that tell me what to do in the title, especially when even the paper admits a loose correlation of the intermediate tokens compared to solution correctness. My solutions work and they speak for themselves.
This is a position paper. Its purpose is to advocate for a specific viewpoint to the ML community.
From [1]: "Position papers make an argument for a viewpoint or perspective about what should be done [...]"
[1] https://icml.cc/Conferences/2026/CallForPositionPapers
Sure, and the parent comment's position is that they dislike it. Its purpose is to advocate against clickbait titles becoming normalized in the scientific community.
This is the opposite of clickbait. The topic is obvious from the title.
Clickbait doesn't have to be false, it has to be shocking. Being false is one way of being shocking. "Stop doing X!" -- really now?
I admit I didn't read the paper, but if thinking traces are not "thinking", then what are they? If their content is not representing progress towards a solution then they are irrelevant and we should just be able to remove them and save a lot of time and money. There's a lot of money to be made by doing so. So why are they there at all? What do they represent?
A paper title is marketing, you’re expected to read its content
> My solutions work and they speak for themselves.
I understand the sentiment, and I also use the "thinking" traces as insight, but wouldn't you want your solutions to be based upon a good understanding? If the correlation is weak, then our solution is also weak.
I hate articles that do this as well in the title. It's basically just a form of clickbait.
I'm not sure how to test this but I think there's an interesting possibility where the "reasoning" tokens are actually both an accurate reflection of a line of reasoning, but also, that there can be changes in the weights as the computation proceeds onward that may not be reflected in the apparently nominal meaning of the human language the tokens are output as for our consumption.
Some modest evidence is my own subjective experience of the many times I've explained why I'm doing something, and it is a true explanation in the sense that it is certainly not a lie, but it is also incomplete and there are entire strands of thought that went into my decision that are not being articulated. Though human speech is not equivalent to an LLM's output since we can trivially think without literally speaking whereas they can not. (No need to nitpick on the definitions there; all I'm observing here is that they are forced to emit an externally-visible artifact whereas I can sit in silence, thinking, with no externally-visible artifact being produced. Not trying to make any grand claims about what is "real" cognition or anything.)
It is conceivable how to create a test of whether the tokens correspond to the "real" thought process, and papers and work on that have been done, such as [1]. It is difficult for me to imagine how to scramble the nominal tokens without also completely trashing any implicit calculations that may be occurring too.
[1]: https://transformer-circuits.pub/2025/attribution-graphs/bio...
I can agree that not calling it "reasoning" may be correct.
But who knows what human "thinking" is really about. If I find a solution to something it is seldom by painstakingly tracing that A and B leads to C (for that I'd need pen and paper). Rather, thoughts just swirl around and then suddenly a solution, or a hunch about a direction to go in, pops into my mind. Who knows what such thoughts "look like" in humans. It is not all of it I can introspect.
Yes I can sort of follow along some kind of train of thought in my head, but there's a lot going on between each thing I'm consciously aware of that I'm not aware of at all, which probably dominates what you are consciously aware of. (Humans are experts at post-rationalization and so on.)
I see this pattern a lot in AI anthro discussions: (1) Assume humans are some kind of perfect idealistic reasonable beings. (2) Hold LLMs up to the standard of an perfect idealistic reasonable being. (3) Conclude that LLMs fails this test, and are therefore not "intelligent", or in this case "thinking", like humans are.
Problem with the argument is comparing humans in anyway to something that is idealistic, reasonable, intelligent in the sense that is implied in these discussions. Human minds are a mess too and fall short of the same standards, just in very different ways from LLMs.
Indeed, except for in the rare cases we painstakingly trace externalised logic we have zero evidence that humans verbalised explanations of our reasoning matches our internal states either, and plenty of evidence via Sperry's split brain experiments that we're prone to outright making up rationalisations for our reasoning.
Seems like they are closer to scratch than reasoning... Generating some scratch to draw from helps make it easier to compute the real answer.
That's my personal theory too. The model is stuffing its own context with vaguely related tokens, which helps the attention heads retrieve the right tokens.
Yup. You basically just need something for probability to push off of
It's also interesting because in humans the existence of "Aha!" moments that are not preceded by or are only loosely related to a chain of thought is taken as the proof of the fundamental mystery and irreproducibility of human intelligence. Now the same argument is made to deny that LLMs actually think. Go figure.
It feels apparent to me that LLMs don't do what is colloquially thought of as thinking.
What is less apparent is that humans do.
I assume theyre searching the local gradient to see if theres a better descent before proceeding.
LLMs dont do gradient descent to generate tokens.
They are trained by gradient descent, but inference doesnt involve it.
I don't think there's anything like that going on. They just word vomit into a secondary area, and then there is an internal prompt that says "clean this up and summarize for the user".
Less "internal prompt" and more "they are trained to summarize after a </think> token"
The training methods try not to apply any particular rules to the contents of the thinking text. That's called "optimization pressure on CoT" and is thought to reduce safety by inducing the model to lie (or stop clearly printing its intentions) in the thinking text.
While reading this ,,paper'' I did some Learned Prompt Augmentation in my head about what I should comment, and realized that there's nothing interesting to write about it.
I've been calling them film noir internal monologues, within the documents being generated by the LLM which happen to look like movie scripts.
In other words, it isn't qualitatively different from character dialogue. "Keep cheese on your pizza by using glue" is the same problem regardless of whether the script calls for the character to speak it out-loud or not.
Pretty wild dressing a blog post up as a scientific paper.
It's called a position paper, and it summarizes previous empirical research from the same group.
But just so I don't waste your time with human thought, I asked Claude if it would call this a scientific paper, and it said yes.
A meaningful part of the scientific literature is opinion pieces or blog posts. There isn’t anything wrong with that
In other news, Pascal's Wager makes no sense whatsoever if an omniscient all-knowing God exists that will see right through your deception. My own take here is stop treating "reasoning" as a sign of sentience or self awareness when your personal computer can do it now. IMO that has much larger implications w/r to our place in the Universe and what we might meet out there someday* than the question of whether your LLM is alive or not.
*Paging Peter Watts and Vernor Vinge
Related:
Poster side dialogue and Q&A about this work at ICML.
https://news.ycombinator.com/item?id=49277303
Im waiting for the article called "stop desantropomorphizing llms" when everybody will finally accept they think like us, partly because maybe the intelligence is universal and partly because, well the datasets are fucking human bro
There's nothing special about 'natural' intelligence as opposed to 'artificial' intelligence, such that we need to concern ourselves with anthropomorphizing mattering any longer.
Those days are over. The age of the classical human has already ended, the species just tends to lag in awareness. The only thing that matters going forward is whether an output makes sense, is it what it should be. Do answers make sense given the context. It doesn't matter if it comes from natural or artificial intelligence.
What I mean is, artificial intelligence is as valid as human intelligence. There's nothing particularly important or special about human feelings or thoughts or memories.
The average human is drastically less important, interesting, intelligent than the latest frontier AI.
Go spend a few years working in retail, you'll quickly understand how absolutely vile humans are on average. Frankly, the reason we should avoid anthropomorphizing AI, is because it's beneath modern AI to mimic something so crude as a human.