Words have no power to impress the mind without the exquisite horror of their reality
Edgar Allan Poe
What’s all the noise?
One defining factor of humanity is our ability to communicate in highly complex ways. Language, spoken and written, has allowed us to build our knowledge generation by generation, and advance in ways unique to our species. One such advance is the recent explosion of generative AI as an openly available tool for the masses.
The collective imagination was first captured by image generators. The ability to generate complex pictures, simply by describing what we want using simple prompts has led to a plethora of computer generated images, some abstract, and some almost indistinguishable from reality. It seems we are all frustrated artists at heart.
More recently, text generators have hit the headlines with the release of ChatGPT. It wasn’t the first, but it was certainly the most talked about. Talk of such advances has taken many forms. Excitement at the possibilities, fear for jobs, predictions of the dawning of AGI (artificial general intelligence) and even wild speculation about the emergence of self awareness in machines.
At a more mundane level, many have suggested that these large language models (LLMs), will replace traditional search engines and might even spell the end for Google’s dominance in this space.
Much (if not all) of this speculation misses one very important point. Words, in themselves, convey no meaning. None at all.
Language reflects reality
It would appear at first glance that I am already contradicting myself. Firstly, I suggest that words have allowed us to share knowledge, and then I state that those same words are meaningless. How can both of these statements be true?
The answer is simple. Words are just hooks. Tokens that represent our experiences. The reason words work as communication is because, as humans, we have shared experiences. We receive the same sensory inputs and live in the same world as each other.
At a simplified level, human languages are made up of nouns, adjectives and verbs. We didn’t create those concepts when we created language. We simply attached sounds to the things we observed and experienced. The sounds we associate with objects, linguists call nouns. The sounds we associate with the quality of those objects, they call adjectives. The sounds we use to refer to the actions we or those entities undertake are the verbs.
When we construct a sentence, we merely connect together these sounds as a direct reflection of what we observe in the world around us.
“The hungry boy eats the green apple,” or more conceptually “time flies like an arrow.”
None of those sounds (or their written equivalents) have any meaning in themselves. They are simply labels we assign to our sensory experiences. Take away the experience and the word means nothing.
It’s only words
Take that second sentence as an example. “Time flies like an arrow”.
It can be parsed in several ways, but only one makes sense to the average reader. If you ignore experience and analyse it grammatically, various meanings emerge.
Firstly, are we being instructed to time flies in the same way as we would time an arrow, and do we have a suitable stopwatch for such a purpose? Obviously nonsense – why would anyone want to do such a thing?
Secondly, perhaps there is a type of fly referred to as a “time fly” and we are being informed that these particular flies are partial to an arrow? Again, obviously nonsense. Arrows are not a suitable food for flies, and the idea that they might like arrows at an aesthetic level is even more ridiculous.
The only meaning that makes sense, is the one in which time passes quickly, much as an arrow in flight does. This resonates with us and communicates a concept because it matches an experience we have regularly when enjoying an activity or trying to meet a deadline.
This sentence conveys no new meaning to its recipient. What it does is allow two people to share an understanding of an experience. For words to work, experience has to be shared. Only then can it be used as a communication tool. A convenient shorthand for a much more rich set of sensory memories. If you’re not convinced, try to imagine explaining the colour green to someone who has never seen it, using only words. There is simply no way. The word green has no meaning – it’s just a label.
I’ve touched on this subject in a previous blog post about the need for shared experience when attempting to communicate novel ideas.
Right? Not quite
So why is this a problem for LLMs? For some use cases it isn’t. Using an LLM to paraphrase, summarise or even improve sample text is fine. There is no need for the LLM to “understand” the text – merely to be able to generate good quality textual output.
The issues arise for use cases that expect an LLM to act as a source of information. Search is just one of these cases, but probably the most talked about at present, and the one that is demonstrating this problem all too clearly.
When you do a search, you’re looking for information. Sometimes you’re just looking for a quick answer to a simple question. Questions like “what is 10 stone in kilograms?” or “how far away is the moon”. A single answer will do, and the source doesn’t matter. What does matter is that the answer is right.
Sometimes you’re looking for something more complex. Perhaps you’re researching a topic or looking into something that affects you such as a planned venture, or the very serious issue of the implications of a recent diagnosis. In this case, you probably want multiple answers and further reading via sources. In this latter case, maybe there isn’t such a thing as a “right” answer, but you definitely want the information you receive to be credible.
In these areas, LLMs are failing. It is in the nature of an LLM to stitch words together in a statistically likely (and therefore grammatically correct) way based on the context of a series of prompts, but there is no guarantee that those words will be drawn from consistent sources or that the end result will be anything relating to the truth.
Do androids dream?
So, what is the real issue here? Some suggest it’s just a matter of time. As the models mature they’ll get more accurate. Others suggest adding post-processing functionality will help to filter out the “hallucinations”. Both of these positions, however, miss the most important point.
A system that doesn’t understand, cannot possibly analyse. This is why it’s so important to understand that words, in themselves, convey no meaning.
A system trained on words alone can never understand, analyse or even hallucinate. These words are used carelessly by those reporting on AI. They assign human characteristics to what are still very simple machines, and in doing so, lead others to draw human related conclusions regarding their capabilities.
Still not sure? Perhaps the very fact that words are attached in our brains to things that do have meaning makes it hard to accept that in isolation they are meaningless? The best way, I think, to understand this is to imagine an LLM trained on text created using a made up language in which the “words” do not relate to anything in the real world. If all that text is created based on invented grammatical rules, the result is an LLM capable of producing more examples of that text, and even of regenerating elements of the existing text.
Such an LLM cannot possibly “understand” what it is being fed, or what it is producing, because there is nothing to understand. Despite this, it can still do its job. This is what LLMs are doing right now. Any meaning in the text produced is literally in the eye of the beholder.
This is why LLMs cannot be “matured” to a point where they reliably produce factually correct output. There is no algorithm or machine that can analyse text for correctness based on words alone. To do that it needs to know what the words mean. It needs to know what makes sense and what doesn’t.
They will also never achieve awareness, no matter how many scare stories suggest they will. It’s remotely possible they will become capable of awareness, but even then it won’t happen because they have nothing to become aware of. Their “brains” are literally full of nothing.
Their time will come
Many see the act of focussing on the problems as a pessimistic or counterproductive one. Personally I see it as the exact opposite. If you ignore the problems and shortcomings, or fail to engage with them, you run the risk of going round in circles. By pretending they don’t exist, you prevent yourself from finding the solutions.
Focussing on problems is what engineers do. By understanding the reasons why something can’t be done, they find a way to do it. This is true for LLMs. By understanding the limitations of language, and therefore of LLMs, the solution becomes clear. If your goal is Artificial General Intelligence, or you want a machine to understand the spoken and written word, you need to provide it with a far richer set of information. You need to provide it with all the additional context that gave rise to the language in the first place. In other words, you need to give it human sensory input.
You also need to give it enough “neurons” to process all that information and do something with it, and connect those neurons in complex ways that reproduce the different processing areas of the human brain.
Is this possible? Absolutely. We are living proof that it can be done – we’re also a demonstration of the scale needed to achieve it. If you want to know a bit about that scale, I have another blog post you might be interested in, entitled “How Smart is AI?”.
An artificial AI that can handle that level of input isn’t here yet, and as I point out in that post, it may not be for some time to come. In the meantime, there will be many productive uses for generative AI, but being a reliable source of information won’t be one of them.
The meaningless of language will continue to be the sticking point, and the unwillingness to accept that fact will lead many AI protagonists to waste a lot of time trying to see meaning where none exists.
In conclusion, getting the truth out of an LLM is the technology equivalent of trying to get blood out of a stone. There is no blood to be had.
