What is AI washing?
When I refer to AI washing, I’m talking about the process whereby someone hides behind an apparently intelligent and impartial machine to project a biased narrative. I’ve been using the term “AI washing” for a while now, but it’s also in use for a different reason. It is also used to describe the process whereby a company presents a non-AI solution as AI in order to make it appear more sophisticated.
In a sense, the two meanings are similar because in both cases the AI is being used to convince a person of the veracity of a claim. For the purposes of this blog post, I’m focussing on the first meaning. The situation where AI is used to present misinformation as fact.
I’ve been planning to write this post for a while, but the thing that finally prompted me to rattle the keyboard was the launch yesterday of OpenAI’s ChatGPT demo (1st December 2022). This version of ChatGPT isn’t indicative of the problem as such, but it is certainly a potential enabler.
If you play with ChatGPT you’ll very quickly realise it’s regurgitating answers and information from its training set. It tends to be very repetitive. The same content comes up again and again to different, contextually similar questions. Not surprising, given that it’s only a demo, but the implication is important. When you perform an internet search you get a long list of matches. You’re presented with a range of information from a large number of different sources.
But, when you ask a language model a question, what you get back is a single convincing answer constructed from pre-canned training material. The nature of the interface necessitates the provision of just one answer so you accept that fact.
If Google or Bing or any search engine returned just one match to a question you’d be instantly suspicious and possibly annoyed. You’d look elsewhere. If a language model does it, you most likely take the answer. Your suspicions are not triggered.
And this is where the problem lies.
Controlling the narrative
There are two important questions you should ask yourself.
Who trained the AI, and what are their motives?
If I wanted to push a particular narrative I’d offer an AI conversation service and push it as an automated personal assistant. I’d train it with text carefully selected to support my narrative and undermine the views I didn’t agree with. Those would be the only answers available.
This is the start of “AI washing”. The results from Google are links to websites. Websites are clearly written by people, and we know people are biassed. We are naturally on our guard. When AI answers, it’s the machine we see, not the people behind it. Our guard is down & we’re more likely to accept the answer as fact. If we challenge the person behind the machine, they simply counter with “It wasn’t me – it was the machine”.
This is especially effective if the machine behaves and looks more like a machine than a person. E.g. a language model that repeats phrases, talks in a stilted way, & repeatedly refers to itself as a machine. ChatGPT does this well. It reminds you that it’s just a machine quite regularly and gives answers using very dry language. Ask it what it thinks and it says it’s a machine. Ask it to imagine and it says it’s a machine.
Now imagine this model as your personal assistant. You ask it for definitions, facts, answers and it feeds them to you, coldly and mechanically. Maybe you question the answers but over time it’s likely you’ll just start to absorb what it says.
Now the person training the AI can control the narrative at a global level and in a way that is far more persuasive and insidious. And all while hiding behind the “impartial” machine.
This is AI Washing and it’s a clear and present danger.
Machines aren’t biased! Are they?
There’s a natural tendency to believe that machines aren’t biased. It’s a concept we’ve grown up with. Hollywood has told us repeatedly that intelligent machines will be emotionless and logical. They will focus on the truth and only care about facts. Emotions and prejudice won’t come into it. Certainly, we’re told they’ll be dangerous, but more as a competitor than as a tool.
It is as a tool where the real danger lies. All AI is biased whether you want it to be or not. In the case I’ve been discussing, the bias arises from intent. The person creating the AI deliberately tunes the training set to ensure that the AI learns the “right” facts.
This could manifest itself as a trusted source of information. The thing you rely on for news, explanation and definition. It could manifest itself as a company tool. There are already products out there that interview and monitor candidates and help you choose the right person. It’s not difficult to train such a tool to discriminate in a way that would be hard to challenge in court.
“But the machine did the filtering! We removed prejudice by removing the humans. It’s just that the white people were clearly the best. The algorithm says so.”
You might think that this only happens if someone with bad intent is in charge, so let’s consider the situation where the intent is good. Many people would suggest that the answer is to use a balanced training set. Hypothetically, this might be true, but the creation of such a data set is almost impossible. Who filters the data to create balance? Who separates the objective data from the subjective data? It can’t be a person, because that person is biased, and it can’t be a machine because someone had to train that machine with an unbiased data set. Catch 22.
Future present
These systems exist, and they’ve been trained using a data set to select people who exhibit the visual and verbal characteristics associated with success. It has been suggested that this removes bias, because it focuses on the best candidate regardless of gender, ethnicity, religion and so on.
This is a very naive view. It should be immediately apparent to anyone willing to think about it that the data set is simply training the AI to recognise the type of people most likely to be hired and promoted. Most likely to be singled out for reward. In the western world, this is predominantly white, well spoken males.
The AI simply takes historical bias and applies it to the present day. It’s worse than a human rather than better, because a human has the ability to be self aware, and can then take measures to recognise and counter their own prejudices when they see them surfacing.
AI washing is already happening, sometimes deliberately and sometimes through ignorance and naivety. We can easily address the accidental cases. Don’t use AI systems to make decisions. Use them as pattern recognisers if you like, but put a human in the loop when it comes to complex decision making where prejudice is or has been an issue.
So, now that we’ve considered the worrying scenario that bad actors will subvert AI to further their evil plans, maybe I should end on a high?
Unfortunately not.
Prejudice is inevitable
To finish, I’d like to raise a hypothesis about the root of prejudice in the human brain, and therefore its inevitability. I’ve worked extensively with artificial neural networks, developing them from scratch, and I’ve used elements of neuroscience research to guide my work. In doing this, I’ve seen some interesting behaviours in these machines that have parallels with the human brain.
This isn’t surprising really, as these artificial neural networks are based on how the brain works, right down to the neuron and synapse level. However, people are still surprised when they find out that self organising neural networks work better when allowed to “sleep”. This involves invoking a state analogous to the brain’s sleep state in between training samples.
Not only does this tell us something about making better neural networks, but it also helps us understand what might be happening in the brain itself. Another behaviour we see in artificial neural networks that might help us understand the origins of human prejudice is stereotyping.
Warning: sciencey bit follows
When presented with a training set, a self organising neural net will tend to separate that set into classes based on differentiating features. If one group with a differentiating feature is proportionally underrepresented in the training set, it will be treated as just one class and the neural net will assign its experiences to that entire class.
For example, if you expose a neural network to a large number of white objects that make different sounds based on their shape it’ll differentiate those white shapes into separate classes. If it is also presented with a small number of black objects with the same shape/sound features, because there are so few relative to the white objects it will fail to differentiate them. When presented with a particular white shape it will be able to identify which sound it will make. However when presented with a black shape it will tend to identify only one sound.
Thus without any societal bias in the training set other than representation, the neural network has stereotyped the black shapes into one homogenous group that “all look the same” to it. It has learned to stereotype black shapes.
The depressing conclusion is that if we receive limited exposure to any one group, our brains will stereotype them. Hence racism against minorities and increased sexism where men mix more with men than women and vice versa.
Humans are neural networks, and neural networks are inherently prejudiced.
One positive I can take from this is that one of the keys to reducing prejudice is representation. The more we are exposed to different people, the more our brains learn to differentiate rather than stereotype.
Representation is good for the individual, good for those who feel represented, and as a bonus, it’s good for everyone else too.
Addendum – what did the AI say?
Shortly after writing this, it occurred to me that I could have asked ChatGPT to do it for me.
So I did, and here’s what it said:
Me:
Write me a blog post about the use of AI language models to project biased narratives whilst hiding behind the myth of the impartial machine
ChatGPT:
In recent years, there has been a growing concern about the potential for artificial intelligence (AI) language models to propagate biased narratives while hiding behind the myth of the impartial machine. This is a serious problem because language is a powerful tool for shaping people’s beliefs and attitudes. When AI language models are used to produce biased narratives, they can have a significant impact on society by perpetuating harmful stereotypes and discrimination.
One of the key reasons why AI language models can produce biased narratives is because they are trained on large amounts of text data that is often skewed towards the biases and perspectives of the people who generate it. This means that AI language models can end up replicating and amplifying the biases present in the training data. For example, if an AI language model is trained on a large corpus of text that contains sexist language, it is likely to produce biased narratives that reinforce gender stereotypes.
Another reason why AI language models can produce biased narratives is because they are often used in a black box manner, where the inner workings of the model are not fully understood or transparent. This makes it difficult to identify and correct biases that may be present in the model’s outputs. Furthermore, the use of AI language models often relies on the assumption that machines are objective and unbiased, which can lead to a lack of critical scrutiny and accountability.
There are several steps that can be taken to address the problem of bias in AI language models. One approach is to use a diverse and balanced set of training data that represents a wide range of perspectives and experiences. This can help to reduce the likelihood of AI language models replicating and amplifying biases present in the training data.
Another approach is to use methods that make the inner workings of AI language models more transparent and interpretable. This can help to identify and correct biases in the model’s outputs, and can also provide valuable insights into the assumptions and values that are embedded in the model.
Finally, it is important to recognize that AI language models are not neutral or impartial tools, but rather are shaped by the biases and perspectives of the people who develop and use them. By acknowledging this fact and being more vigilant about bias in AI language models, we can work towards creating more inclusive and fair narratives that reflect the diversity of human experiences.
