Introduction
In this series of blog posts, I’m going to try to explain in reasonably plain English how artificial neural networks work. To be honest, “reasonably” is doing a lot of heavy lifting in that statement of intent. Please, bear with me if you can, and I’ll do my best.
I started working on artificial intelligence (AI) in the late eighties and have dabbled with software based neural networks ever since. I’ve used them to predict outcomes, analyse text, images and music, and control machinery, physical and simulated. Throughout that time, I’ve read many research papers about AI and also about its longer lived biological counterpart. Those documents are highly informative, but they can make for arduous reading. Hopefully these posts will be more accessible, without being any less practical.
There might be some maths, but I’ll try to keep it simple. If you want complicated proofs there are papers for that and I’ll signpost them where I can (and when I remember). However, those proofs aren’t necessary either to understand or to implement your own artificial neural networks. There will be very little history, because I’m very much not a historian. This series will focus on what and how, rather than who, when or why.
From now on I’ll be referring to artificial neural networks as ANNs for brevity.
The inspirational brain
To explain ANNs I’m going to start from the biological brains on which they are very loosely based. These organs are complicated structures made up of thousands, millions and even billions of neurons. The actual size very much depends on the animal in which they sit, but even the smallest are capable of amazing things. Each of the neurons in these brains is connected to thousands of other neurons via chemical connections called synapses.
And that’s pretty much it. Neurons and synapses.
Neurons are the decision makers of the brain. They take a set of inputs, combine them together and produce a single output as a result. Synapses provide the mechanism by which these neurons are connect to one another. They dictate the strength of those connections and it is by changing that strength that the brain learns and remembers.
To those early explorers who looked into the structure of the brain, the fact that just two components, connected together in huge numbers could give rise to intelligence and self-awareness seemed too good to be true. Miraculous even. It also seemed too tempting an opportunity to ignore. Many people (me included) have, as a result, sought to replicate that amazing piece of biology.
So far, no one has come close to succeeding.
The germ of an idea
It was from these early discoveries that the idea of an artificial brain was conceived, and the artificial neuron was invented. Look anywhere for information about ANNs and you’ll be presented with something similar to these two diagrams.

When drawn like this there are clear similarities, as follows:
- The inputs to the ANN represent nerve signals or outputs from other neurons.
- The weights represent the strength of the synapses that connect these signals to the dendrites of the biological neuron. Not all inputs are treated equally and so neurons can behave differently to one another, even with the same inputs.
- The transfer function and the activation function represent the cell body (the Soma). All the inputs combine here, and if the result is large enough, the neuron “fires” and an output is produced.
- The output activation of the artificial neuron represents the Axon and its branching terminals. This is how the neuron passes its output on to other neurons.
I’ll save an explanation for why this artificial model is far from being equivalent to a biological neuron for a later post. For now, let’s just take that as read and explain what it does.
The sum of its parts
Essentially, all an artificial neuron does is take each of its inputs, multiplies it by a weight, and then adds them all together. (In jargon for mathematicians: it takes the dot product of the weight matrix and the input matrix). The result of that sum (the excitation) is then passed through a mathematical equation known as an activation function. This function is chosen such that it creates an output that is limited to a chosen range. In networks that attempt to match the biological neuron, this function limits the output to between 0 and 1. In simple terms, these outputs can be considered to represent “yes” and “no”.
The output from a neuron is then either passed on to become an input for other artificial neurons or used as an overall network output. Learning is achieved by adjusting the value of the weights that are applied to the inputs until you get the desired outcome.
In biological neurons, synapses can be inhibitory (reduces the excitation of the receiving neuron) or excitatory (increases the excitation of the receiving neuron). To reflect this, the weights in ANNs can be negative or positive. What this means in practice is that the output from one neuron can act to either encourage another neuron to fire (excitatory – positive) or discourage it from firing (inhibitory – negative).
To understand how this works in practice, let’s limit ourselves to a neuron with just two inputs, X1 and X2, and with an activation function that means its output approaches 0 for an excitation level of -1, 0.5 for an excitation of 0 and and 1 for an excitation level of 1. This is commonly referred to as a sigmoidal activation function (see graph below).
If we also start with input weights of W1 = 1.0 and W2 = -1.0, this means that whenever X1 is greater than X2 we will get an output greater than 0.5, when X2 is greater than X1 we get less than 0.5 and when they’re equal we get 0.5.
If we consider an output of 0 to represent “no”, 1 to represent “yes” and 0.5 to represent “don’t know”, this simple neuron can therefore tell us if one input is greater than the other. Changing the weights changes the combinations of inputs that will result in an output occurring, and when we connect many of these neurons together we can get more complex outcomes.
Most guides to ANNs will talk about a “bias input” to each neuron. Basically, this is an additional input that is always set to 1 and has a weight associated with it. This is used to “move” the activation function to the left and to the right, changing the point at which the output goes from a no to a yes. In practice it isn’t needed as long as the function is centred on the origin (as this one is).
A deceptively tricky problem
To demonstrate this in practice, let’s take a simple problem and try to solve it with these simple artificial neurons. For our example, we want to create an ANN that will turn a light on when just one of two inputs is present (but not both). It turns out that we could achieve this with the following network:
In this diagram, the squares are inputs and outputs, the circles are artificial neurons, the lines are connections and the numbers are the weights on each of these connections. All the connections feed forward to the next layer until the output is reached. Not surprisingly, it is therefore called a “feed-forward network”. This is the simplest type of ANN and is severely limited in what it can do, but it is still enough to solve many machine learning problems.
It’s not immediately obvious from the weights in this diagram that it will achieve the desired outcome, so I’ve created a table below that shows how it works. As you can see, the final outputs are not exactly 1, but if we assume anything close to 1 is a yes, then the light is on when one and only one of the inputs is on.
This is known as the exclusive OR (XOR) problem and is used as a machine learning test because it is simple to reproduce but surprisingly difficult to get a machine to learn the solution to.
| X1 | X2 | N1 | N2 | N3 | N4 | N5 | Light on |
| 0 | 0 | 0.5 | 0.5 | 0.5 | 1.0 | 0.0 | No |
| 0 | 1 | 1.0 | 0.0 | 0.9 | 1.0 | 0.98 | Yes |
| 1 | 0 | 0.0 | 1.0 | 0.13 | 0 | 0.9 | Yes |
| 1 | 1 | 0.5 | 0.5 | 0.5 | 1.0 | 0.0 | No |
You might wonder where those weights came from. How did I work out that they would give the right answer? Simple. I trained a neural network to solve the problem and copied down the values from that.
A promising start – almost
With these very simple building blocks, and this basic “feed-forward” layout, you can achieve some quite surprising things. Two examples from my work that use this layout with less than 100 artificial neurons are:
- Sentiment analysis
Analysing news articles to determine if they are about a specific topic, and if the overall sentiment of the article is negative or positive.
- Personalised recommendations
Training a personalised AI to Identify which pieces of music or which films a specific person will like and dislike and recommending playlists to them based on this.
These are tricky problems, and solving them with such a simple setup is very encouraging. When people first experiment with ANNs, these early successes are very encouraging. It feels like you’ve found the magic box that can do anything with very little effort.
However, as you experiment further with this type of network, you suddenly start to hit brick walls. There are problems that feed-forward networks are very bad at dealing with, and they’re problems that biological networks excel at solving.
One such problem is motion control. If you try to train a network like this to control a robotic arm, for example, you will really struggle. The answer that many people pursue is simply to add more neurons and more layers. This can yield some improvements, but it feels a lot like using a sledgehammer to crack a nut. (In my opinion, large language models are an extreme example of this phenomenon.)
The reason it feels like this is because that’s exactly what it is. This is not just an opinion. I have successfully trained an ANN with less than 100 neurons to control a robotic arm and get it to track an object. Not only is it capable of solving the problem, but it can also learn the task in a matter of seconds. Small can be beautiful and big is not necessarily clever.
You might not be surprised to hear that the answer to this problem comes from the world of the biological brain. It involves a different approach to learning and a different way of connecting the neurons.
This brings us to the subject of neural network layouts, and the tricky issue of feedback connections. Two topics I’ll be covering in Part 2 – Hidden neurons, troublesome feedback.
