AI can write an email, recognise an image, or help you build software. But what actually happens underneath those abilities? What does it mean when someone says a model has been “trained”?
For this first article in TayoTechTalk’s How AI Works category, I want to begin with something you can follow using ordinary arithmetic. We will build a tiny delivery-fee example, examine the numbers inside it, and then connect those ideas to neural networks.
You do not need to know Python or advanced mathematics. You only need to be comfortable multiplying and adding a few numbers.
First, which part of AI are we explaining?
Artificial intelligence is a broad field. This article focuses on machine learning: systems whose behaviour is shaped by learning patterns from data. Neural networks are one family of machine-learning models. Not every AI system uses a neural network.
We will start with supervised learning, where training examples include the answer we want the model to predict. For a delivery-fee model, an example might contain a journey’s distance and its recorded fee. Other learning approaches exist, but this gives us a manageable starting point. [1]

A tiny example: estimating a delivery fee
Imagine a fictional delivery company with the following records:
| Distance | Recorded delivery fee |
|---|---|
| 2 km | ₦900 |
| 5 km | ₦1,500 |
| 8 km | ₦2,100 |
These numbers are deliberately made up to illustrate a clean pattern. They are not market prices, and three examples would not be enough to validate a real delivery-pricing system.
When distance rises from 2 km to 5 km, the fee rises by ₦600. That is ₦200 for each additional kilometre. The same relationship appears between 5 km and 8 km.
A formula that fits all three records is:
Estimated fee = (distance × ₦200 per km) + ₦500
For a 6 km delivery, that gives:
(6 × ₦200) + ₦500 = ₦1,700
If we write those numbers into an application ourselves, we have created a fixed pricing rule. In a machine-learning version, we choose a form of model and a training procedure; the procedure fits its adjustable numbers from examples. In this case, the form is called linear regression. [2]
What are inputs, weights, and bias?
Our example gives us a way to separate three terms that can sound intimidating.
| Term | Meaning in this example | Value |
|---|---|---|
| Input | The distance for the delivery being estimated | 6 km |
| Weight | The multiplier applied to distance | ₦200 per km |
| Bias | The offset added to the result | ₦500 |
| Output | The model’s estimated fee | ₦1,700 |
The mathematical form is y = wx + b. Here, x is the input, w is the weight, b is the bias, and y is the prediction.
The input changes when we ask about a different journey. The weight and bias are the model’s parameters: the adjustable values fitted during training. In our fictional pricing example, the bias acts like a base fee. In other models, it is simply an offset and may have no everyday interpretation. [2]

Suppose we changed the weight to ₦250 per kilometre while keeping the bias at ₦500. Our 6 km estimate would become ₦2,000. Changing the bias to ₦700 while keeping the original weight would instead give ₦1,900.
Notice the difference: increasing the weight changes how rapidly the fee grows with distance. Increasing the bias shifts every estimate upward by the same amount in this model.
Also, “bias” here means a mathematical offset. It is a different use of the word from unfair or prejudiced outcomes in AI systems.
How does a model learn those numbers?
Let us start with deliberately poor settings: a weight of ₦100 per kilometre and a bias of ₦0.
For our 5 km example, the prediction would be ₦500. The recorded answer is ₦1,500, so the prediction is ₦1,000 too low.
A loss function turns prediction errors into a numerical measure. One common choice for regression is squared error. For this example, the squared error is (500 − 1,500)² = 1,000,000. That number is an error score, measured in squared naira, not a charge owed to anyone. Across several examples, we can average their squared errors. [3]
Now imagine trying a weight of ₦150 and a bias of ₦300. For 5 km, we get ₦1,050: still too low, but closer. These are hand-picked settings to illustrate improvement, not a claim about the exact updates a training algorithm would produce.
Training automates the adjustment process. A method such as gradient descent uses gradients—measurements of how the loss changes when parameters change—to choose updates. A setting called the learning rate controls the size of those updates. Updates are based on the training objective across the selected examples, rather than simply increasing a number whenever one prediction is low. [4]

For our invented dataset, the settings ₦200 per kilometre and ₦500 fit every record exactly. Real data is usually messier: traffic, route choice, parcel size, and pricing changes can all affect a delivery fee. A perfect fit to a few tidy records says little about performance in the real world.
What is an artificial neuron?
An artificial neuron is a mathematical unit. In a common neural-network design, it multiplies its inputs by weights, adds those results and a bias, and applies an activation function.
For example, take two inputs: 2 and 3. Give them weights of 4 and −1, and add a bias of 1. The intermediate value is:
(2 × 4) + (3 × −1) + 1 = 6
If the activation is ReLU, which returns zero for negative values and leaves positive values unchanged, the output is 6. If the intermediate value had been −2, ReLU would produce 0.
Nonlinear activation functions allow networks to represent relationships that a single straight-line model cannot capture. Adding layers of purely linear calculations would still give an overall linear calculation. [5]
The name “neuron” is an analogy. This arithmetic unit is not a biological brain cell.
How do neurons become a neural network?
A neural network connects these units in layers. An input layer receives numerical features; hidden layers transform them; an output layer produces the result. Each hidden layer receives values produced by the preceding layer.
“Hidden” means these layers sit between the inputs and outputs. A deep neural network has multiple hidden layers. The picture below shows a small illustrative network, not the architecture of a particular commercial AI system. [6]

During training, backpropagation efficiently calculates gradients through the network. An optimizer then uses those gradients to update its parameters. The arithmetic grows larger, but the connection to our small example remains: predictions produce a loss, and training changes adjustable numbers. [1]
Training and using a model are different activities
Training fits parameters from data. Inference uses a model to produce an output for an input. Asking our fitted delivery model to estimate a fee for 6 km is inference.
Producing an answer does not, by itself, mean the model has changed its weights. A system can store information or include earlier messages in its input without performing another training step. [1]
How do we know it learned something useful?
We need to evaluate the model on examples it did not train on. The ability to perform well beyond the training examples is called generalization. A model can fit its training data closely and still fail on new cases; this is a central concern in overfitting. [7]
For our delivery example, imagine that all the training records describe small parcels on quiet roads. What happens with a large parcel during rush hour? Distance alone may not provide enough information. The model needs suitable inputs, representative examples, and evaluation that reflects the service we actually want to provide.
There is another limitation: our formula predicts ₦500 even at zero kilometres. That is mathematically consistent with its base fee, but it does not establish whether the fictional company should charge for a cancelled trip. Business rules and model predictions answer different questions.
Try the arithmetic yourself
Using our fictional model, estimate the fee for 10 km. Then ask what would happen if the bias increased by ₦100.
The first answer is (10 × ₦200) + ₦500 = ₦2,500. With a bias of ₦600, the answer becomes ₦2,600. The distance-related part stays the same.
Understanding this small calculation gives you a useful foundation for reading about much larger models. You can now ask concrete questions: What information goes in? Which numbers are learned? How is an error measured? How do we test the result on something new?
Those are the questions I will keep returning to in this series as we explore neural networks, language models, and the resources needed to run them.
Sources and further reading
The examples and illustrations in this article are original educational material. The technical explanations were checked against Google’s Machine Learning Crash Course and glossary: