Machine learning is a way of building software in which the behaviour is learned from data rather than written by a programmer as explicit rules. You supply examples — and in supervised learning, the correct answers — and an algorithm adjusts itself until its predictions match.
The traditional formulation inverts: ordinary programming takes rules and data to produce answers. Machine learning takes data and answers to produce rules.
A working definition
Arthur Samuel coined the term in 1959 for a program that could play checkers better than the person who wrote it. Tom Mitchell later formalised it: a program improves its performance at a task T as measured by a measure P, with experience E.
The practical content of that definition is that you specify the objective and the evaluation, not the procedure. A spam filter that learns is one where nobody wrote a list of spam rules — someone showed it thousands of labelled emails.
The three main types
1. Supervised learning — labelled data
You provide inputs and the correct outputs. The algorithm learns the mapping between them. This covers the majority of deployed commercial machine learning.
- Regression — predicting a continuous value: house prices, demand forecasts, energy consumption
- Classification — predicting a category: spam or not spam, fraudulent or genuine, a tumour’s likely type
Common algorithms: linear and logistic regression, decision trees, random forests, gradient boosting (XGBoost and successors — still dominant on structured tabular data), support vector machines, k-nearest neighbours, and neural networks.
2. Unsupervised learning — unlabelled data
No answers provided. The algorithm finds structure that was already there.
- Clustering — k-means, DBSCAN, hierarchical. Customer segmentation, document grouping
- Dimensionality reduction — PCA, t-SNE, UMAP. Collapsing thousands of variables into a few that still carry the information; also used to visualise high-dimensional data
- Association — “customers who bought X also bought Y”
- Anomaly detection — identifying entries that do not fit, which is how most fraud detection actually works
3. Reinforcement learning — an agent learning by acting
An agent takes actions in an environment and receives rewards or penalties. It learns a policy — which action to take in which state — rather than mapping input to output. The central tension is exploration versus exploitation: try something new for potentially better information, or take the best-known option.
Applications include game playing (AlphaGo), robotics, recommendation optimisation, and resource scheduling. Unlike the other two, the feedback arrives after a sequence of decisions, not immediately after each one.
In between
Semi-supervised learning combines a small labelled set with a large unlabelled one. Self-supervised learning generates its own labels from the data — predict the next word, the masked word, the next frame — and it is the foundation of modern large language models. Their training signal comes from the text itself, without a human labelling anything.
Deep learning
Neural networks composed of many layers. Each layer transforms its input into progressively more abstract representations: edges, then shapes, then objects; characters, then words, then meaning in context. Because the representation is learned rather than hand-designed, deep learning removed a large amount of prior feature engineering work.
Architectures by data type
- Convolutional neural networks (CNNs) — images. Sliding filters detect local patterns regardless of position
- Recurrent networks (RNNs, LSTMs) — sequences. Largely superseded for language
- Transformers — introduced in 2017 with “Attention Is All You Need”. Self-attention lets every element in a sequence weigh every other, capturing long-range relationships in parallel. BERT, GPT and essentially all current large language models are transformers
How a model actually learns
- Initialise weights randomly — the model starts knowing nothing
- Forward pass — run an example through, producing a prediction
- Loss function — quantify the error. Mean squared error for regression, cross-entropy for classification
- Backpropagation — work backward through the network to determine how much each weight contributed to the error
- Update — an optimiser (gradient descent, SGD, Adam) nudges each weight in the direction that reduces loss. The learning rate controls step size: too large and training diverges, too small and it crawls
- Repeat over many epochs across the training set, until validation performance stops improving
That is the entire mechanism. Everything impressive in the output is the accumulated effect of millions of small gradient steps.
The problems that actually occur
Overfitting and underfitting
An underfit model has not learned enough — high error on training and test alike. An overfit model has memorised the training data, including its noise, and fails on anything new. It typically shows excellent training accuracy and poor validation accuracy.
Countermeasures: more data, regularisation (L1/L2, dropout), early stopping, and simpler architectures. Cross-validation — splitting data into folds and rotating which is held out — gives a far more honest estimate than a single split.
Bias in, bias out
A model trained on historically biased data reproduces that bias, often more consistently than a human would. Documented examples include hiring systems penalising resumes associated with women because of historical hiring patterns, and facial recognition systems with markedly higher error rates on darker-skinned faces. This is a data and problem-framing issue far more often than a modelling one — and it is why the label “unbiased algorithm” should be treated as a claim requiring evidence.
Misleading evaluation
Accuracy is nearly useless when classes are imbalanced. A model that always predicts “not fraud” is 99% accurate on a dataset with 1% fraud and completely worthless. Use precision (of those flagged, how many were real), recall (of the real ones, how many did you catch), their harmonic mean F1, and ROC-AUC. The right metric depends entirely on the cost of each type of error.
Other failure modes
- Distribution shift — the world changes and the model’s training data does not. Fraud patterns, consumer behaviour and disease presentation all drift
- Data leakage — information from the future or from the target sneaks into the features, producing spectacular validation scores and useless production performance
- Adversarial examples — tiny, imperceptible perturbations that flip a classification
- Opacity — deep models are difficult to interpret, which matters in lending, medicine and criminal justice
- Cost — training large models requires substantial compute and energy
Where you already use it
Spam filtering; recommendation systems; route and traffic prediction in maps; card-fraud detection; voice assistants and speech-to-text; autocorrect and predictive text; translate tools; photo search and organisation; credit and insurance scoring; medical imaging triage; predictive maintenance on machinery; ad targeting; and the large language models doing summarisation and drafting.
Its most reliable wins are problems with abundant data, a clear objective, and a cost of error that is understood — fraud detection being almost the textbook case.
What machine learning is bad at
- Causal reasoning — it learns correlations. Knowing that hospital admissions correlate with ice cream sales does not tell it what causes what
- Extrapolating far outside the data — distribution shift breaks it silently
- Working without data — the constraint is usually data availability and quality, not algorithms
- Guarantees — output is probabilistic, not certain, and there is no mechanism by which it knows it is wrong
That last point matters for generative systems specifically: they are trained to produce plausible text, which is not the same objective as producing true text. Confidently stated fabrication follows directly from the optimisation target rather than being an incidental bug.
Frequently asked questions about machine learning
Is machine learning the same as artificial intelligence?
Machine learning is a subset of AI — one approach among several, and historically the one that succeeded. Classical AI built systems on explicit logic and search; modern systems lean heavily on learning from data. When people say “AI” today they usually mean machine learning, and often specifically deep learning.
Do I need to be a mathematician to use it?
To use existing tools, no — frameworks and managed services handle the mechanics. To do it well, you need genuine understanding of statistics, probability and linear algebra, because the difficult part is never fitting the model; it is knowing whether your data, split, metric and objective are correct. Most real failures are design errors, not coding errors.
Will machine learning replace programmers?
It changes the work more than it eliminates it. Code generation handles boilerplate and well-specified fragments effectively, and raises the value of specifying problems precisely. What remains distinctly human — system design, defining the right objective, understanding the domain, testing whether the output is correct — is precisely the part where model-generated code most needs scrutiny.
Why does my model score perfectly in testing and fail in production?
Almost always one of four causes: data leakage inflating your test score, distribution shift between training and production, an evaluation metric that does not reflect the actual decision, or a mismatch between how the test set was constructed and how real inputs arrive. It is worth checking in that order — leakage is the most common and the least obvious.
Is more data always better?
Up to a point, and the point depends on the problem. Beyond it, returns flatten while cost rises linearly. Data quality and representativeness dominate quantity: a smaller clean dataset that resembles production will outperform a vastly larger mislabelled one. Incorrect labels are worse than absent ones, because they teach the model actively wrong things.
This article explains general technology concepts. For where these systems are deployed in everyday life, see our guides on AI tools for students and writing prompts with AI.















