Most of what runs inside a modern model is old. The artificial neuron dates from 1943, the Widrow-Hoff least-mean-squares rule from 1960, and backpropagation was popularised in 1986. What changed is the data, the hardware, and a few architectural ideas that made very large networks trainable. Knowing which parts are old and which are new is practical. It tells you which claims to doubt, why some ideas keep returning, and why progress in this field arrives in bursts separated by long disappointing stretches.
This article follows the field era by era. For each one it asks the same questions: what problem people were trying to solve, what stopped them, what removed the obstacle, and what survived into today's stack. It includes a small demo you can run, which reproduces the most famous turning point in miniature. Dates are given only where they are well established. Attributions that are still argued over are described as such.
Why an engineer should read the history
Three patterns recur often enough to plan around. First, capability usually arrives when a bottleneck in data or compute is removed, not when a cleverer algorithm appears. Second, every era produced overconfident forecasts followed by funding collapses, so a field-wide mood is weak evidence about any particular method. Third, ideas that failed once often worked later at a larger scale. Neural networks were dismissed at least twice before they took over.
Each pattern becomes a habit: ask what a technique needs in data and compute, and judge a claim by its evaluation, not by the excitement around it.
The eras at a glance
Foundations, 1943 to 1956
In 1943 Warren McCulloch and Walter Pitts described an idealised neuron: it sums weighted binary inputs and fires if the sum crosses a threshold. They showed that networks of such units can compute logical functions. Every unit in a modern network is a direct descendant of this model, with a smooth activation in place of the hard threshold.
In 1950 Alan Turing published "Computing Machinery and Intelligence". He replaced the question "can machines think?" with a behavioural test, the imitation game, and discussed the idea of a "child machine" that is taught rather than programmed. In the summer of 1956 the Dartmouth workshop, proposed the year before by John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon, gathered the founders of the field under the name "artificial intelligence". The proposal's assumption that the key problems could be cracked by a small group in a summer set a pattern of optimism that the field has repeated ever since.
Symbols and the first learners, 1956 to 1974
Early AI split into two traditions. The symbolic tradition treated intelligence as manipulating symbols with rules. Newell, Shaw and Simon's Logic Theorist proved theorems from Principia Mathematica, and their General Problem Solver searched for sequences of operators that reduce the difference between a current state and a goal. Search with heuristics, a central idea of this era, still runs in planners, compilers and game engines.
The learning tradition adjusted numbers from data. Arthur Samuel's checkers program, described in a 1959 paper, improved by playing against itself, and that paper helped popularise the phrase "machine learning". Frank Rosenblatt's perceptron (1958) learned the weights of a threshold unit from labelled examples. Widrow and Hoff's ADALINE and its least-mean-squares rule (1960) adjusted weights in proportion to the error, which is gradient descent on squared error in all but name.
In 1969 Minsky and Papert's book Perceptrons analysed precisely what single-layer perceptrons can compute. A single layer can only separate classes with a straight line, or a hyperplane in higher dimensions, so it cannot represent XOR. Multi-layer networks could, but no one had a practical way to train them. Historians still debate how much the book caused the decline of neural network research. The underlying problem, training hidden layers, was real either way.
The first winter
By the early 1970s the promises had outrun the results. In 1966 the ALPAC report concluded that machine translation had fallen far short of its goals, and US funding for it was cut. In 1973 James Lighthill's report to the UK Science Research Council argued that AI methods suffered from combinatorial explosion: techniques that worked on toy problems collapsed when the problem grew. Funding fell in both countries.
The technical lesson: search over symbolic states grows exponentially unless knowledge prunes it. The next era tried to supply that knowledge by hand.
Expert systems and the second winter
Expert systems encoded a specialist's knowledge as hundreds or thousands of if-then rules, with an inference engine to chain them. MYCIN, developed at Stanford in the 1970s, recommended antibiotics for blood infections and attached certainty factors to its conclusions. R1, later called XCON, configured VAX computer orders at Digital Equipment Corporation and became one of the first commercially successful AI systems. Through the early 1980s companies built knowledge-engineering teams, and specialised Lisp machines were sold to run the software.
The limits were operational rather than theoretical. Knowledge acquisition was slow, because experts could not articulate everything they knew. The rule bases were brittle: an input slightly outside the rules produced confident nonsense. Maintenance grew harder with every rule, because rules interacted in ways nobody could predict. When cheaper general-purpose workstations overtook Lisp machines in the late 1980s, that market collapsed, and a second winter followed. Rule systems did not disappear. They live on in business-rule engines and in the guardrails around today's model deployments.
Backpropagation and the connectionist return
Training hidden layers needs the gradient of the loss with respect to every weight. Reverse-mode automatic differentiation, the general technique, was described by Seppo Linnainmaa in 1970, and Paul Werbos proposed applying it to neural networks in his 1974 thesis. The 1986 Nature paper by Rumelhart, Hinton and Williams showed convincingly that backpropagation learns useful internal representations, and it brought the method into wide use. The same chain rule runs in every training step today; see backprop through a transformer for the modern form.
Other lasting architectures appeared in the same period. Yann LeCun and colleagues trained convolutional networks to read handwritten digits at the end of the 1980s. Weight sharing and local receptive fields, the ideas behind convolutional nets, still underpin vision models. Hochreiter and Schmidhuber's LSTM (1997) used gated memory cells so that gradients could flow across long sequences, and it dominated sequence modelling for the next two decades.
The demo below reproduces the turning point. A perceptron trained with Rosenblatt's rule cannot learn XOR however long it runs. A network with one hidden layer, trained by backpropagation, learns it in a few thousand steps.
import numpy as np
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]], dtype=float)
y = np.array([0, 1, 1, 0], dtype=float) # XOR
# 1958: one threshold unit, perceptron learning rule
w, b = np.zeros(2), 0.0
for epoch in range(100):
for xi, ti in zip(X, y):
pred = float(w @ xi + b > 0)
w += (ti - pred) * xi
b += (ti - pred)
print("perceptron:", [float(w @ xi + b > 0) for xi in X]) # never all four right
# 1986: one hidden layer trained with backpropagation
rng = np.random.default_rng(0)
W1, b1 = rng.normal(0, 1, (2, 4)), np.zeros(4)
W2, b2 = rng.normal(0, 1, (4, 1)), np.zeros(1)
sig = lambda z: 1 / (1 + np.exp(-z))
for step in range(5000):
h = sig(X @ W1 + b1) # forward pass
out = sig(h @ W2 + b2).ravel()
d_out = (out - y)[:, None] # dLoss/dlogit for cross-entropy
d_h = (d_out @ W2.T) * h * (1 - h) # chain rule through the hidden layer
W2 -= 0.5 * h.T @ d_out; b2 -= 0.5 * d_out.sum(0)
W1 -= 0.5 * X.T @ d_h; b1 -= 0.5 * d_h.sum(0)
print("mlp:", np.round(out, 3)) # close to [0, 1, 1, 0]When run, the perceptron's predictions keep cycling and never match all four targets, while the network's outputs end close to 0, 1, 1, 0. The d_out line uses the fact that sigmoid plus cross-entropy gives a gradient of prediction minus target, derived in the cross-entropy derivation.
The statistical learning era
Through the 1990s and 2000s much of the field favoured methods with clearer theory. Support vector machines (Cortes and Vapnik, 1995) found the maximum-margin separator by solving a convex problem, and kernels let them draw non-linear boundaries. Boosting, notably Freund and Schapire's AdaBoost, combined weak learners into strong ones. Breiman's random forests (2001) averaged many decorrelated trees. Hidden Markov models drove speech recognition, and statistical machine translation learned from aligned bilingual text instead of hand-written grammars.
This era left the field its working discipline: train, validation and test splits, cross-validation, regularisation, and the idea of generalisation as something to measure rather than assume. Shared benchmarks such as MNIST and the UCI repository made results comparable. Gradient-boosted trees remain the strongest default for tabular data. In 1997 IBM's Deep Blue beat Garry Kasparov at chess, mostly through specialised search hardware and hand-tuned evaluation rather than learning. It was a reminder that search and compute can carry a long way.
The deep learning turn, 2006 to 2017
In 2006 Hinton and colleagues showed that deep belief networks could be trained by pre-training one layer at a time, and the phrase "deep learning" came into wide use. Two changes outside the algorithms mattered more. NVIDIA released CUDA in 2007, making graphics processors programmable for general matrix arithmetic. ImageNet, introduced by Deng and colleagues in 2009, supplied labelled images at a scale earlier datasets could not.
In 2012 AlexNet, from Krizhevsky, Sutskever and Hinton, won the ImageNet challenge by a wide margin. It was a convolutional network trained on two GPUs with ReLU activations and dropout. The architecture was not radically new. The data, the hardware and a few training techniques were. Within a few years most of the field had switched.
The following years produced most of the toolkit still in use: word2vec embeddings (2013), sequence-to-sequence models and attention for translation (2014), generative adversarial networks (2014), the Adam optimiser (2014), batch normalisation and residual networks (2015). Residual connections are what let networks hundreds of layers deep train at all. In March 2016 DeepMind's AlphaGo, combining deep networks, reinforcement learning and tree search, beat Lee Sedol at Go. The reinforcement learning side of that story is covered in reinforcement learning.
Transformers and scale, 2017 to now
"Attention Is All You Need" (Vaswani and colleagues, 2017) removed recurrence and built a sequence model entirely from attention and feed-forward layers. Because every position is processed in parallel during training, transformers use GPUs far more efficiently than recurrent networks. The mechanism is worked through in multi-head attention math. BERT (2018) showed the value of pre-training a bidirectional encoder on unlabelled text and fine-tuning it. The GPT line showed that a decoder trained only to predict the next token gains broad capabilities as it grows.
In 2020 Kaplan and colleagues reported scaling laws: test loss falls predictably as a power law in parameters, data and compute. In 2022 the Chinchilla work (Hoffmann and colleagues) revised the recipe, finding that many large models had been trained on too little data for their size. Instruction tuning and reinforcement learning from human feedback, described in the InstructGPT paper in 2022, made models follow requests rather than merely continue text. ChatGPT's release in November 2022 brought the result to a mass audience.
Read Richard Sutton's 2019 essay "The Bitter Lesson" alongside this period. It argues that general methods which scale with computation, namely search and learning, have repeatedly beaten methods built on human knowledge. Its argument is the contrast between expert systems and deep learning, stated as a principle.
Patterns across seventy years
| Era | Bottleneck | What removed it | What survived |
|---|---|---|---|
| Symbolic search | Combinatorial explosion | Not removed; work moved to narrower domains | Heuristic search, planning |
| Expert systems | Hand-acquired knowledge, brittleness | Learning rules from data | Rule engines, guardrails |
| Early neural nets | Training hidden layers | Backpropagation | The chain rule in every framework |
| Statistical ML | Hand-designed features | Learned representations | Evaluation discipline, boosted trees |
| Deep learning | Data and compute | ImageNet-scale data, GPUs | Convolutions, residuals, Adam |
| Transformers | Sequential computation | Attention, parallel training | Pre-training then adaptation |
Mistakes history warns against
- Forecasting from mood. Each era predicted general intelligence soon. Judge progress by evaluations you can reproduce, not by announcements.
- Overfitting to a benchmark. Once a benchmark becomes a target, results improve faster than capability does. Keep held-out and production-like tests.
- Brittleness outside the training distribution. Expert systems failed on inputs just outside their rules, and learned models fail on inputs just outside their data. The failure mode is the same, so the defence is the same: monitor inputs.
- Dismissing an idea because it failed at small scale. Neural networks, attention and self-supervision all looked marginal before they had the data and compute they needed.
- Ignoring cost. Expert systems died partly on maintenance cost, and Lisp machines on hardware cost. Training and serving cost shape today's choices in the same way.