Nobel Prize in Physics 2024: Hopfield Networks and the Boltzmann Machine
On this page
What was the Nobel Prize in Physics 2024 awarded for?
The Royal Swedish Academy of Sciences gave the award with this exact citation: "for foundational discoveries and inventions that enable machine learning with artificial neural networks".
In plain words, the two laureates did not build a computer chip or discover a new particle.
Instead, they took mathematical tools from physics, especially the physics of magnets, and used them to build networks of simple artificial "neurons" that can store patterns, recognise patterns and learn from examples rather than from step-by-step instructions.
This is the basic machinery behind most of what we now call artificial intelligence.
The official name of this award is the Nobel Prize in Physics. It was announced on 8 October 2024 and is one of six Nobel Prizes awarded each year, this one specifically by the Royal Swedish Academy of Sciences in Stockholm.
Who are the laureates?
John J. Hopfield
John J. Hopfield was born on 15 July 1933 in Chicago, Illinois, USA. At the time of the award he was a professor at Princeton University, New Jersey, USA, and he received one half of the prize. He had earlier earned his PhD in 1958 from Cornell University.
Hopfield's contribution was the Hopfield network, published in 1982: a model of "associative memory" that can store patterns and recreate the closest stored pattern when given a messy or incomplete version of it. He built this using ideas borrowed from the physics of magnetic materials.
Geoffrey Hinton
Geoffrey Hinton was born on 6 December 1947 in London, United Kingdom. At the time of the award he was a professor at the University of Toronto, Canada, and he also received one half of the prize. He completed his PhD in 1978 at the University of Edinburgh.
Working with Terrence Sejnowski and other colleagues between 1983 and 1985, Hinton extended Hopfield's network into the Boltzmann machine, a network that learns to recognise and generate patterns using the mathematics of statistical physics. He later helped drive the rise of deep learning in the 2000s.
| Laureate | Born | Affiliation at award | Share |
|---|---|---|---|
| John J. Hopfield | 15 July 1933, Chicago, IL, USA | Princeton University, USA | 1/2 |
| Geoffrey Hinton | 6 December 1947, London, UK | University of Toronto, Canada | 1/2 |
What problem were they trying to solve?
Early neuron models and learning rules
From the 1950s, researchers wanted computers to do what brains do effortlessly: recognise a face, finish a half-remembered word, or sort objects into categories without being told the rules.
Early computers worked like a recipe, taking fixed instructions and producing a fixed result, which suits arithmetic but not messy, ambiguous tasks such as reading a handwritten digit.
In 1943, Warren McCulloch and Walter Pitts (a neuroscientist and a logician) proposed that a brain neuron could be modelled as adding up weighted signals from other neurons and firing a yes-or-no output once the sum crossed a threshold.
This simple idea became the starting point for decades of work on both biological and artificial neural networks.
In 1949, psychologist Donald Hebb proposed that learning happens because the connection, or synapse, between two neurons that are active at the same time grows stronger. This became known as the Hebb rule, and it is still used as a basic principle for training artificial networks today.
From perceptrons to the physics-inspired revival
Researchers explored two broad network shapes: feedforward networks, where signals move in one direction from input to output, and recurrent networks, which allow feedback loops among nodes.
In 1957, Frank Rosenblatt built a feedforward network called the perceptron for interpreting images, with three layers of nodes and weights that could be adjusted systematically.
The perceptron drew wide attention, but it struggled with non-linear problems, a simple example being the "one or the other but not both" (XOR) puzzle.
In 1969, Marvin Minsky and Seymour Papert published an influential book spelling out these limitations, and this led to a pause in funding for neural network research for much of the 1970s.
Alongside this, a separate line of work took its inspiration not from logic but from magnetic systems in physics, building recurrent network models to study how many simple interacting parts can produce collective behaviour.
It was this physics-rooted thread, picked up and developed by this year's laureates in the 1980s, that reawakened interest in neural networks after the quiet years.
How does the Hopfield network store memories?
Building the energy landscape
John Hopfield was a theoretical physicist whose earlier work, in the 1970s, had looked at how electrons move between biomolecules and at error-correcting steps in biochemical reactions.
After attending a meeting on neuroscience, he became fascinated with how the brain's structure might give rise to computation, and in 1982 he published his associative memory model.
Hopfield's network has N nodes, each holding a binary value of 0 or 1, rather like black or white pixels in a picture. Every node is connected to every other node by a weighted connection.
At random times, a node's value is updated: the network adds up a weighted sum of the values of all other nodes, and if that sum is positive the node is set to 1, otherwise it is set to 0.
The weights are chosen using the Hebb rule, so that they reflect correlations between nodes in the patterns to be remembered, and crucially the weights are kept symmetric, which guarantees that the network settles down rather than oscillating forever.
Hopfield then defined an energy for the whole network, built from all the node values and all the connection weights, that always decreases or stays the same as the dynamics run.
This mathematics closely mirrors how physicists describe a magnet: the first equation plays the same role as the "molecular field" that aligns atomic magnetic moments in a solid, and the second plays the role of the energy of a magnetic configuration.
Hopfield, with a background in physics, recognised and exploited this correspondence directly.
- Feed a pattern (for example, an image) into the network by setting each node to black (0) or white (1).
- Adjust the connection strengths using the Hebb rule so that this pattern has low overall energy.
- Later, present the network with a distorted or incomplete version of a stored pattern.
- Go through the nodes one by one, flipping each node's value whenever doing so lowers the network's energy.
- Stop when no further change lowers the energy; the network has typically settled into the nearest stored pattern.
The stationary states that the network settles into, at the bottom of energy valleys, represent the stored memories.
Correcting and extending the memory
Hopfield used this property as a method for error correction or pattern completion: a system started with an incorrect or misspelled pattern is drawn toward the nearest energy minimum, where the correction occurs.
In Hopfield's original demonstration, the number of patterns that could be reliably stored and told apart was limited, and later researchers, including Hopfield himself, developed methods to store more patterns and to let each node take more than just two values, so that it could stand for a pixel of any colour rather than only black or white.
Hopfield also built an analogue version of the model with continuously varying node values rather than strict 0s and 1s, and showed that the same collective, pattern-correcting behaviour survived this change, which answered doubts that the effect was merely an artefact of the simple binary design.
Hopfield likened searching this network for a saved pattern to a ball rolling down a bumpy slope of hills and valleys, its motion slowed by friction: dropped near a valley, the ball rolls down and settles at the bottom, just as the network settles into its nearest stored memory.
Draw and label
Energy landscape of a Hopfield network
Draw a wavy line representing "energy" with several dips (valleys). Label each valley as a stored memory/pattern.
Draw a small ball placed partway up a slope near one valley, with an arrow showing it rolling down into that valley, to show how a distorted input pattern is corrected toward the nearest stored memory.
How does the Boltzmann machine learn to classify patterns?
From Hopfield network to Boltzmann machine
Storing and recreating one remembered pattern is one task, but recognising what a new, never-seen pattern represents needs something more.
When Hopfield published his associative memory paper, Geoffrey Hinton was working at Carnegie Mellon University, wondering whether machines could learn to sort information into categories the way people learn to recognise a cat or a dog after seeing only a few examples, without being given formal definitions.
Between 1983 and 1985, working with Terrence Sejnowski and other co-workers, Hinton took Hopfield's network as a starting point and built a stochastic, or probability-based, extension of it called the Boltzmann machine.
Instead of settling on one fixed low-energy pattern, each possible overall state of the network is assigned a probability of occurring, following the Boltzmann distribution, an equation from nineteenth-century physicist Ludwig Boltzmann that describes how probable a state is given its energy and a notional temperature.
The network's nodes are divided into visible nodes, which correspond to the patterns being learned, and additional hidden nodes, which allow the machine to represent more general and complex probability distributions than the visible patterns alone would allow.
Because the machine aims to match statistics of a whole training set rather than memorise one pattern, it is described as a generative model.
- Feed example patterns from the training data into the visible nodes.
- Adjust the connection weights so that these training patterns become the most probable states of the network.
- Run the trained network by repeatedly updating node values according to the probability rule.
- Let the network settle into a state; this state is a new pattern resembling the training examples.
- Use the network to recognise new, unseen examples that share traits with the training category, or to generate fresh examples of that category.
Restricted Boltzmann machines and pretraining
Hinton and his colleagues worked out a mathematically elegant gradient-based learning algorithm for setting the weights, but each training step required lengthy simulations, so the original Boltzmann machine was theoretically interesting yet impractically slow.
The breakthrough came with a slimmed-down version, the restricted Boltzmann machine, which removes connections within the visible layer and within the hidden layer, leaving weights only between the two layers.
For this restricted version, Hinton devised a much faster approximate training method, called contrastive divergence.
With Simon Osindero and Yee Whye Teh (and, in closely related work, Ruslan Salakhutdinov), he then showed in 2006 how stacking several restricted Boltzmann machines and training them layer by layer could give a deep, many-layered network a good starting point for its weights, after which the whole network could be fine-tuned using backpropagation.
This layer-by-layer pretraining picked up basic structures in raw data, such as corners in images, without needing labelled examples, and this pretraining was a milestone on the path to today's deep learning.
A related feedforward technique, backpropagation itself, had been demonstrated in 1986 by David Rumelhart, Hinton and Ronald Williams, who showed that networks with a hidden layer could learn tasks that were impossible for simpler networks without one.
Draw and label
Structure of a Boltzmann machine
Draw two rows of circles. Label the bottom row "visible nodes" (where training data enters) and the top row "hidden nodes".
Draw lines connecting nodes between the two rows to represent weighted connections, showing how information flows between the visible and hidden layers.
How did the discovery unfold?
| Year | Event |
|---|---|
| 1943 | Warren McCulloch and Walter Pitts proposed an early mathematical model of how brain neurons might cooperate. |
| 1949 | Donald Hebb proposed that learning strengthens connections between neurons that are active together, the "Hebb rule". |
| 1957 | Frank Rosenblatt proposed and built the perceptron, an early feedforward network for image interpretation. |
| 1969 | Marvin Minsky and Seymour Papert published a book highlighting limitations of networks like the perceptron, which led to reduced funding for the field. |
| 1982 | John Hopfield published his associative memory model, the Hopfield network, based on ideas from magnetic materials. |
| 1983-1985 | Geoffrey Hinton, with Terrence Sejnowski and others, developed the Boltzmann machine, extending Hopfield's model using statistical physics. |
| 1986 | David Rumelhart, Hinton and Ronald Williams demonstrated backpropagation for training multilayer feedforward networks. |
| 2006 | Hinton, with Simon Osindero and Yee Whye Teh, developed a layer-by-layer pretraining method using restricted Boltzmann machines, enabling deep networks. |
| 2024 | Hopfield and Hinton were jointly awarded the Nobel Prize in Physics for these foundational contributions, announced on 8 October. |
Why does this work matter?
Uses within physics itself
The explosion of machine learning over roughly the past fifteen to twenty years relies on the artificial neural network structures that Hopfield and Hinton helped establish.
Today's large networks, called deep neural networks, can contain more than one trillion parameters, compared with the fewer than 500 parameters in Hopfield's original 1982 model of 30 nodes, a striking illustration of how far the field has scaled up.
Within physics itself, there are several concrete uses. Artificial neural networks made searches for the Higgs boson at CERN's Large Electron-Positron collider more sensitive during the 1990s, and were used in analysing the data that confirmed the Higgs boson's discovery at the Large Hadron Collider in 2012.
They have also helped reduce noise in gravitational wave measurements from colliding black holes, identify exoplanet transits in data from the Kepler Mission, and process the imaging data behind the Event Horizon Telescope's picture of the black hole at the centre of the Milky Way.
Beyond particle physics and astronomy, neural networks are increasingly used to model and predict the properties of molecules and materials, including calculating protein structures that determine their function and searching for new materials, such as those suited to more efficient solar cells.
Uses in everyday life and open questions
The most spectacular example is AlphaFold, a deep learning tool for predicting three-dimensional protein structures from their amino acid sequences.
In everyday life, the same underlying methods power image recognition, language generation, and film or series recommendations. In healthcare, a randomised study showed that machine learning improved breast cancer detection from mammography screening images, and to motion correction techniques for magnetic resonance imaging (MRI) scans.
The committee chair, Ellen Moons, said the laureates' work "has already been of the greatest benefit," noting its wide use in physics, such as developing new materials with specific properties.
In her award ceremony speech, she added that while such tools can give fast answers, it remains humanity's "collective responsibility" to ensure they are used safely and ethically.
Although computers cannot think, machines can now mimic functions such as memory and learning, and two matters remain open: which further applications of machine learning will prove most valuable remains to be seen, and there is ongoing, wide-ranging discussion of the ethical issues surrounding the technology's development and use.
How does this connect to what you study?
If you study physics, the idea of an "energy landscape" in the Hopfield network is the same mathematics used to describe how magnetic materials settle into stable configurations, linking statistical mechanics directly to computer science.
If you study biology, the original inspiration, the brain's neurons and synapses, connects this prize to how real nervous systems are thought to learn, following Donald Hebb's rule that neurons which fire together wire together.
If you study computer science or mathematics, concepts like probability distributions, optimisation and gradient-based learning (used in backpropagation) that you meet in statistics or calculus are exactly the tools used to train these networks.
Seeing how a nineteenth-century physics equation from Ludwig Boltzmann ended up inside modern image-recognition software is a useful example of how older, pure scientific ideas can find completely new and powerful uses many decades later.
Quick facts for exams
The Nobel Prize in Physics 2024 was awarded jointly to John J. Hopfield (Princeton University, USA) and Geoffrey Hinton (University of Toronto, Canada), each receiving half the prize, for foundational work enabling machine learning with artificial neural networks.
It was announced by the Royal Swedish Academy of Sciences on 8 October 2024. Hopfield, born in Chicago in 1933, created the Hopfield network, an associative memory inspired by magnetic materials.
Hinton, born in London in 1947, created the Boltzmann machine using statistical physics, and later helped pioneer deep learning. The total prize amount that year was 11 million Swedish kronor, shared equally.
| Fact | Detail |
|---|---|
| Prize | Nobel Prize in Physics 2024 |
| Date announced | 8 October 2024 |
| Awarding body | Royal Swedish Academy of Sciences |
| Laureates | John J. Hopfield; Geoffrey Hinton |
| Country of birth | Hopfield: USA (Chicago); Hinton: United Kingdom (London) |
| Affiliation at award | Hopfield: Princeton University, USA; Hinton: University of Toronto, Canada |
| Shares | 1/2 each |
| Citation | "for foundational discoveries and inventions that enable machine learning with artificial neural networks" |
| Prize amount | 11,000,000 Swedish kronor |
Note: Source. The prize facts in this note are from the Nobel Prize's official site, nobelprize.org.
Glossary
- Artificial neural network — a computing structure made of nodes and weighted connections, loosely inspired by brain neurons and synapses.
- Node — an individual unit in a neural network holding a value, representing a neuron.
- Weight (connection strength) — a number describing how strongly one node's value influences another, representing a synapse.
- Hebb rule — the principle that connections between simultaneously active nodes should be strengthened, proposed by Donald Hebb in 1949.
- Hopfield network — John Hopfield's 1982 recurrent network that stores patterns as low-energy states and recreates the nearest stored pattern from an incomplete input.
- Energy (in a network) — a quantity, borrowed from physics, that is low when the network's node values match a stored pattern.
- Boltzmann machine — Hinton's network, developed with Sejnowski in 1983-1985, which uses statistical physics to learn probability distributions of patterns.
- Visible nodes — the nodes in a Boltzmann machine where training data is fed in.
- Hidden nodes — the internal nodes in a Boltzmann machine not directly fed data, used to capture more complex patterns.
- Restricted Boltzmann machine — a faster, simplified Boltzmann machine without connections within the same layer, used for pretraining.
- Backpropagation — an algorithm, advanced by Rumelhart, Hinton and Williams in 1986, for training multilayer networks by adjusting weights to reduce error.
- Deep learning — training of large, many-layered ("deep") neural networks, an area Hinton helped pioneer.
- Generative model — a model, such as the Boltzmann machine, that can produce new examples resembling its training data.
- Statistical physics — the branch of physics describing systems made of many similar interacting components, such as gas molecules or network nodes.
Common errors and misconceptions
- Misconception: The Nobel Prize in Physics 2024 was for inventing artificial intelligence itself. Correct: It was for foundational physics-based methods, the Hopfield network and the Boltzmann machine, that underpin today's machine learning.
- Misconception: Hopfield and Hinton worked together on one invention. Correct: Hopfield created the associative memory model in 1982; Hinton, with Terrence Sejnowski, extended it into the Boltzmann machine in 1983-1985.
- Misconception: The Hopfield network and Boltzmann machine are the same thing. Correct: The Boltzmann machine is a statistical extension of the Hopfield network, using probability rather than a single fixed low-energy state.
- Misconception: This prize means computers can now truly "think" like humans. Correct: The source material states that although computers cannot think, machines can mimic functions such as memory and learning.
- Misconception: Hinton's deep learning breakthroughs happened in the 1980s. Correct: His 1980s work was the Boltzmann machine; the pretraining method enabling deep networks came later, in 2006.
- Misconception: The prize amount goes entirely to one laureate. Correct: The 11 million Swedish kronor prize was shared equally between Hopfield and Hinton.
Exam-style questions with model answers
Q1. Who were the two laureates of the Nobel Prize in Physics 2024, and what was their joint citation? [2 marks]
- The laureates were John J. Hopfield (Princeton University, USA) and Geoffrey Hinton (University of Toronto, Canada). The citation was "for foundational discoveries and inventions that enable machine learning with artificial neural networks".
Q2. In which year was the prize announced, and by which body? [1 mark]
- It was announced on 8 October 2024 by the Royal Swedish Academy of Sciences.
Q3. Explain how the Hopfield network uses the idea of "energy" to store and recall patterns. [4 marks]
- Hopfield represented his network's overall state using a quantity equivalent to the energy in a magnetic spin system in physics. The network is trained by setting connection weights so that each pattern to be remembered corresponds to a low-energy state. When a distorted or incomplete pattern is fed into the network, the system updates its nodes one by one, changing a node's value whenever doing so lowers the overall energy. This process continues until no further change reduces the energy, at which point the network has typically settled at the nearest stored pattern, much like a ball rolling downhill into the nearest valley.
Q4. What is the Boltzmann machine, and how did Hinton develop it from Hopfield's model? [4 marks]
- The Boltzmann machine is a network developed by Geoffrey Hinton with Terrence Sejnowski between 1983 and 1985. It extends Hopfield's network using statistical physics, specifically the Boltzmann equation, which assigns a probability to each possible state of a system of many interacting parts. The machine has visible nodes, where training data is fed in, and hidden nodes, which help capture more complex patterns. It is trained by adjusting weights so that the training examples become the network's most probable states, allowing it to later recognise or generate similar patterns.
Q5. Describe how the Boltzmann machine evolved into a tool for deep learning. [5 marks]
- The original Boltzmann machine was theoretically interesting but slow in practice because training required time-consuming calculations across the whole network. Hinton developed a faster, simplified version called the restricted Boltzmann machine, which removes connections within the visible layer and within the hidden layer. For restricted Boltzmann machines, Hinton created an efficient approximate learning method. With Simon Osindero and Yee Whye Teh, in 2006 he then developed a procedure for pretraining multilayer networks by stacking restricted Boltzmann machines and training each layer in turn. This pretraining gave deep networks a good starting point, after which the connections could be fine-tuned using backpropagation. This approach was an important milestone towards what is now called deep learning, in which large, many-layered networks are trained on vast amounts of data.
Q6. Discuss why the Nobel Committee considered this work significant enough for a Physics prize, including at least two examples of its impact given in the source material. [6 marks]
- The committee's reasoning was that both laureates used fundamental tools from physics, the physics of magnetic spin systems for Hopfield and statistical physics for Hinton, to build methods that became the foundation of today's machine learning using artificial neural networks. This represents physics extending beyond traditional subject matter into computation and even life-like functions such as memory and learning. Beyond the historical importance, there are concrete uses within physics itself: artificial neural networks made searches for the Higgs boson at CERN's colliders more sensitive and helped analyse the data that confirmed its discovery in 2012; they have also been used to reduce noise in gravitational wave measurements, to identify exoplanet transits in Kepler Mission data, and to help produce the Event Horizon Telescope's image of a black hole. Beyond physics, the same deep learning methods underlie AlphaFold's prediction of protein structures and everyday tools such as image recognition and language generation. Committee chair Ellen Moons said the laureates' work "has already been of the greatest benefit," citing its use in developing new materials with specific properties.
Key takeaways
- The Nobel Prize in Physics 2024 went jointly to John J. Hopfield and Geoffrey Hinton for foundational work enabling machine learning with artificial neural networks.
- Hopfield's 1982 network stores patterns as low-energy states, borrowing mathematics from magnetic materials in physics.
- Hinton's Boltzmann machine, developed with Terrence Sejnowski from 1983 to 1985, uses statistical physics to learn and generate patterns.
- Both laureates' networks were inspired by how brain neurons and synapses are thought to work, including Donald Hebb's 1949 learning rule.
- Restricted Boltzmann machines and layer-by-layer pretraining, developed by Hinton, helped make modern deep learning practical from the 2000s onward.
- Artificial neural networks are now widely used within physics itself, including in discovering the Higgs boson and processing gravitational wave data.
- The prize carried 11 million Swedish kronor, shared equally between the two laureates.
- Open questions remain about which applications of machine learning will prove most valuable and about the ethics of its use.
Test yourself
What did John Hopfield's 1982 model allow a network to do?
It allowed a network to store patterns and recreate the closest stored pattern when given a distorted or incomplete version.
Which nineteenth-century physicist's equation gives the Boltzmann machine its name?
Ludwig Boltzmann, whose equation from statistical physics describes the probability of different states of a system.
Who proposed the rule that neurons which fire together strengthen their connection?
Donald Hebb proposed this learning rule in 1949.
What are the two types of nodes in a Boltzmann machine?
Visible nodes, where training data is fed in, and hidden nodes, which help model more complex patterns.
What development did Hinton lead in 2006 that helped enable deep learning?
With Simon Osindero and Yee Whye Teh, he developed a method for pretraining multilayer networks by stacking restricted Boltzmann machines layer by layer.
Name one application of artificial neural networks within physics.
Helping discover the Higgs boson at CERN, or reducing noise in gravitational wave measurements, or identifying exoplanets, or imaging a black hole.
How much was the Nobel Prize in Physics 2024 worth in total, and how was it shared?
It was worth 11 million Swedish kronor, shared equally between John J. Hopfield and Geoffrey Hinton.
