Mechanistic interpretability

A neural network is trained, not programmed. When training ends it can add numbers or turn objects in space, but nobody has written down how it does it. We work out the algorithm a network has learned, and then we test whether that explanation is right.

Start where the answer is known

Large language models are too big to take apart one neuron at a time. On natural data there is also no ground truth to check an explanation against. So we start small. We train networks on tasks from group theory, the mathematics of symmetry, where every possible solution can be written down before training begins. Adding numbers on a clock is one such task. Combining the rotations and flips of a polygon is another.

Once a network has learned the task, we can ask which of the known solutions it uses, or whether it found a new one. Each paper below has a figure you can play with.

How a network adds

Take addition modulo 59. Numbers wrap around after 58, the way hours wrap around after 12. A trained network does not keep a table of the 3,481 possible sums. It places each number on a few circles, at an angle that depends on the number, and it adds two numbers by adding their angles.

Score of every possible answer, from 0 to 58

(23 + 41) mod 59 = 5 Highest score: 5

Fig. 1 An idealized version of what trained networks do. Each circle turns a number into an angle, at its own speed k. Adding a and b adds their angles. The bars score every possible answer. With one circle, several wrong answers score almost as high as the right one. Switch on more circles and only the right answer is left.

One circle is not enough, because several wrong answers sit almost as close to the right angle as the right answer does. The network uses a few circles turning at different speeds, and only the right answer scores well on all of them. In our NeurIPS 2025 paper we showed that this is an approximate form of the Chinese remainder theorem, an old method from number theory that splits a calculation into smaller ones and then combines the results. The two most common kinds of network, multilayer perceptrons and transformers, both learn it. In deep networks the number of circles grows only with the logarithm of the modulus.

The same shape in different networks

Two earlier studies found two different circuits for this task in two kinds of transformer. They called them Clock and Pizza, and concluded that the design of a network decides which algorithm it learns.

At TAG-DS 2025 we showed what really decides the shape. Each neuron in a cluster has two phases, one for each input. If the two phases are equal in every neuron, the cluster’s activity fills a disc. If they vary independently, it covers a torus, the surface of a doughnut. Clock and Pizza both keep the two phases nearly equal, so both learn a disc. A multilayer perceptron that sees the two inputs side by side, which we called MLP-Concat, learns the torus.

Phases of 40 neurons

Their joint activity is: a disc β = (1, 0, 0)

Fig. 2 The phases of 40 neurons in one cluster, for four architectures, and the shape their joint activity takes. Points on the dashed diagonal have equal phases. The Betti numbers β count the shape’s connected pieces, its loops and the voids it encloses.

Our ICLR 2026 paper made the comparison exact. Instead of reading neurons one at a time, we gathered the neurons that work together and studied their combined activity as one geometric object, across hundreds of trained circuits. The disc is a flattened view of the torus, the same algorithm seen through a projection. Clock and Pizza turned out to be the same circuit.

Colour by Drag to turn it
Fig. 3 Every pair (a, b) is a point on the torus. Flattening the torus gives the disc that transformers learn. Coloured by a + b, the disc splits into the slices that gave the Pizza circuit its name.

Grids of numbers

Addition modulo a prime can be stacked. In the elementary p groups, each element is a list of numbers modulo a prime p, added position by position. Our TAG-DS 2025 paper found that every neuron responds to a coset, a set of elements lying on one line through the grid, and that it orders the neighbouring lines by their distance around the torus, called the Lee metric. The network intersects the lines of a few neurons to find the answer, a multidimensional version of the Chinese remainder theorem.

Neurons switched on

Cells left: 7

Click a cell to choose the answer.

Fig. 4 The pairs of numbers modulo 7, drawn as a grid whose edges wrap around. Each neuron direction ξ lights up one line of cells, its coset, with the neighbouring lines fading by Lee distance. One direction leaves seven candidates. Two leave one.

Divide and conquer

The rotations and flips of a regular polygon form what mathematicians call the dihedral group. In our NeurIPS 2025 workshop paper we looked at single neurons in networks trained on it. A neuron is a wave over the group’s elements. When its frequency shares a factor with the number of sides, its values fall into a few exact levels, and each level is a coset. When it does not, every element gets its own value, an approximate coset.

Activation levels: 6, exact cosets

Outer ring: rotations. Inner ring: reflections. Solid lines multiply by r, dashed lines by s. Hover over an element to see the others at the same level.

Fig. 5 One neuron on the Cayley graph of the 36 symmetries of an 18-sided polygon. Colour is the neuron’s activation. At frequency 6 the values fall into three levels for the rotations and three for the reflections, six cosets in all. At frequency 5, which shares no factor with 18, there are no exact levels.

In our ICML 2026 paper we followed these neurons up to the whole network. Neurons group into clusters, the activity of each cluster has the shape of a Cayley graph, and each cluster solves an easier problem: which coset the answer lies in. The network adds the clusters’ votes at the output, and the element they all agree on is the answer. Only a logarithmic number of clusters is needed, and multilayer perceptrons and transformers learn the same solution.

Candidates left: 1

Outer ring: the rotations 0 to 14. Inner ring: the reflections 15 to 29.

Fig. 6 The example from the paper, in the 30 symmetries of a 15-sided polygon. One cluster narrows the answer to the five elements of its class that match it modulo 3. Another narrows it to the three that match modulo 5. Only the answer is in both.

Beyond Cayley graphs

In the groups above, the activity of each cluster is a Cayley graph. In the alternating groups, the even rearrangements of a list, that is not always so. Our TAG-DS 2026 paper, an oral presentation, found that when a cluster’s subgroup is normal, its activity is still a Cayley graph. When the subgroup is not normal, its cosets do not form a group, and the activity takes the shape of a Schreier coset graph, the more general object. This holds in A4, A5 and A6, for multilayer perceptrons and for transformers.

Fig. 7 The 12 elements of A4, collapsing onto the cosets of a subgroup. For a normal subgroup the result is the Cayley graph of a smaller group. For a subgroup that is not normal it is a Schreier coset graph.

Rotations

All of these groups are finite. Rotations in space are not. There are infinitely many of them, and they change smoothly from one to the next. In a paper now under review, we trained networks to compose rotations in three dimensions and higher.

Networks with one or two hidden layers did not find a clean solution. Deeper networks learned Rodrigues’ rotation formula, the standard way to turn a vector about an axis by a given angle. Single neurons respond to rotations about particular axes, and the later layers combine them term by term. In higher dimensions the neurons respond to rotations within planes instead of about axes.

Drag the sphere to turn it
Fig. 8 Rodrigues’ rotation formula, which deeper networks learn when they are trained to compose rotations. A vector v turned about the axis k by the angle θ is the sum of the three coloured parts.

The weights between layers

Most of this work describes what neurons do while the network runs. A second paper under review describes the weights that connect one layer to the next in deep networks trained on modular addition. The difficulty is that the neurons in a layer can be put in any order without changing the network, so there is no natural way to line the weights up and read them. We found an order hidden in the network’s own activity, in the phase of a two-dimensional Fourier transform. Sorted in that order, the weights between layers turn out to be sine waves.

Weights from one hidden layer to the next

  • positive weight
  • negative weight
Fig. 9 Weights between two hidden layers, in the form the paper describes: clusters of neurons, and within each cluster a cosine of the difference between their phases. In the network’s own order they look like noise. Sorted by cluster and by phase, the sine waves appear.

What interpretability is for

An explanation that is right is useful well beyond research. We see three uses.

Smaller models

Once the algorithm is known, the network can be collapsed into something much smaller. In our NeurIPS 2025 paper, replacing the trained neurons with the simple neurons the theory predicts left the network’s accuracy unchanged. A network that has learned a known algorithm can, in the end, be replaced by the algorithm.

Explanations for operators

A mine planner or a mineralogist has to act on what a model reports. If we know which mechanism produced an output, we can show them the reason, in terms they can check against what they know. The other two lines of our research all have someone like this at the end.

Models that can be gated

A model whose mechanisms are known can be gated. A capability that should not be used can be switched off where it is computed, and a model can be held back from deployment until it is shown to reach its answers through the right mechanism.

All three depend on networks trained in different ways arriving at the same algorithms, so that a tool built for one network works on the next. Every new group we study tests this. So far it has held, with one twist: for rotations, networks need depth before they split the problem into pieces.

Papers

  1. Interpreting SO(n) Multiplication: Deep Networks Generalize by Learning an Algorithm

    Arthur Ayestas Hilgert, Xiangzhuo Zeng, Sihui Wei, Gabriela Moisescu-Pareja, Vincent Létourneau, Gavin McCracken

    International Conference on Learning Representations (ICLR 2027) Under review

    Interactive figure

  2. The Form of the Weights in Deep Networks Trained on Modular Addition

    Sihui Wei, Arthur Ayestas Hilgert, Xiangzhuo Zeng, Gavin McCracken

    International Conference on Learning Representations (ICLR 2027) Under review

    Interactive figure

  3. Deep neural networks divide and conquer dihedral multiplication

    Sihui Wei, Gavin McCracken, Gabriela Moisescu-Pareja, Harley Wiltzer, Doina Precup, Irina Rish, Jonathan Love

    International Conference on Machine Learning (ICML 2026)

    Interactive figure

  4. On the Geometry and Topology of Representations: The Manifolds of Modular Addition

    Gabriela Moisescu-Pareja, Gavin McCracken, Harley Wiltzer, Vincent Létourneau, Colin Daniels, Doina Precup, Jonathan Love

    International Conference on Learning Representations (ICLR 2026)

    Interactive figure

  5. Toward a general understanding of neural representations learned by deep neural networks on group multiplications

    Arthur Ayestas Hilgert, Sihui Wei, Doina Precup, Gabriela Moisescu-Pareja, Gavin McCracken

    Topology, Algebra, and Geometry in Data Science (TAG-DS 2026) Oral

    Interactive figure

  6. Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks

    Gavin McCracken, Gabriela Moisescu-Pareja, Vincent Létourneau, Doina Precup, Jonathan Love

    Neural Information Processing Systems (NeurIPS 2025)

    Interactive figure

  7. Interpreting deep neural networks trained on elementary p groups reveals algorithmic structure

    Gavin McCracken, Arthur Ayestas Hilgert, Sihui Wei, Gabriela Moisescu-Pareja, Zhaoyue Wang, Jonathan Love

    Topology, Algebra, and Geometry in Data Science (TAG-DS 2025) Flash talk

    Interactive figure

  8. The Geometry and Topology of Modular Addition Representations

    Gabriela Moisescu-Pareja, Gavin McCracken, Harley Wiltzer, Vincent Létourneau, Colin Daniels, Doina Precup, Jonathan Love

    Topology, Algebra, and Geometry in Data Science (TAG-DS 2025)

    Interactive figure

  9. The Representations of Deep Neural Networks Trained on Dihedral Group Multiplication

    Gavin McCracken, Sihui Wei, Gabriela Moisescu-Pareja, Harley Wiltzer, Irina Rish, Jonathan Love

    NeurIPS 2025 Workshop on Symmetry and Geometry in Neural Representations

    Interactive figure