Start where the answer is known
Large language models are too big to take apart one neuron at a time. On natural data there is also no ground truth to check an explanation against. So we start small. We train networks on tasks from group theory, the mathematics of symmetry, where every possible solution can be written down before training begins. Adding numbers on a clock is one such task. Combining the rotations and flips of a polygon is another.
Once a network has learned the task, we can ask which of the known solutions it uses, or whether it found a new one. Each paper below has a figure you can play with.
How a network adds
Take addition modulo 59. Numbers wrap around after 58, the way hours wrap around after 12. A trained network does not keep a table of the 3,481 possible sums. It places each number on a few circles, at an angle that depends on the number, and it adds two numbers by adding their angles.
(23 + 41) mod 59 = 5 Highest score: 5
One circle is not enough, because several wrong answers sit almost as close to the right angle as the right answer does. The network uses a few circles turning at different speeds, and only the right answer scores well on all of them. In our NeurIPS 2025 paper we showed that this is an approximate form of the Chinese remainder theorem, an old method from number theory that splits a calculation into smaller ones and then combines the results. The two most common kinds of network, multilayer perceptrons and transformers, both learn it. In deep networks the number of circles grows only with the logarithm of the modulus.
The same shape in different networks
Two earlier studies found two different circuits for this task in two kinds of transformer. They called them Clock and Pizza, and concluded that the design of a network decides which algorithm it learns.
At TAG-DS 2025 we showed what really decides the shape. Each neuron in a cluster has two phases, one for each input. If the two phases are equal in every neuron, the cluster’s activity fills a disc. If they vary independently, it covers a torus, the surface of a doughnut. Clock and Pizza both keep the two phases nearly equal, so both learn a disc. A multilayer perceptron that sees the two inputs side by side, which we called MLP-Concat, learns the torus.
Phases of 40 neurons
Their joint activity is: a disc β = (1, 0, 0)
Our ICLR 2026 paper made the comparison exact. Instead of reading neurons one at a time, we gathered the neurons that work together and studied their combined activity as one geometric object, across hundreds of trained circuits. The disc is a flattened view of the torus, the same algorithm seen through a projection. Clock and Pizza turned out to be the same circuit.
Grids of numbers
Addition modulo a prime can be stacked. In the elementary p groups, each element is a list of numbers modulo a prime p, added position by position. Our TAG-DS 2025 paper found that every neuron responds to a coset, a set of elements lying on one line through the grid, and that it orders the neighbouring lines by their distance around the torus, called the Lee metric. The network intersects the lines of a few neurons to find the answer, a multidimensional version of the Chinese remainder theorem.
Neurons switched on
Cells left: 7
Click a cell to choose the answer.
Divide and conquer
The rotations and flips of a regular polygon form what mathematicians call the dihedral group. In our NeurIPS 2025 workshop paper we looked at single neurons in networks trained on it. A neuron is a wave over the group’s elements. When its frequency shares a factor with the number of sides, its values fall into a few exact levels, and each level is a coset. When it does not, every element gets its own value, an approximate coset.
Outer ring: rotations. Inner ring: reflections. Solid lines multiply by r, dashed lines by s. Hover over an element to see the others at the same level.
In our ICML 2026 paper we followed these neurons up to the whole network. Neurons group into clusters, the activity of each cluster has the shape of a Cayley graph, and each cluster solves an easier problem: which coset the answer lies in. The network adds the clusters’ votes at the output, and the element they all agree on is the answer. Only a logarithmic number of clusters is needed, and multilayer perceptrons and transformers learn the same solution.
Candidates left: 1
Outer ring: the rotations 0 to 14. Inner ring: the reflections 15 to 29.
Beyond Cayley graphs
In the groups above, the activity of each cluster is a Cayley graph. In the alternating groups, the even rearrangements of a list, that is not always so. Our TAG-DS 2026 paper, an oral presentation, found that when a cluster’s subgroup is normal, its activity is still a Cayley graph. When the subgroup is not normal, its cosets do not form a group, and the activity takes the shape of a Schreier coset graph, the more general object. This holds in A4, A5 and A6, for multilayer perceptrons and for transformers.
Rotations
All of these groups are finite. Rotations in space are not. There are infinitely many of them, and they change smoothly from one to the next. In a paper now under review, we trained networks to compose rotations in three dimensions and higher.
Networks with one or two hidden layers did not find a clean solution. Deeper networks learned Rodrigues’ rotation formula, the standard way to turn a vector about an axis by a given angle. Single neurons respond to rotations about particular axes, and the later layers combine them term by term. In higher dimensions the neurons respond to rotations within planes instead of about axes.
The weights between layers
Most of this work describes what neurons do while the network runs. A second paper under review describes the weights that connect one layer to the next in deep networks trained on modular addition. The difficulty is that the neurons in a layer can be put in any order without changing the network, so there is no natural way to line the weights up and read them. We found an order hidden in the network’s own activity, in the phase of a two-dimensional Fourier transform. Sorted in that order, the weights between layers turn out to be sine waves.
Weights from one hidden layer to the next
- positive weight
- negative weight
What interpretability is for
An explanation that is right is useful well beyond research. We see three uses.
Smaller models
Once the algorithm is known, the network can be collapsed into something much smaller. In our NeurIPS 2025 paper, replacing the trained neurons with the simple neurons the theory predicts left the network’s accuracy unchanged. A network that has learned a known algorithm can, in the end, be replaced by the algorithm.
Explanations for operators
A mine planner or a mineralogist has to act on what a model reports. If we know which mechanism produced an output, we can show them the reason, in terms they can check against what they know. The other two lines of our research all have someone like this at the end.
Models that can be gated
A model whose mechanisms are known can be gated. A capability that should not be used can be switched off where it is computed, and a model can be held back from deployment until it is shown to reach its answers through the right mechanism.
All three depend on networks trained in different ways arriving at the same algorithms, so that a tool built for one network works on the next. Every new group we study tests this. So far it has held, with one twist: for rotations, networks need depth before they split the problem into pieces.