Adding up the orders

The frozen read hands back a pile of per-direction facts: this order here, that gauge there, a count. The theory of why singular models generalise#5 wants a single number, the global $\lambda$ that prices the whole model. Between the pile and the number sits an assembly problem, and it is not a sum by default: how per-direction orders combine depends on how their pieces of the singular set meet.1

Sum versus crossing

Two dead directions can relate in two basic ways, and each has a rule.

If their singular loci are independent, separate flat pieces that the model can occupy simultaneously, each contributes its own coefficient and the contributions add: $\lambda = \tfrac{1}{2k_1} + \tfrac{1}{2k_2} + \cdots$. Each independent direction is a separate charge on the model's complexity bill.

If their loci cross, meeting transversally at the point being read, the flattest piece dominates and the assembly takes the minimum of the per-direction coefficients, with the multiplicity $m$ counting how many directions attain it. A crossing is not two structures side by side; it is one structure that several directions describe at once, and it is priced once, at its cheapest rate.

Compose a model from a few dead directions, toggle each pair between independent and crossing, and watch the global $(\lambda, m)$ react. The presets load the calibration cells, the scalar deep-linear networks, whose closed-form answer is known exactly.

The deep-linear preset is worth unpacking, because it shows the two reads playing different roles. A generic line through the origin of the scalar depth-$L$ network reads the collapse at order $k = L$, one factor of flatness per layer; that is the per-direction diagnosis, and it is what the frozen read reports. The assembly works on the singular locus itself, which is $L$ hyperplanes (each "this layer's weight is zero") meeting at the origin: each hyperplane is crossed transversally at coefficient $\tfrac{1}{2}$, the crossing rule takes the minimum over the $L$ of them, and out comes $\lambda = \tfrac{1}{2}$ with $m = L$, Aoyagi's closed form. The steep $k = L$ read and the assembled $\lambda = \tfrac{1}{2}$ are not in tension; the first is the signature of looking through a crossing along a generic line, and the read uses exactly that steepness as the signal to route the direction to the structured assembly rather than price it naively.

That is not a toy validation. Across twelve enumerable cells, normal crossings, separable sums, their composites, and the scalar deep-linear family at depths two through five, the assembled $(\lambda, m)$ reproduces the closed-form thresholds to machine precision.1 Where the structure can be enumerated, the cheap per-direction reads add up to the exact global answer the deep theory demands.

Where the assembly breaks

One family of structures defeats blind assembly: the wide matrix deep linear networks, whose singular locus is determinantal. There the coordinate-by-coordinate count mis-books the structure and the assembled coefficient leaves the closed form; a better rotated read does not fix it, because the failure is in the assembly, not the per-direction order. Partial reads cross the boundary, a rigorous bracket that contains the true coefficient, a detector that flags the determinantal locus, a structured rule that resolves the shallow cases exactly, but a general resolver stays open. For one calibrated global number on such a network, the posterior sampler remains the practical tool, exactly as the cost ladder#11 said: the reads and the sampler are complements, the sampler for the one number where enumeration fails, the reads for the per-direction structure the sampler's number silently sums over.

What the count really counts

The same care applies to $m$. The volume read#8 counts collapsing directions, and that count equals the analytic multiplicity only when the dominant directions form a single crossing at one order. The assembly types each direction first, so the count is robust to padding, lower-order directions, higher-order ones, and gauges all excluded, and it under-counts in exactly one known way: when the multiplicity splits across several loci with equal thresholds, which then need enumerating separately. Gauge directions never enter $m$ at all; they are booked as an architectural multiplicity of their own, the $\sim(L-1)\,d^2$ reparametrisations plus the per-site normalisation kernels, present from initialisation and constant through training.

The last constant: fixed by the order, absorbed by the network

The Watanabe triple#5 has a third member, the singular fluctuation $\nu$, and it delivers this arc's parting surprise twice over.

The first half is a gift: $\nu$ is universal in the order. Isolate a single order-$k$ dead direction and its fluctuation is a fixed number depending on $k$ alone, $\nu(2) \approx 0.173$ and $\nu(3) \approx 0.278$, confirmed by exact-posterior sampling on isolated cells to within a few percent.1 Read the order and you have read the direction's fluctuation for free; no new instrument needed.

The second half is the correction. On a trained network the realized $\nu$ comes in below the universal value. The dead direction does not fluctuate in isolation; the live structure around it soaks up part of the data fluctuation the dead direction would otherwise carry, and the suppression grows with the overlap between the dead and live subspaces. Measured trained vision transformers sit at effective overlaps of $\rho_{\mathrm{eff}} \approx 0.62$ to $0.81$, squarely inside the suppression band, so their realized fluctuation is a fraction of the clean theory's.

The left panel is the gift: the estimated $\hat\nu$ converging onto the universal $\nu(k)$ as data grows. The right panel is the correction: the dead direction's contribution collapsing as the dead–live overlap rises, with the band real transformers occupy shaded. Together they are the current status of $\nu$: exactly predictable per direction, systematically suppressed in company.

The arc, closed

This is where the program's promise lands. A dead direction is a flat direction with an integer; the integer is readable from a slope, layer by layer, from either side, off the weights, along a trajectory or at a standstill; the reads classify what they find and refuse what they cannot; and where the structure can be enumerated, the per-direction integers assemble into the very constants, $\lambda$, $m$, $\nu$, that singular learning theory proved govern generalisation. The deep theory's quantities are not just estimable; on the right structures they are computable from parts, and the parts are things you can point to in the network.

Still open / underspecified. The determinantal boundary is the assembly's known edge, and a general exact resolver there is an open problem in the program's register. The multiplicity's equal-threshold splitting case needs enumeration by hand. And the $\nu$ suppression is isolated and measured, but a predictive law, given a network, compute the realized $\nu$ from its dead–live overlap without sampling, is not yet established.


  1. The sum-versus-crossing assembly, the twelve-cell machine-precision validation against Aoyagi's closed forms, the determinantal boundary, the typed multiplicity, and the ν-universality and absorption measurements are from Tejas Pradeep Shirodkar, Measuring Dead Directions: Decomposing and Classifying Singular Structure off Canonical Alignment, arXiv:2607.00603 (2026). ↩︎