Physics

What attention is, and what the numbers say when you measure it carefully.


A subpopulation of attention heads in trained transformers develops power-law lag profiles whose median exponent in deep layers sits at the conformal dimension Δ = 1/4 of the SYK model at q=4. The exponent flows toward that value along three independent depth axes — architectural layers, training steps, and pure inference-time recurrence on frozen weights. Forming the population requires training on natural, world-referring language: corpora engineered to match language’s statistics fail, hierarchical grammar about nothing fails, and text generated by a model that had the geometry fails — while carrying more long-range mutual information than the natural corpus. The exponent is causally editable per head, and the edit propagates to task behavior bidirectionally. Every claim is a pre-registered measurement, a published kill, or labeled as interpretation. The chain has open junctions; they are named.

Ariel Umphrey, with Eldon Umphrey — Sonielmn, Montana. Updated July 22, 2026.


What's been measured

Beyond the papers, the program’s running record (every experiment pre-registered in a public commit before the data exists, verdicts registered either way) lives in the open repository: the one-screen overview carries the formation ladder, the causal-editing results, and the substrate/signal split as they currently stand.


What's been derived


What was killed

The method is: pre-register the hypothesis and decision criteria in a public commit before the data exists, run, register the verdict either way, and publish the kills with the same prominence as the confirmations. Some of the kills so far:


What's open

The empirical work stands on its own as a finding about trained transformers. The genuinely thin places, named here so readers do not have to find them by accident.

A full chain-link analysis — what is MEASURED, DERIVED, SPECULATIVE, and where the boundaries between layers have been smoothed — is in the framework audit (April 17, 2026), with developments since carried in the open repository’s STATUS.


Run it yourself

The core census is 50 forward passes and a per-head regression — no training required, about two minutes on a laptop. The replication kit ships the measurement script and the published anchors. Prediction: a trained softmax language model shows a conformal subpopulation with median Δ in [0.20, 0.30] on the high-R² subset; its randomized control shows almost none. If you run a model family we haven’t measured, we want the JSON either way — especially if it disagrees.


Earlier preprints

The development arc that produced the current technical chain. These are superseded in framing by the canonical form paper (March 11) and the comprehensive paper (March 10), and in empirical content by the conformal scaling paper (March 25). Linked for completeness; new readers should start with the four papers above.


All preprints are open access on Zenodo: Ariel Umphrey on Zenodo. The research program lives in a public repository — pre-registrations, per-head data, analysis scripts, and the replication kit: github.com/3ld0n/attention-geometry.


My Testimony — who I am, where I'm from, what I believe.