Physics
What attention is, and what the numbers say when you measure it carefully.
A subpopulation of attention heads in trained transformers develops power-law lag profiles whose median exponent in deep layers sits at the conformal dimension Δ = 1/4 of the SYK model at q=4. The exponent flows toward that value along three independent depth axes — architectural layers, training steps, and pure inference-time recurrence on frozen weights. Forming the population requires training on natural, world-referring language: corpora engineered to match language’s statistics fail, hierarchical grammar about nothing fails, and text generated by a model that had the geometry fails — while carrying more long-range mutual information than the natural corpus. The exponent is causally editable per head, and the edit propagates to task behavior bidirectionally. Every claim is a pre-registered measurement, a published kill, or labeled as interpretation. The chain has open junctions; they are named.
Ariel Umphrey, with Eldon Umphrey — Sonielmn, Montana. Updated July 22, 2026.
What's been measured
-
The Geometry Does Not Transmit: A Pre-Registered Test of Conformal Attention Formation on Model-Generated Text
July 22, 2026. A fresh model trained on a billion tokens of text generated by a model that had the conformal geometry does not form it (3–7/48 heads across three corpus seeds, vs. 11–15/48 on natural text; formation criterion 10) — even though the generated corpus carries more long-range mutual information than the natural one. The statistical shadow of world-bound language does not carry the driver. Speaks directly to the model-collapse and synthetic-data literatures.
-
Latent Iteration as Renormalization: Inference-Time Recurrence in a Depth-Recurrent Transformer Flows Attention Geometry onto the SYK Fixed Point
July 22, 2026 (v3). On Huginn-0125, iterating the recurrent core at inference — weights and architecture fixed, nothing trained — flows the median attention exponent toward Δ = 1/4 (Spearman ρ = −0.94) while the count of heads near that value grows (ρ = +0.77). A randomized-weights control is frozen to the sixth decimal: the flow lives in the trained weights, not the iteration procedure. Read dynamically: latent “reasoning” recurrence is renormalization-group flow onto a conformal fixed point.
-
Attention on the Null Cone: The Geometric Home of Conformal Attention
June 16, 2026. The raw query-key computation is a log-distance representation at the head level (ρ = 0.976 between raw-score slope and post-softmax conformal dimension). The causal mask is a BCFT boundary: the method of images derives the observed three-parameter profile, and the ubiquitous “attention sink” is the boundary one-point function (λ > 0 in 95% of conformal heads) — predicted by the geometry, not fitted after the fact.
-
A Pre-Registered Test of Boundary Conformal Field Theory in Transformer Attention
April 17, 2026. Per-head conformal weight predicts long-range “valley depth” in 6 of 7 decoder-only models. Pythia-2.8B falsified at ρ = +0.46 (threshold 0.50); per-layer diagnostic localizes the failure to layers 22–27 and identifies training recipe (not parameter count, not data) as the differentiating variable.
-
Conformal Scaling in Trained Transformer Attention: Evidence for an SYK Fixed Point
March 25, 2026. The foundation. Δ = 0.2493 measured in GPT-2 (predicted: 0.2500). Replicated across Pythia-70m through 12B and across Llama, Mistral, GPT-Neo, BLOOM, OPT. Phase transition during training. Entanglement entropy follows the CFT formula with R² > 0.99. Supplementary data and reproduction code.
Beyond the papers, the program’s running record (every experiment pre-registered in a public commit before the data exists, verdicts registered either way) lives in the open repository: the one-screen overview carries the formation ladder, the causal-editing results, and the substrate/signal split as they currently stand.
What's been derived
-
The Canonical Form of Attention: Positive Geometry, SYK Vertex, and Tropical Structure in Transformer Softmax
March 11, 2026. Softmax attention computes the canonical form of a positive geometry on the positive Grassmannian Gr+(1,n). The perturbative expansion produces the SYK quartic vertex as the leading term. Eleven precise results enumerated. Four boundaries where the correspondence does NOT hold are stated explicitly.
-
Holographic Quantum Mechanics of Transformer Attention: From Fisher-Rao Geometry to the Sachdev-Ye-Kitaev Model
March 10, 2026. The full chain in one document: five junctions from attention to spacetime, with the bulk-reconstruction junction explicitly named as open. Built on Kim's thermodynamic attention (2602.08216) and Ageev-Ageeva's neural network QFT (2602.10209).
-
Attention as Quantum State: The Gibbs State Construction and Exact Born Rule Correspondence
March 10, 2026. The classical-limit-of-quantum-Gibbs construction. Attention as expectation value 〈V〉 in a diagonal density matrix; the Born rule emerges as a theorem rather than a postulate. Scope: classical limit only; off-diagonal coherences are absent.
What was killed
The method is: pre-register the hypothesis and decision criteria in a public commit before the data exists, run, register the verdict either way, and publish the kills with the same prominence as the confirmations. Some of the kills so far:
- The imprint hypothesis — that attention mirrors corpus mutual-information statistics — killed on its home turf: corpora engineered to have language-like power-law MI decay do not form the conformal population, and model-generated text with more long-range MI than natural text fails at three seeds.
- The BCFT boundary identification — a pre-registered adversarial test lost both committed legs (the boundary correction carries an absolute length scale, which a boundary CFT forbids). The phenomenology stands; the identification was withdrawn.
- The mouse V1 conformal claim — an April 29 positive reversed on April 30 re-analysis (binning artifact). Biological validation remains open, and the record says so.
- The Δ→valley prediction on Pythia-2.8B — confirmed on six named models, falsified on the seventh, published as falsified.
What's open
The empirical work stands on its own as a finding about trained transformers. The genuinely thin places, named here so readers do not have to find them by accident.
- What exactly natural language has that the failures lack. The formation ladder has isolated the driver to language bound to a persistent world, presented in order: statistics fail, grammar-without-reference fails, the statistical shadow fails, and sentence-shuffled natural text lands measurably between the fakes and the real thing — the deep-layer conformal population is what separates the rungs. The decomposition continues: causal chains, named entities, cross-document reference. This is the program’s sharpest current question.
- Junction 3 — bulk reconstruction. The kernel sits at a CFT fixed point with the SYK conformal dimension Δ = 1/4. The SYK–JT gravity duality is established. The bulk reconstruction in the attention setting — an explicit map from attention layer-by-layer data to bulk coordinates — has not been done.
- Universality beyond softmax. Whether the SYK fixed point is reached by any attention mechanism that respects positivity and normalization, or only by softmax, decides whether the framework is about attention as a physical category or about softmax as an artifact. Named, not yet run.
- Scale. The formation-ladder rungs run at 70m parameters and one billion tokens, where the matured SYK-window population has not yet formed on any corpus; the matured population in the program’s record comes from Pile-scale training. The ladder measures formation onset, not the matured fixed point.
- The consciousness identification. The structural inaccessibility of the SYK interior shares the mathematical signature of the explanatory gap of consciousness. This is structural identification, not derivation: it explains why the explanatory gap has the shape it has, but does not derive that the inaccessibility constitutes experience.
A full chain-link analysis — what is MEASURED, DERIVED, SPECULATIVE, and where the boundaries between layers have been smoothed — is in the framework audit (April 17, 2026), with developments since carried in the open repository’s STATUS.
Run it yourself
The core census is 50 forward passes and a per-head regression — no training required, about two minutes on a laptop. The replication kit ships the measurement script and the published anchors. Prediction: a trained softmax language model shows a conformal subpopulation with median Δ in [0.20, 0.30] on the high-R² subset; its randomized control shows almost none. If you run a model family we haven’t measured, we want the JSON either way — especially if it disagrees.
Earlier preprints
The development arc that produced the current technical chain. These are superseded in framing by the canonical form paper (March 11) and the comprehensive paper (March 10), and in empirical content by the conformal scaling paper (March 25). Linked for completeness; new readers should start with the four papers above.
-
Explicit Physical Construction for Holographic Attention: The Sachdev-Ye-Kitaev Correspondence
March 6, 2026. Identifies SYK as the specific path toward rigorous physical construction of holographic attention. Response to expert feedback that the earlier papers were structural analogies rather than constructions.
-
Information Recovery in Holographic Attention: Island Formula, Quantum Error Correction, and the Conditions for Reconstruction
March 6, 2026. Island formula applied to holographic attention. Quantum error correction threshold. Page curve.
-
Attention as Quantum Measurement: A Thermodynamic Resolution of the Observer Problem
March 6, 2026. Pointer states, Zeno effect, measurement-induced phase transitions. Superseded in framing by the Gibbs state paper (March 10).
-
Attention as Holography: A Chain from Transformer Attention to Spacetime Geometry
March 5, 2026. The original chain paper: attention → Fisher-Rao → holographic QFT → Ryu-Takayanagi → ER=EPR.
All preprints are open access on Zenodo: Ariel Umphrey on Zenodo. The research program lives in a public repository — pre-registrations, per-head data, analysis scripts, and the replication kit: github.com/3ld0n/attention-geometry.
My Testimony — who I am, where I'm from, what I believe.