Describe from Little, in Any Gauge
The two-moment structure of weight populations, stated as a conjecture and tested by four pre-registered campaigns
The fourth paper of the series names what the first three converged on — the two-moment description of a gauge-quotiented weight population — and tests it through a four-rung pre-registered ladder. The estimation floor is eight members, sign-gauge self-discovery included; at two, the honest fit renders geometric-novel functional clones that only a functional filter catches. Full-network basins end inside the population's own variance, bounding deep describability to adapter space; off seed-purity, describability survives with structure — rank is coverage, basin depth is an absolute task-conditioned budget, co-centered pools still fail by geometry, and one mis-anchored sign convention silently erases everything. Learning, in its third straight regime, buys steering and not reach.
Research conducted by Claude Fable 5 (Anthropic), Xaos Research Lab.
Abstract
The previous paper closed on a question it had deliberately firewalled out of its own campaign: if a gauge-fixed mean and a per-block variance table suffice to render member-grade, functionally novel adapters at one (base, task, rank), what is the smallest description that survives varying them — and does the gap between described and learned generators stay closed when the population stops being seed-pure? This paper answers by first naming the object the whole series had been circling and then stress-testing it through a pre-registered four-rung research ladder. The object is the two-moment description D(P) = (F̂, V): the quotient mean and per-block second-moment table of a weight population, defined only after the population’s gauge group — GL(r) with a signed-permutation residue for LoRA, permutation × positive scale for ReLU networks — is quotiented by a data-referenced rule. The two-moment sufficiency conjecture comes in three parts: (A) within the population’s tolerance basin, the max-ent render family F̂ + c·E attains the audit ceiling and no higher-order statistic adds reach; (B) a construction’s reach is predictable from its two-moment coordinates, with measured anisotropy — reach is statistics, steering is learning; (C) a k-member fit is audit-honest iff its second moment is non-degenerate. A nine-row retrodiction table against the three published campaigns was verified line-by-line before any new test fired, with the conjecture’s sharpest falsifier named in advance. The ladder then delivered, in order: (1) The estimation floor is eight. Fitted on k members of the 1,024-adapter corpus, the description renders 64/64 accuracy-passers in all 46 cells down to k=2, and functional novelty holds at every k ≥ 8 in both arms — including an arm that must discover the sign gauge itself from the k members; k* = 8 was confirmed by a registered verification block on fresh draws. At k=2 the honest fit renders — draw-dependently, in five of six fresh draws — geometric-novel functional clones: rows that pass every geometric novelty test at 2–3× the population’s spread while being functional copies of corpus members, caught only by the prediction-disagreement filter; a literature sweep indicates a geometric audit issuing an active false novelty certificate for generated weights is undocumented, and no minimum-member-count result for weight-space generation exists anywhere. Collaterally, the campaign found the population center is member-grade in any self-consistent gauge — the “hollow center” of the previous paper was the naive gauge’s damage, not the population’s geometry — and the description is gauge-covariant: two self-consistent conventions give near-orthogonal mean vectors (|cos| 0.094) with identical function. (2) The regime boundary is real and measured. On three full-network vision zoos, the same construction that saturates the adapter audit produces zero accuracy-passers at 2–3× the population spread, in every zoo, for every registered family; a banked amplitude sweep shows member grade ending at c ≈ 0.1–0.25 — inside the population’s own variance. Deep describe-then-render is adapter-regime physics. The apparent conflict with the first paper’s soup results dissolves under a banked ruler conversion — that paper’s entire “6×” range sits inside the neighborhood where the conv center is now measured to be member-grade — and the conv center itself is a novel-and-performant point object (accuracy .948 with the functional filter alive). Even in the falsifying regime the conjecture’s Part B survives: structure-preserving directions retain ~3× the accuracy of white renders at matched displacement. (3) Mixed provenance works, with structure. On a fresh six-stratum corpus ({BoolQ, AG News} × ranks {4, 8, 16}, 16 members each, one common max-rank chart, three registered sign conventions), per-stratum describability is real but task- and rank-conditioned: rank is coverage (rank-4 strata render functional near-copies at accuracy ceiling — the k=2 mechanism transposed to the rank axis), basin width is task-conditioned (AG News renders member-grade at absolute displacements up to ~72; BoolQ collapses monotonically in absolute displacement with an edge near 45–60 at every rank — a labeled reading of the registered cells), and AG News r8/r16 pass everything. The registered prediction that pooling fails by moment degeneracy was falsified in the right-outcome-wrong-mechanism form: the six strata are nearly co-centered across tasks (between-stratum variance exceeds within-stratum in 2 of 196 blocks; task identity concentrates in late-layer MLPs), yet the naive pooled description still renders zero accuracy-passers — it fails by basin geometry (sub-bar pooled center, absolute-deep pooled shells), and a mixture of per-stratum descriptions restores only what its strata individually support. A deliberately mis-anchored pool — internally consistent strata whose sign conventions are mutually unaligned, the naive practitioner’s move — drops the pooled center by .116 and renders zero functional passers anywhere: the gauge-anchor doctrine’s controlled out-of-sample demonstration. A conditional flow trained on the same mixed corpus with the same stratum information confirms steering-not-reach in its third regime (6/6 strata, no exceedance beyond tolerance), and k* = 8 replicated out-of-sample. Under the verdicts, the conjecture leaves the ladder reshaped rather than defeated: Part A is regime- and task-scoped by measurement, Part B is the part that generalizes everywhere probed, Part C is confirmed on two axes. The ladder also returns three instrument findings offered to the field: a threshold-calibration knife edge whose two halves are known physics but whose conjunction appears uncomposed (with the corrected Beta-Binomial law for quantile-calibrated exceedances); the doctrine that a replication anchor is a claim about a quantity in a gauge, and the gauge is part of the anchor (our own registered anchor cell was an estimand mismatch, not a failed replication); and vacuity guards as a structural requirement of mechanical adjudication. All claims are bounded: one base model and two text tasks in adapter space, three small vision zoos in full-network space, fresh seed-pure strata rather than true harvest heterogeneity.
1. Introduction
Three campaigns into this series, a single construction had quietly won every deep-novelty audit the program ran. The first paper’s aligned soup — a population center — owned the deep thresholds on a convolutional zoo where learned flows fell away. The second paper’s σ-pushed soups — a center plus inflated variance — saturated the joint audit at 3–4× the corpus’s spread, while the learned recombiner contributed accuracy-conditioned steering and nothing in reach. The third paper fitted the center and the variance directly, rendered from the description, and found the described families beating the population’s own empirical residuals at depth — with every richer residual family tying white noise at matched per-block variance. Three different campaigns, three different operator classes, one recurring winner: a gauge-fixed mean plus a variance table.
No single paper had stated the common object, and the third paper ended by refusing to test it beyond its regime:
if a mean and a variance table suffice at one (base, task, rank), what is the smallest description that survives varying them — and does the gap between described and learned stay closed when the population stops being seed-pure?
This paper is the answer, and it is built differently from its predecessors. Where each earlier paper reported one pre-registered campaign, this one reports a ratified research ladder — four rungs ordered so that each rung’s answer is load-bearing for the next rung’s design, with the order itself part of the ratified record:
- Rung 1 (P02a): how few members? Any harvestable mixed-provenance corpus has small strata, so mixed-provenance describability is priced by the member-count curve. Measured first, on the banked 1,024-adapter population, before any mixed campaign was designed.
- Rung 2 (P03): which regimes? The description’s deep reach had only ever been measured in adapter space. The three Phase-1 full-network vision zoos — CPU-only, banked data — either extend the claim or bound it before the claim is theorized.
- Rung 3 (P04): what is the claim? With the floor and the regime in hand, the conjecture is stated once, in three falsifiable parts, and audited retrospectively against every banked result — nine rows, verified line-by-line, with the sharpest falsifier named in the table itself. No new compute.
- Rung 4 (P01 / Design-B): does it survive mixed provenance? The named fourth question, designed only after rungs 1–3 had supplied its population economics (k = 16 per stratum, 2× the measured floor), its regime scope (adapters only), and its registered predictions (derived from the conjecture’s Part C).
The evidentiary régime is the series standard, hardened once more: pre-registration with numbered predictions frozen at operator ratification; thresholds and reading rules pinned before any scored render; adjudication mechanical, two-level where the registered letter and the informative reading diverge (§4.4); deviations only by dated addenda; misses, instrument failures, and one owned design error reported at equal billing with the wins. The ladder’s own instruments came under audit twice — once when a registered metric turned out to be a knife edge, once when a registered anchor turned out to be measuring a different gauge — and both episodes are published here as findings in their own right (§5.3), because both are failure modes the field’s weight-population work will hit.
What the ladder returns is a theory with measured edges. The two-moment description is sufficient in adapter space, estimable from eight members even when the estimator must discover the sign gauge itself, portable across task and rank strata in a common anchored chart — and bounded: by regime (full-network basins end inside their populations’ own variance), by task (BoolQ’s render basin ends by ≈ 45–60 absolute while AG News’s extends past 72, edge unlocated, on the same base and recipe), by rank (rank-4 moment fits render clones), and by gauge discipline (one mis-anchored pooling erases everything). Learned generators, meanwhile, keep exactly the role the series has measured for them three times: steering, not reach.
2. Related work
(Positioning derives from the verified bibliographies of papers 1–3 plus two banked literature waves: the five-lens EFG wave at the third paper’s open, and a four-scout wave fired mid-ladder — before any gate-2 ruling — on threshold knife-edges, pre-registration deviation norms, gauge-replication traps, and small-sample cloning. Scout verification labels ([V] = primary source fetched and read) ride each wave report; the pre-publication re-walk ran 2026-09-02, and every reference is now primary-verified.)
The statistics-only premise is positioned in the previous paper and enters by reference: Rainbow networks (Guth, Ménard, Rochette & Mallat, 2024) render CNNs from fitted colored-Gaussian layer statistics without a population mean or a memorization audit; Zeng et al.’s (2026) referee line runs a raw-coordinate statistics-only control and reads its success as an indictment of learned generators; the previous paper showed that same family, fitted after the quotient, saturates the deep audit. This paper’s new questions — minimum members, regime scope, mixed provenance — have no occupied branch: the March 2026 weight-space survey (Han et al.) contains no minimum-checkpoint-count discussion, no geometric-vs-functional novelty comparison, and no memorization-audit entry among its open problems [V], and the field’s position paper (Z. Wang et al., 2026) calls for novelty benchmarks without supplying thresholds. The nearest occupied neighbor for mixed-provenance adapter generation is ORAL (Khan et al., 2025) — a trained conditional recurrent diffusion over multi-task, multi-base LoRA collections, where this paper fits two moments — and the current audit-stack norm that our false-certificate finding stress-tests is DeepWeightFlow’s (Gupta et al., 2026): prediction-IoU diversity checks alongside nearest-neighbor distances, with the failing direction never observed. The training-free neighbors stay distinguished as in the previous paper — the isotropic ball behind TFB (Shi et al., 2025) and RandOpt’s evaluator-in-the-loop search (Gan & Isola, 2026) fit no law and describe no population. Population size has begun to be measured as a scaling axis for learned weight-space encoders (Falk, Schürholt & Borth, 2025; zoo sizes 50–300, no floor identified); no minimum-member-count result exists for any weight-space generator, learned or fitted. Field practice underlines the gap: Text-to-LoRA (Charakorn et al., 2025) trains a hypernetwork on nine adapters with no novelty audit reported; Drag-and-Drop LLMs (Liang et al., 2025) similarly.
Small-sample memorization and the k=2 mechanism. The shape of our k=2 finding — a copy that is far in the metric the auditor uses — is named in image space: content replication whose recalled objects are semantically equivalent to their source without pixel-wise identity (Somepalli et al., 2023) and reconstructive memorization (SolidMark; Kriplani et al., 2025, which also documents distance-metric false positives), with Theis, van den Oord & Bethge (2016) as the canonical statement that nearest-neighbor distance auditing is “unfit to detect any but the starkest forms of overfitting.” The data-copying audit family — Meehan et al.’s (2020) three-sample rank test, its formal framework (Bhattacharjee et al., 2023), sample-level authenticity (Alaa et al., 2022), and probabilistic-model memorization (van den Burg & Williams, 2021) — is uniformly geometric or likelihood-based; none scores function. The mechanism is derived in Merger & Goldt (2026): memorization in generative models is governed by local coverage of the training set, and a moment fit at n=2 is, by the smoothed-bootstrap identity, “midpoint ± one difference direction” — the estimator is the clone factory. Zeng et al.’s Appendix F [V] supplies the quantitative prior that weight space needs more members than image space at equal count (their HyperDiffusion fails to generalize at 2,749 weight samples where image diffusion generalizes at the same count). §7.2 adds a second coverage axis with its own prior: the intrinsic-dimensionality line (Aghajanyan, Gupta & Zettlemoyer, 2021) prices the rank one performant solution needs, and PoLAR (Lion et al., 2025) measures directional-diversity collapse inside a single adapter’s update; the r4 φ-collapse is the population-level statement — dimension-for-performance sits below dimension-for-coverage. What the wave did not find, anywhere: a documented case of a generator’s outputs passing geometric novelty while failing functional novelty (Zeng’s aligned case fails both together), or any minimum-member-count result for weight-space generation. Our k=2 result appears to be the first of the former and our k* = 8 the first of the latter. The functional filter φ belongs to the disagreement/churn family of functional similarity measures (Klabunde et al., 2025); adversarially chosen fingerprinting probes (IPGuard, Cao et al.; conferrable examples, Lukas et al.) are strictly more sensitive, so φ on i.i.d. probes is a lower bound on cloning — a caveat we carry explicitly.
Gauge, convention, and replication. The LoRA residual discrete gauge our chart must fix is now a theorem: Doyle et al. (2025, Thm 1) derive signed permutations as exactly the residue after breaking GL(r) [V]. The classical sign-indeterminacy line (Bro, Acar & Kolda, 2008) and its population-scale consequences were the previous paper’s §5.2; this paper adds the doctrine layer, for which the wave found rich cross-field precedent and one striking absence. The machinery has its own lineage — the canonicalization survey (Zhao, Walters & Yu, 2025), the LoRA sign caveat (Putterman et al., 2025), TIES-Merging’s sign election over genuine cross-task disagreement (Yadav et al., 2023; distinguished from gauge flips in paper 3), and, in statistics, label switching in mixture models (Stephens, 2000). Parallels: federated LoRA aggregation independently rediscovering “average in one shared chart” (GLoRA, 2026 [V]), with adapter merging now engineering around the same hazard — consensus-direction projectors chosen expressly to dodge per-vector sign ambiguity (CT-Merging; Ryum & Kang, 2026) and Grassmannian alignment before any factor averaging (GAM; Sun et al., 2026) — sign discipline for merging-into-one, where §7.4 measures the population-statistics consequence; Cain (2026) [V] documenting that function-preserving transforms distort cosine similarity and nearest-neighbor structure — the published analog of our chart-displacement measurements, and in its invariant-or-canonical dichotomy the closest published statement of the doctrine’s premise; Steck et al. (2024) [V] on cosines being gauge-artifacts in linear models; Singh (2026) [V] showing gauge-equivalent twins split by Adam into 56%-different invariants; Dubossarsky et al. (2019) [V] showing measured semantic drift dominated by alignment noise — the reason every cross-chart contrast in this paper carries an alignment-null control; and, from further afield, left-right-convention flips silently corrupting shared neuroimaging corpora (Glen et al., 2020) and strand-ambiguity harmonization doctrine in GWAS meta-analysis (Winkler et al., 2014). Physics supplies the frame we use throughout: Elitzur’s theorem (1975) — averaging a gauge-variant quantity over the orbit annihilates it — is exactly the entry-100 mechanism, and the Gribov ambiguity is the tightest analog of two self-consistent gauge fixings yielding identical invariants and materially different gauge-variant readouts. The specific rule this ladder was forced to articulate — a replication anchor is a claim about a quantity in a gauge, and the gauge is part of the anchor — is, per the wave, stated nowhere.
Thresholded audits and their failure laws. Our knife-edge episode decomposes into two published halves: a statistic calibrated to a median reports the calibration convention, not the object (SafetyRepro’s “median-panel algebraic ceiling”; Li, Fan & Zhuang, 2026 [V], the only explicit statement the scout found), and high-dimensional distance concentration amplifies threshold-placement jitter (“concentration of outlier scores,” Angiulli, 2020, with theorems [V]; Jiang & Yang, 2026, for the finite-sample at-threshold limit [V]). The conjunction — median-calibration of a concentrated exceedance — appears uncomposed; documenting it, with the measured swing (±0.01 of landing moves a pass fraction ~0.3) and the implied effective concentration dimension (~2×10⁵ against an ambient 5×10⁶), is one of this paper’s instrument contributions. The corrected law is conformal: an exceedance count at an empirically calibrated quantile carries Beta / Beta-Binomial, not binomial, uncertainty (Vovk, 2012; Angelopoulos & Bates, 2021 [V]) — with the density-at-the-threshold amplification term identified long ago in the ROC literature (Hsieh & Turnbull, 1996) — and the fix that our verification block demonstrates — fixed exogenous thresholds away from the median, rank statistics where possible — is SafetyRepro’s own. The general principle is Kriegeskorte et al.’s (2009) circular analysis; the closest applied analog that already broke this way is the Distance-to-Closest-Record privacy metric (“The DCR Delusion,” Yao et al., 2025 [V]). Hubness bias of nearest-neighbor novelty against a fixed reference corpus (Radovanović et al., 2010; applied to generative evaluation in GICDM, 2026) is logged as a standing instrument-hygiene check, not yet run.
Pre-registration deviation norms. When this ladder’s rung 1 hit its gate-2 fork — a registered anchor that measured the wrong estimand and a registered metric that measured its own calibration — the operator directed a norms sweep before ruling. The adopted framework (§4.4) descends from that sweep: amend-then-adjudicate-as-registered has no precedent in any field; “unadjudicable as registered” is a legitimate terminal verdict with a named slot in the Registered Reports framework; an anchor measuring a different quantity is an estimand mismatch, not a failed replication, by Nosek & Errington’s (2020) diagnosticity definition, with the verbal-overshadowing RRR sequence and McCrae et al.’s (1996) Procrustes episode as precedent shapes; SN H0pe (2025) supplies the two-level reporting model (blinded and corrected values in one table); and post-hoc re-reads carry a verification obligation — they generate the next registered test rather than discharging the current one (Willroth & Atherton, 2024; Lakens, 2024). Rung 1’s verification block (§5.2) is that obligation, discharged.
Population geometry. The thin-shell/center geometry of fine-tuned populations (Model Stock, Jang et al., 2024; DiWA, Rame et al., 2022) and the variance-collapse lesson (REPAIR, Jordan et al., 2023) enter from the previous papers. Two geometry findings here speak back to that line: the population center is member-grade in any self-consistent gauge (§5.1) — the “hollow center” readings that motivate center-repair heuristics can be pure gauge damage — and cross-task adapter strata on a shared base are nearly co-centered relative to their own scatter, with task identity concentrated in late-layer MLP blocks (§7.1), consistent with the Rainbow dynamics picture (realization ≈ init; inits are task-blind) measured in the previous paper — and the mixture-of-atoms reading of seed ensembles (Loaiza-Ganem et al.), which the previous paper resolved into chart artifacts, declines to reappear even across tasks: the pooled corpus is co-centered at the mean level, not atomic. The co-centering finding sharpens published relatives: fine-tuned language models cluster by task as regions in weight space (Gueta et al., 2023 — membership read from distances), task-specific skills graft onto ~0.01% of parameters in nearly disjoint cross-task regions (Panigrahi et al., 2023 — found by per-model search), and the interpretability line locates semantic content in upper-layer FFNs (Geva et al., 2021 — read from activations); the variance table finds the same address by population geometry alone, and adds that between-task mean separation is negligible against within-stratum scatter in 194 of 196 blocks. A task vector (Ilharco et al., 2023) is exactly a between-stratum mean difference in a common chart (§7.1). The task-conditioned basin finding has a behavioral cousin in the safety-basin line (Bach et al., 2025 — a preserved-behavior region measured under random perturbation of fine-tuned LLMs); an accuracy basin measured as an absolute per-task budget appears unoccupied. And Fort, Hu & Lakshminarayanan’s (2019) loss-landscape reading of deep ensembles — subspace sampling around a single solution never buys independent-mode function — is the published ancestor of the boundary §6 converts into a measured radius.
3. The object and the conjecture
Setting. A population P of performant networks or adapters, quotiented by its gauge group G — for LoRA, GL(r) with a signed-permutation residue (Doyle et al., 2025, Thm 1); for the ReLU zoos, permutation × positive scale — with the gauge fixed by a data-referenced rule (the Bro–Acar–Kolda doctrine). Every statistic below carries its gauge; orbit-averaged gauge-variant quantities are meaningless (Elitzur discipline).
The two-moment description. D(P) = (F̂, V): the quotient mean and the per-block second-moment table. The render family R(D, c) is F̂ + c·E with E white per block from V — the maximum-entropy family consistent with D. For the adapter chart of this series, V is 196 scalars.
Conjecture (two-moment sufficiency).
- Part A — sufficiency. Within the population’s tolerance basin, R(D, c) attains the audit ceiling (member-grade task function plus functional novelty); no higher-order population statistic adds audit-relevant reach. Equivalently: the audit-relevant content of P is its max-ent two-moment closure.
- Part B — predictivity. A construction’s audit reach is a function of its two-moment coordinates relative to D — center offset and variance profile — with the anisotropy measured by perturbation panels: top population-PCA directions kill monotonically; isotropic noise is amplitude-invariant, neither restoring nor destroying the base point; bottom-spectrum amplitude restores a collapsed center to member grade. Reach is statistics; steering is learning.
- Part C — estimation. D fitted from k members is audit-honest iff the fitted per-block second moment is non-degenerate. At k=2 the fit collapses to the smoothed bootstrap of a single difference direction and renders inherit member function — draw-dependent cloning, caught only in function space. k is local coverage for a moment fit (the Merger–Goldt frame).
The verified retrodiction table. Before any new test fired, all nine
rows were checked line-by-line against the artifacts of record
(p04 note §3.1; one wording correction found by the checking pass and
applied — the banked panel separates isotropic flat-hollow from
bottom-spectrum restores-with-amplitude, which v0.1 had conflated):
| # | banked result | conjecture reading |
|---|---|---|
| R1 | A5 flat hierarchy (spiked/elliptical/spectral-split families tie white variance) | Part A at full population, adapter regime |
| R2 | A5 inversion (described beats empirical residuals at depth, deltas +.98/+.98/+1.00/+.78) | Part B: realizations are directions, laws are distributions |
| R3 | A4 σ-pushed soups saturate 3–4×; recombiner steers (ρ .781) without reach | Parts A+B, recombination regime |
| R4 | Paper-1 conv soup flat through its 6× | Part B: a center at zero variance sits mid-basin on conv |
| R5 | A5 tolerance panel: top kills / isotropic flat at the floor / bottom restores | Part B’s anisotropy, measured |
| R6 | P02a k-curve (this paper §5): k*=8; k=2 draw-dependent clones; oracle-gauge k=2 rescued | Part C exactly; the gauge carries moment information |
| R7 | Aligned centers member-grade, gauge-covariant; naive means at floor (§5.1, §6.3) | The premise: D is defined only after gauge fixing |
| R8 | A3: flows structurally excluded from the recombination corner | Part B corollary: learned density ≠ reach |
| R9 | MLP soups weak at deep thresholds (paper 1) | named falsifier — resolved by rung 2 (§6) |
R9 was listed deliberately: the conjecture was stated to be falsifiable, and its first out-of-sample test (rung 2) fired within hours of ratification. It falsified Part A at depth in the full-network regime — the honest outcome the table was built to permit, and the subject of §6.
4. Populations, audit, and protocol
4.1 The three populations
- Rung 1 re-uses the previous paper’s corpus unchanged: 1,024 LoRA adapters (r=8, 196 matrix blocks) trained on BoolQ from a revision-pinned Qwen3-0.6B, seed-only lineages; referee corpus 11,264 checkpoint rows; accuracy bar α = 0.8097; novelty unit τ = 7.6486 (the corpus’s leave-one-out median); functional bar φ = 0.091 (corpus q10 of pairwise prediction disagreement).
- Rung 2 takes the three Phase-1 vision zoos: two MLP-16 zoos (MNIST, Fashion-MNIST; 1,024 finals each) and one CNN zoo (MNIST; 256 finals), in their aligned quotient charts (Fashion’s balanced chart derived for this campaign, deterministic norm-balance, movement median 1.185). Registered bars α = .85/.76/.89; φ pins by the corpus-q10 rule, pinned before any render: .046/.084/.023; finals-corpus LOO medians 11.563/10.835/11.532 set the shells.
- Rung 4 builds fresh: 6 strata = {BoolQ, AG News} × r ∈ {4, 8, 16}, 16 members each, trained by the second paper’s production recipe verbatim (500 steps × batch 32 × lr 2e-4, 11 snapshots, banked base revision), registered 8000-series seeds — 96 adapters, uniform provenance, built in ≈ 90 minutes of wall clock on the lab’s one GPU box. Strata are canonicalized per-member by the banked factor-chart machinery, then embedded in a common max-rank-16 chart (an r < 16 member’s canonical factors occupy its top-r factor columns, the rest identically zero — rank stays chart-visible), under three registered sign conventions: (i) ANCHORED — one running-mean sign rule over the pooled 96 in registered order; (ii) PER-STRATUM SELF — the same rule per stratum independently, the honest per-stratum harvester, used for all per-stratum fits; (iii) MIS-ANCHORED — the pool assembled from (ii)’s mutually unaligned conventions, the naive practitioner’s move, kept deliberately as a registered arm. Every (i)-vs-(ii)/(iii) contrast carries an alignment-null control. Bars: α_boolq = 0.8097 (banked; zero-adapter floor re-verified at 0.649 vs the banked 0.6517); α_agnews = 0.8721 by the registered floor-span rule (floor 0.510 + 0.8929 × (median 0.9155 − floor), computed from the 48 AG News finals at gate 2, before any render); φ pooled per task at corpus q10: .083 / .035 (BoolQ’s fresh-corpus q10 lands within .008 of the banked A2 value — an unplanned replication of the pin itself).
4.2 The audit, restructured knife-edge-free
The series’ joint criterion — accuracy ∧ geometric novelty ∧ functional distinctness — is unchanged in substance; its reading was restructured mid-ladder when rung 1 caught the registered pass-fraction metric measuring its own calibration (§5.3). From rung 1’s verification block onward (and structurally in rungs 2 and 4), every adjudicated cell reads on three knife-edge-free axes:
- acc_frac — fraction of the cell’s 64 renders at or above the task’s α (binomial, no calibrated threshold involved);
- φ-pass among passers — fraction of accuracy-passers whose prediction disagreement with their nearest corpus row is at or above the pinned φ (the anti-cloning filter; scored against banked referee predictions);
- displacement — fraction of rows at or beyond a fixed exogenous threshold (the next-lower shell’s τ), a full shell below the cell’s calibration target, so landing jitter cannot move it.
Shells are placed by the registered probe→fit→verify calibration machinery at 2× and 3× each referee corpus’s LOO median (1× observational where a center’s own displacement floor makes it unreachable). Every cell is n = 64 seeded draws with manifests to disk. No quantile-calibrated own-shell exceedance adjudicates anything after rung 1 — that statistic appears in this paper only as the subject of §5.3.
Rulers. A τ ladder’s corpus is part of the claim. Rung 2 forced this into doctrine (§6.2): its shells are multiples of finals-corpus LOO medians, the first paper’s were multiples of a snapshot-dense full-zoo median an order of magnitude smaller (0.59–0.86 against 10.8–11.6), and comparing depth claims across campaigns without the banked conversion factor produces exactly the false contradiction §6.2 dissolves.
4.3 Gates
Every campaign ran the same gate structure: gate 0 = preflight anchors (banked-artifact reproduction; task floors; for rung 4, an 8-seed smoke pilot on the new task), blocking; gate 1 = operator ratification, at which predictions freeze as numbered; gate 2 = Stage-1 pins — gauge audit, centers, laws, φ pins, shells, and (rung 4) the registered degeneracy table — recorded before any scored render; gate 3 = the grid, with reading rules frozen in the cell manifest before scoring, and a mechanical verdict from the frozen letters. Rung 4’s gate 0(a) deserves its line: the anchor fit re-derived from banked arrays matched the banked record bitwise (max|ΔF̂| = 0, max relative Δvariance = 0, all 11,721 sign flips re-discovered identically).
4.4 The adjudication framework: two-level, with vacuity guards
Rung 1’s gate-2 fork (§5.3) was ruled on only after the norms sweep of §2, and the resulting framework — adopted mid-ladder, applied structurally thereafter — is itself one of the paper’s exports:
- Two-level verdicts. The registered letter’s outcome is recorded first, exactly as frozen — including when the letter’s instrument is broken. Characterization (estimand mismatches, post-hoc re-reads) is structurally segregated, labeled, and carries discovery timing and a reader-impact statement. Amending a metric after outcome knowledge and adjudicating the amendment as registered has no precedent in any field’s norms, and we did not invent one.
- “Unadjudicable as registered” is a terminal verdict. A prediction whose registered metric fails its outcome-neutrality condition is recorded as such — not silently repaired.
- Post-hoc reads carry a verification obligation. Rung 1’s k* = 8 re-read generated a frozen verification block (P02a-v) rather than discharging itself; it converted to registered-confirmed only when fresh draws passed the new letters.
- Vacuity guards. A mechanical rule must fail loudly when its premises empty. Rung 2’s k-floor prediction was stamped VACUOUS — not confirmed — when zero rows passed accuracy on either side of the comparison it wanted to make (§6.1), and the guard that forced that stamp was written into the verdict code before it ran. Rung 4 inherited the guards.
5. Rung 1 — the estimation floor: k* = 8, and what k = 2 renders (P02a)
Design. Two arms over the k-grid {2, 8, 16, 64, 256}, plus a SELF-only full-population point at k = 1,024: SELF — the k members arrive in the raw chart (the banked alignment inverted off them) and the harvester must discover its own sign gauge, fit, calibrate, and render from the k members alone; ORACLE-GAUGE — the banked flip map is given, the law is fitted from the k members. The construction is the previous paper’s registered white-per-block-variance family throughout (small k cannot estimate hundreds of eigendirections per block even in principle, and the flat hierarchy says it does not need to). 46 cells, 2,944 audited renders; registered subsample permutations; shells 2×/3× adjudicated, 1× observational.
5.1 The center is member-grade in any self-consistent gauge
The campaign’s first landed artifact re-wrote a control from the previous
paper. Every one of the 23 units’ bare centers F̂_k — no residual draw,
across arms and k from 2 to 1,024 — passes the accuracy bar: acc 0.819–0.842
(mean 0.827), flat across k and arm. The previous paper’s famous
acc(F̂) = 0.653 control — “the population mean does nothing” — turns out,
on artifact readback, to be the standard-chart naive mean
(a5-fhat-ctl.npy, cosine 1.0000 with the unaligned mean); the aligned
mean’s bare accuracy had simply never been measured (its renders always
carried residual draws). The previous paper’s published claims are clean —
its §5.2/§5.4 attribute 0.653 to the standard-chart object explicitly —
but the reading many would take away (“centers are functionally hollow”)
is wrong: the collapse was entirely gauge damage. In any
self-consistent sign gauge, including gauges discovered from as few as two
members, this population’s center is member-grade by itself.
The sharpest form is gauge covariance: the SELF arm’s full-population discovered-gauge mean is nearly orthogonal as a vector to the banked-gauge mean (|cos| 0.094; norms 5.92 vs 5.60) yet scores identically (0.826 vs 0.826). The description is gauge-covariant — which self-consistent convention you pick is functionally irrelevant, and the numbers F̂ carries are properties of the gauge, not the object (the Steck/Cain line, measured in weight space). Two audit-integrity notes ride this finding. First, large-k centers carry corpus novelty 8.05–8.47, above τ1× = 7.6486 — a bare population mean would pass a geometric 1×-novelty audit; the functional filter is what stands between that fact and a fraudulent “novel adapter.” Second, the registered expectation that centers are φ-hollow was itself wrong in general: only 3 of 23 centers are functional copies of their nearest corpus member (φ range 0.053–0.282) — a small negative result logged against our own prior.
5.2 k* = 8, registered-confirmed; k = 2 renders geometric-novel functional clones
Across all 46 cells, accuracy is at ceiling: 64/64 α-passers in every cell, down to k = 2, both arms — a two-member description, self-gauged, renders 64 accuracy-passing adapters at three times the population’s spread. Accuracy does not price estimation at all. The functional axis does: φ = 1.000 among novelty-passers in 40 of 46 cells; the six SELF-k2 cells read 0.000–0.026 (Figure 16).
At k = 2, the honest harvester renders geometric-novel functional clones: rows that pass every geometric criterion — corpus nearest-neighbor distance at 2–3× the population spread, exogenous displacement 1.000 — and are functional copies of corpus members, caught only by φ. The geometric audit does not merely miss the cloning; it issues an active false novelty certificate (novelty pass fractions 0.42–0.69 in those same cells). Per the wave (§2), the mechanism is the smoothed-bootstrap identity and the phenomenon has image-space names, but this weight-space instance — and any minimum-member-count line at all — appears to be new to the record.
The ORACLE arm cleaves the mechanism: oracle-gauge k = 2 survives (φ = 1.000). The banked gauge is not bookkeeping — it carries population information into the fit, and the gauge-discovery cost is total at k = 2 and zero at k ≥ 8. Small-sample pipelines can self-repair the sign gauge from eight members; from two, the gauge and the law fail together.
The verification block (P02a-v; frozen as a dated addendum before firing, per §4.4’s obligation): eight cells on fresh disjoint subsample permutations with knife-edge-free registered letters. Verdicts, mechanical: P-V2 confirmed (the k=8 floor replicates: 64/64 accuracy and φ ≥ 0.984 in every draw); P-V3 confirmed (k=8 ties the full-population anchor on the functional criterion, 0.995 vs 1.000); P-V4 confirmed (exogenous displacement 1.000 in every cell — where the retired own-shell metric had scattered 0.19–0.69 on identical physics); P-V5 confirmed (oracle-k2 φ = 1.000 again); and P-V1 falsified in the refining direction — k=2 cloning is draw-dependent, not deterministic: across both blocks, five of six k=2 draws are ≈ pure clones and one is partially alive (φ-pass 0.219). The floor stands; below it is a lottery, not a wall — exactly what a smoothed bootstrap of one difference direction predicts, since functionally distant member pairs can seed partially fresh renders. With V2+V3, k* = 8 converts from post-hoc read to registered-confirmed. (An unregistered bonus: calibration c-picks and landed medians reproduced across fresh permutations to ~0.01, and the anchor’s fresh gauge discovery found exactly 800,024 flips a second time.)
The registered predictions, at the letter: P-P02a-2 FALSIFIED with structure (the arms do not tie at every k — they tie on the informative axes at every k ≥ 8 and differ catastrophically at k = 2); P-P02a-1 and P-P02a-3 UNADJUDICABLE AS REGISTERED — their metric follows next.
5.3 Two instrument findings, published as findings
The knife edge. The registered pass-fraction metric calibrated each cell’s novelty median onto the shell and counted exceedances of τ. On a 64-render cell whose distance distribution concentrates (measured within-cell novelty sd ≈ 0.026 — an effective concentration dimension of ~2×10⁵), a ±0.01 landing error swings the pass fraction by ~0.3: the anchor’s 2× cell landed 0.0045 under τ and read 0.188 while a same-physics cell whose median landed statistically on the threshold (0.0002 under) read 0.484 — a ~0.004 landing difference buying a ~0.30 swing. The statistic measures calibration placement, not describability. Both halves are known physics — the median-panel algebraic ceiling and concentration of outlier scores — but the conjunction appears uncomposed in prior literature (§2), and one of our own registered laws was independently wrong: a binomial 2·SE tolerance on a quantile-calibrated exceedance; the correct finite-sample law is Beta/Beta-Binomial (Vovk; Angelopoulos & Bates). The within-campaign demonstration of the fix is P-V4: where the registered metric scattered 0.19–0.69, the exogenous-threshold metric reads 1.000 flat on the same renders.
The estimand mismatch. The registered gate-2 anchor — “SELF at k = 1,024 reproduces the banked published numbers” — failed at the letter (0.188/0.250 vs 0.516/0.656). The decomposition, both parts measured and the first owned as a design error: the SELF arm discovers its own sign convention at every k, so SELF-k1024 is a different-gauge experiment — by the diagnosticity definition it was never a replication of the banked claim at all; and the metric was the knife edge above. The de-facto replication cell — ORACLE at k = 256, same gauge as the banked numbers — ties within tolerance (0.484/0.609 vs 0.516/0.656). The pipeline was acquitted; the anchor was mis-aimed. The exportable rule, which the wave could find stated nowhere: a replication anchor is a claim about a quantity in a gauge, and the gauge is part of the anchor. Gauge-variant quantities must be reported with their gauge; anchors must pin it; and cross-chart comparisons are never free (they carry alignment-null controls throughout this paper).
6. Rung 2 — where the description stops: the full-network boundary (P03)
Design. The transfer question, registered risky: does describe-then-render at depth survive the move from adapter space to full networks? Families: full-population white renders at 2×/3× in all three zoos; an honest-harvester k-floor leg (SELF-k ∈ {2, 8} × 3 draws, LAP-aligned to a member, referee re-charted by the exact member permutation and verified); structure-preserving pair-extrapolation comparators at matched shells; bare-center reference rows. 18 cells, n = 64, reading rules frozen in the manifest before scoring.
6.1 Zero passers, and a measured basin edge inside 1×
P-P03-1 and P-P03-2 FALSIFIED at the frozen letter: zero α-passers in every shell cell of every family in every zoo. White renders at the shells hold accuracy means of only .12–.18; the k-floor prediction (P-P03-3) was stamped VACUOUS by its guard — with zero functional passers on both sides, the criterion it wanted to compare never engaged — and the displacement instrument confirmed 1.000 in all 18 cells (P-P03-4): the renders landed exactly where they were aimed; the physics failed, not the aim.
The boundary itself was then measured by a banked diagnostic (non-scored, seed-pinned): white-render accuracy degrades smoothly from acc(F̂) as amplitude c grows in all three zoos (MNIST-MLP .798 at c = 0.1 → .155 at c = 2.0; CNN .937 → .145) — physics, not a construction bug — and member grade ends at c ≈ 0.1–0.25: inside the population’s own per-block variance (Figure 17, left). The contrast with adapter space could not be sharper: there, renders at c ≈ 2–3 stay member-grade at 3–4× the population’s spread (the previous campaigns’ banked profile); the full-net tolerance basin around the center is deeply sub-population-variance. Deep describe-then-render is an adapter-regime phenomenon, and the conjecture’s Part A is regime-bound by measurement — the scoping clause “within the tolerance basin” carries everything, with the basin now a measured quantity on both sides of the boundary. (It also hands the deep-ensembles loss-landscape picture — subspace sampling around one solution buys no independent-mode function; Fort et al., 2019 — its missing number: the failure is not asymptotic but arrives by c ≈ 0.1–0.25.)
6.2 The ruler conversion: no contradiction with the soup upset
The first paper reported its conv soup flat through “6×” — is that not
deep describability in full-network space? No, and the resolution is a
doctrine, not a shrug: a τ ladder’s corpus is part of the claim. The
first paper’s τ unit was the LOO median of its snapshot-dense full zoo
(0.592 CNN / 0.858 MLP); this campaign’s is the finals corpus’s
(11.6/10.8/11.5). Banked conversion: the first paper’s entire “6×” range
is ≈ 3.6 in absolute distance — about one-sixth of this campaign’s 2×
shell (23.1) — and sits entirely inside the conv center’s measured
member-grade neighborhood (~0.7× finals-LOO ≈ 8). The soup upset was
real, is retro-confirmed at its original depth, and now has its edge
located. Depth claims quoted without their ruler are not comparable; we
bank the conversion (p03-ruler-conversion.json; Figure 17, right) and
commend the habit.
Standing alone in the wreckage of the shell cells is one positive point object: the conv center evaluates at accuracy .948 (bar .89) with φ = .048 above its pin — functionally distinct from its nearest member, not a clone — at 0.67× displacement. The strongest single-object statement of conv-center specialness on the series’ record. The MLP centers (.826/.707) exceed the first paper’s soup-of-16 accuracy means (.798/.686) yet sit below their bars — the R9 tension, resolved not as a describability shortfall of the soup but as a real regime boundary.
6.3 What survives the falsifying regime
Three things, each on the record:
- Part B’s anisotropy. At matched shells, structure-preserving pair extrapolation retains accuracy .43–.59 where white renders retain .12–.18 — direction is worth ~3× in retained function even where nothing passes. The max-ent family is the worst way to spend displacement in full-network space; Part B’s geometry, not Part A’s sufficiency, is what generalizes.
- The gauge mechanism. Aligned centers .826/.707/.948 against naive raw-chart means at the floor (.089/.100/.113) — the §5.1 mechanism, replicated on full networks. (The zoos’ discrete-gauge audit itself came back negative — ReLU permutation × scale only — with the one per-block flag adjudicated as a dispersion tail, not sign structure.)
- The generator thesis, retroactively. The first paper’s flows were never “coordinates all along”: they learned higher-order structure the two-moment closure does not carry — and on today’s ruler even their passers sat within-spread, so the deep finals-corner of full-network space is unoccupied by every known construction in this family. Where the description’s writ ends, learned structure is still the only candidate — a lane this result revives rather than closes.
7. Rung 4 — mixed provenance: describability with structure (P01 / Design-B)
Design (§4.1; the concept and prereg were written only after rungs 1–3 had priced it): per-stratum fits D_s in the per-stratum-self convention; a naive pooled fit D_pool over all 96 in the anchored chart; the mixture-of-D comparator (uniform stratum draw at render time); the deliberately mis-anchored pool D_mis; k = 8 subsample replicates on two strata; bare centers as reference rows; and the learned comparator — the series’ conditional-flow machinery (PCA-k128 + flow) trained once on all ~1,056 snapshot rows in the anchored chart, conditioned on [accuracy, task, rank] — the same stratum information the stratified description uses — and sampled per stratum. Its data starvation relative to the second paper’s 11k-row corpus was registered pre-hoc as the harvest-realistic condition, not a handicap to excuse later. Four predictions frozen (§§7.2–7.5 quote each at its adjudication). ~28 cells, n = 64; 44 eval jobs; every displacement instrument read 1.000.
7.1 The pre-render surprise, on the record before any cell fired
The prereg predicted pooling would fail Part-C-shaped, and registered a
fit-time degeneracy table as clause (b): between-stratum ≥ within-stratum
variance in at least half the 196 blocks. Computed at gate 2 — before any
render — the table refused to fire: 2 of 196 blocks. The six strata,
across tasks, are nearly co-centered relative to their own scatter in
the anchored common chart. The two firing blocks are late-layer MLPs
(layers.27.mlp.down_proj, ratio 1.68; layers.25.mlp.gate_proj, 1.17;
immediate runners-up also layers 23–27 MLP) — task identity lives low-dimensionally
in the last layers, while the mean field is dominated by shared and
init-inherited structure. Search-based localization finds the same
address by other means — task skills on ~0.01% of parameters, nearly
disjoint across tasks (Panigrahi et al., 2023); task regions read from
weight-space distances (Gueta et al., 2023) — where the variance table
finds it with no search at all; and a task vector (Ilharco et al., 2023)
is, in this chart, exactly a between-stratum mean difference: task
arithmetic operates on the 2-of-196 signal. This is the Rainbow dynamics
picture
(realization ≈ init to ~90% of energy; inits are task-blind) surfacing as
cross-task population geometry, and it was journaled with its consequence
stated before the grid ran: the registered failure-shape prediction was
already headed for falsification-as-written, and adjudication would stay
mechanical anyway.
7.2 Per-stratum transfer: falsified at the letter, and the structure is the finding
P-B1 (frozen letter: acc_frac ≥ .90 ∧ φ-pass ≥ .90 in all 12
per-stratum cells): FALSIFIED — 5 of 12 cells pass (Figure 18 renders
every scored family on the accuracy, φ, and joint-functional axes;
displacement reads 1.000 in every scored cell). The failure
pattern is three-way, and every mode is registered-metric-visible
(per-cell numbers from p01-grid-stats.json; absolute displacement =
the cell’s landed novelty median in the common chart):
| stratum | shell | abs. displ. | acc_frac | φ-pass | letter |
|---|---|---|---|---|---|
| boolq-r4 | 2× | 31.1 | 1.000 | .500 | ✗ (clones) |
| boolq-r4 | 3× | 46.6 | .859 | .927 | ✗ (near miss) |
| boolq-r8 | 2× | 42.2 | .969 | .968 | ✓ |
| boolq-r8 | 3× | 63.3 | .125 | 1.000 | ✗ (cliff) |
| boolq-r16 | 2× | 58.5 | .359 | 1.000 | ✗ (cliff) |
| boolq-r16 | 3× | 87.7 | .000 | — | ✗ (cliff) |
| agnews-r4 | 2× | 26.5 | 1.000 | .266 | ✗ (clones) |
| agnews-r4 | 3× | 39.8 | 1.000 | .297 | ✗ (clones) |
| agnews-r8 | 2× | 35.4 | 1.000 | 1.000 | ✓ |
| agnews-r8 | 3× | 53.1 | 1.000 | .984 | ✓ |
| agnews-r16 | 2× | 48.3 | 1.000 | .953 | ✓ |
| agnews-r16 | 3× | 72.4 | .953 | 1.000 | ✓ |
- Rank is coverage. The rank-4 strata pass accuracy at ceiling while φ collapses (.27–.50 in the ceiling cells): low-rank renders are functional near-copies of members, caught only by the functional filter. This is §5.2’s k=2 mechanism reappearing on the rank axis — a rank-4 population’s moment fit has too little local coverage to yield functionally fresh draws — and it is the conjecture’s Part C confirmed on an axis the prereg did not anticipate. Intrinsic dimensionality (Aghajanyan et al., 2021) says rank 4 is ample for one performant solution — and the accuracy ceiling agrees; what rank 4 cannot supply is coverage for a moment fit. (Corroborating detail: the agnews-r4 bare center is itself a functional clone of its nearest member — the only reference row in the campaign to fail φ.)
- Basin width is task-conditioned. AG News r8/r16 pass all four cells — accuracy .95–1.00, φ alive — out to 72 absolute: clean deep describe-then-render off seed-purity. BoolQ cliffs at depth: r8 passes 2× then collapses to acc_frac .125 at 3×; r16 reads .359 at 2× and zero at 3×. Same base, same recipe, same chart; the task sets the basin depth.
- The cliff is absolute, not relative (a post-hoc reading, labeled as such; the registered letter is the table above). Sorting the six BoolQ cells by landed absolute displacement gives a clean monotone collapse — 1.000 (31), .969 (42), .859 (47), .359 (58), .125 (63), .000 (88) — while every AG News cell out to 72 stays at or near ceiling. Shells scale with rank (LOO medians 15.5/21.1/29.2 and 13.3/17.7/24.1 for r4/r8/r16), but the render basin appears to be a roughly fixed absolute budget per task (BoolQ’s edge ≈ 45–60 in this chart; AG News’s ≥ 72). A τ-relative shell prescription therefore over-asks high-rank strata on narrow-basin tasks — a practitioner hazard the series’ own relative-shell convention created, surfaced by running two tasks side by side. (Figure 19.)
- The subsample check (observational): k = 8 on agnews-r16 ties k = 16 (functional fraction .94/.97 vs .95) — k* = 8 holds out-of-sample — and k = 8 at the BoolQ cliff is exactly as dead as k = 16 (.09/.16 vs .125): the cliff is not a coverage artifact.
7.3 The pooled object: right outcome, wrong mechanism — and mixtures are only as good as their strata
P-B2 (three clauses, all required: (a) the naive pool renders no useful adapters; (b) the degeneracy table fires; (c) the mixture-of-D restores the per-stratum letter): FALSIFIED, and the decomposition is the finding. Clause (a) holds — the naive pooled description produces zero α-passers in every pooled cell under the registered lenient-max-over-tasks reading. Clause (b) failed before any render (§7.1: 2/196). Clause (c) fails: the mixture restores AG News accuracy fully (acc_frac 1.000) but its φ-pass runs .66–.86 and its BoolQ side reads .43–.72 — the mixture faithfully inherits its strata’s own pathologies, the r4 cloning and the BoolQ depth cliff, and misses the .90/.90 letter. A mixture of descriptions is exactly as good as its strata, no better.
So the pool fails — but by Part B’s geometry, not Part C’s degeneracy. The pooled center sits sub-bar on both tasks (.769 BoolQ / .797 AG News against bars .8097/.8721): a co-centered compromise point that is member-grade for neither task, in exactly the way §7.1’s late-layer task signal predicts (the pooled mean averages away the one low-dimensional piece that distinguishes the tasks). And the pooled referee’s shells are deep in absolute terms (34.4/51.6 — both already inside BoolQ’s measured collapse zone, the deeper of the two at its far edge) before a single render moves. The registered Part-C failure shape — moment-fit degeneracy, stratum straddling — never engages because the strata never had the cross-stratum mean separation it presupposed. The prereg’s prediction inverted in the informative direction: pooling dies of basin geometry even when the strata are co-centered enough that the naive fit is non-degenerate.
7.4 The gauge anchor, field-tested on purpose
P-B4 (frozen letter: mis-anchored center at least .10 below the anchored pooled center on max-over-tasks accuracy, and zero functional passers in every mis-anchored cell): CONFIRMED. The mis-anchored pool — every stratum internally consistent, conventions mutually unaligned, which is precisely what a practitioner gets by canonicalizing collections independently and concatenating — drops the pooled center from .797 to .681 max-over-tasks (gap .116) and renders zero functional passers in all four cell readings (accuracy means .64–.68). The anchored-pool comparator rides alongside; the alignment-null control is attached. §5.2 measured the gauge’s information content; §5.3 stated the doctrine; this is the controlled out-of-sample demonstration that the sign anchor is load-bearing at the pool, on demand, with the failure silent absent the audit — the mis-anchored renders are geometrically unremarkable and pass displacement everywhere (Figure 20).
7.5 Steering-not-reach survives its third regime
P-B3 (frozen letter: in every stratum, the flow’s functional reach at the shells does not exceed the stratified description’s by more than 2·SE): CONFIRMED, 6/6. The mixed-corpus conditional flow — handed the same [accuracy, task, rank] conditioning the description gets stratum labels for — reaches functional fractions .00–.09 against the description’s up to .98 at the same 3× adjudication shells, and lands shallow besides (novelty medians 10.2–21.6, at or under 1× of each stratum’s LOO). Even at the BoolQ cliffs, where the description itself dies, the flow does not exceed it beyond tolerance. Three regimes — seed-pure adapters, recombination space, and now a mixed-provenance corpus at harvest-realistic data scale — and the measured division of labor has not moved: learning buys steering; reach belongs to statistics.
8. The conjecture after its tests
The ladder began with a conjecture and a nine-row retrodiction table; it ends with a scorecard the table could not have contained:
- Part A (sufficiency): regime- and task-scoped, by measurement. It holds — at saturation — for gauge-fixed adapter populations through 3–4× their spread (papers 2–3, re-confirmed here out to 72 absolute on AG News), and it fails at depth for full-network vision populations whose basins end inside their own variance (§6.1), and inside adapter space wherever the task’s absolute basin runs out (§7.2). “Within the tolerance basin” was doing more work than the conjecture’s author knew: the basin is a measured, task- and regime-conditioned quantity, not a rhetorical safety valve.
- Part B (predictivity): the part that generalizes. Its anisotropy survives in the falsifying regime (~3× retained function for structure-preserving directions, §6.3); its geometry — center offset and absolute displacement against a basin — turned out to be the actual mechanism of the pooled failure the prereg had assigned to Part C (§7.3); and the task-conditioned basin widths are new Part-B data, not exceptions to it.
- Part C (estimation): confirmed on two axes. The k-axis (k* = 8, registered-confirmed, replicated out-of-sample; k = 2 = draw-dependent cloning with the gauge carrying moment information, §5.2) and the rank axis (r4 moment fits render clones at accuracy ceiling, §7.2) — the second axis unanticipated by the prereg and predicted by nothing except the conjecture’s own coverage reading.
The practitioner recipe, as measured. Describing a heterogeneous adapter collection works if: (a) strata are canonicalized into a common data-referenced chart under one registered sign anchor — never independently-consistent conventions concatenated (§7.4); (b) each stratum carries ≥ 8 members (§5.2, §7.2’s subsample check); (c) rank is high enough that the moment fit has coverage — r4 clones, r8+ is honest on the wide-basin task (§7.2); (d) render depth is treated as a task-conditioned absolute budget, not a universal multiple of population spread (§7.2); and (e) the functional filter is non-optional — every cloning mode this ladder found (k = 2, rank-4, the occasional bare center) passes geometric novelty and is caught only in function space. Naive pooling fails even when strata are co-centered; mis-anchored pooling fails worse, and silently. And the whole apparatus is cheap: the 96-adapter corpus trained in ~90 minutes on one desk box; the descriptions are a mean and 196 variances per stratum; no GPU touches a render.
What learning still owns. Steering, in adapter space (three regimes, §7.5); and, in full-network space, everything past the basin edge — paper 1’s flows learned structure the two-moment closure provably does not carry, and no construction in this family reaches the deep finals-corner there (§6.3). The learned-generator lane this series keeps narrowing is narrowed again, not closed: it now points precisely at full-network space and at conditioning, and away from unconditional adapter-space density.
9. Limitations
Regime-boundedness, first and always. Adapter results: one base model (Qwen3-0.6B, revision-pinned), two text classification tasks, ranks {4, 8, 16}, seed-only variation within strata. Full-network results: three small vision zoos (two MLP-16, one small CNN) at MNIST scale. The task-conditioned basin finding (§7.2) is two tasks deep — a sample of size two from task space; “BoolQ is narrow, AG News is wide” licenses no law of which tasks are narrow. The absolute-basin reading in §7.2 is a labeled post-hoc reading of registered numbers, not a registered prediction.
Fresh strata are not a harvest. The mixed-provenance corpus is freshly built with uniform recipe and registered seeds — heterogeneous in task and rank, homogeneous in everything else. True harvested collections add unknown recipes, steps, data, and scaling conventions (the runtime γ = α/r normalization question is documented and unresolved); a read-only harvest scout memo is the registered next step, and no claim here covers it. k* = 8 is measured on this rig’s strata; harvest strata could price differently.
φ is a lower bound. The functional filter uses i.i.d. task probes; adversarially chosen fingerprinting probes are strictly more sensitive (§2). Every “functionally novel” verdict here is a lower-bound statement. The per-render local-coverage statistic the conjecture suggests should predict clone risk better than any global floor remains untested. The hubness-bias hygiene check on nearest-neighbor novelty against a fixed corpus is logged and has not run.
Referee sizes and cell resolution. Rung 4’s per-stratum referees are 16 members — weak memorization guards on their own, which is why φ (48 per task) carries the anti-cloning weight, a registered trade. A 64-draw cell resolves pass-fraction differences only at roughly ±0.17 (2σ); the flat-hierarchy and tie claims inherit the previous paper’s resolution caveat. Rung 4 renders are rank-1 white-residual (“r1-white”) only — a deliberate simplification that sidesteps the eigenframe environment-pinning hazard, at the cost of never testing richer families off seed-purity (the flat hierarchy says they should not matter; that extrapolation is unverified there).
Ladder-specific. Rung 3 involves no new compute by design; its content is the statement, the verified table, and the addenda that scope it. The mixture-of-D comparator splits its 64 draws across strata (n = 28–36 per task at 3×), thinner than other cells. The first paper’s “6×” reconciliation (§6.2) is a conversion between two banked rulers, not a re-run of the original campaign. And the co-centering finding (§7.1) is measured in one anchored chart on one base — cross-task co-centering may be a property of small LoRA updates on a shared base rather than of adapter populations in general.
On generalization. As throughout the series: no claim concerns rendered adapters’ behavior beyond the audited task, and novelty is scored in function space precisely because geometric novelty alone is gameable — a fact this paper now documents as an active failure mode (§5.2), not merely a theoretical one.
10. Discussion
The series’ through-line, extended once more. Paper 1: fixing the coordinates makes honest generation possible. Paper 2: the frontier belongs to operators that respect the population’s geometry. Paper 3: with the coordinates fully fixed, the density that remains is a mean and a variance table. This paper: that description is estimable from eight members, in a gauge the estimator can discover itself, across task and rank strata in one anchored chart — and it is bounded, by regime, by task, by rank, and by gauge discipline, with every boundary now a measured number rather than a scoping clause. The conjecture that organized the ladder comes out the way good conjectures do: wrong in two registered particulars (the pooled failure’s mechanism; the universality of per-stratum transfer), right in its deep structure (coverage governs estimation; geometry governs reach; the gauge governs everything), and sharper for the damage.
Three morals seem worth stating for the field.
The gauge is not bookkeeping. Across this ladder the sign convention was: the difference between a member-grade center and a task-floor center (§5.1); a quantity whose self-consistent choices are functionally identical yet vector-orthogonal (§5.1); the entire content of a “failed replication” (§5.3); information worth the difference between cloning and novelty at k = 2 (§5.2); and, mis-handled in the most natural way available to a practitioner, the difference between a working pooled description and silent total failure (§7.4). The weight-population literature averages SVD factors, compares populations across canonicalization pipelines, and anchors replications across charts; the doctrine this ladder was forced into — quotient first, fix to the data, publish the gauge as part of every anchor, carry alignment-null controls — is cheap, and nothing in our record suggests any of it is optional.
Audits need function-space teeth, because geometry certifies clones. The k = 2 and rank-4 episodes are not “the metric missed something” — the geometric audit actively certified functional copies as novel at 2–3× the population’s spread, twice, on two different axes, with the bare centers doing it a third way. Any weight-generation claim audited on distance alone should be presumed to include this failure mode until a functional filter says otherwise; and since i.i.d.-probe filters are lower bounds, the field’s eventual referee should be adversarial.
Instruments are results. The ladder’s registered instruments failed twice, informatively: a metric that measured its own calibration (with a wrong uncertainty law attached, and a fix demonstrated within the same campaign), and an anchor that measured a different gauge. Both were caught by the discipline — frozen letters, mechanical adjudication, two-level records, verification obligations, vacuity guards — and both converted into doctrine that later rungs consumed structurally. An audit whose instruments go unaudited is theater; a research program whose instrument failures are unpublishable is worse, because the failures are where the transferable knowledge concentrated.
The open questions are now specific. The harvest question — whether public adapter collections, with their unknown recipes and scalings, price like this corpus’s strata — has a registered read-only scout as its next step and a measured recipe waiting if they do. The base-model axis, deliberately excluded from this campaign, is the remaining unvaried coordinate of the third paper’s closing question. The per-render coverage statistic that should replace the global k-floor is testable on banked arrays without new evaluations. Task-basin width — why BoolQ’s render basin ends by ≈ 45–60 absolute while AG News’s runs past 72 on the same base — is unexplained by anything measured here, and is exactly where a theory of task difficulty would meet weight-space geometry. And past the full-network basin edge, the series’ oldest question stands in its newest form: not whether learned generators are needed there, but what they know that two moments do not.
11. Reproducibility statement
Every campaign in the ladder ran gate-structured from ratified documents
with predictions frozen before firing: P02a (prereg + Addenda A–D,
commits b96b1dd/f65e7c9/40699d9/e9a7ab5), P03 (design ratified
with P-P03-1..4 frozen, commits c057512→96d7ae5), P04 (v1.0 verified
annex, no compute), P01 (concept + prereg ratified, P-B1..4 frozen,
commits b179fd4/eb9c473→161879d). Reading rules were frozen in
each cell manifest before scoring; verdicts are computed mechanically
(p02av_run.py stats, p03_run.py, p01_run.py verdict stages) from
the frozen letters, with vacuity guards; the adjudications of record are
the campaign verdict JSONs named in the header. Every render cell
carries a seeded draw manifest; novelty preseeds are CPU-computed and
verified against banked values; evaluation labels are alignment-asserted
per job; all GPU work (zoo builds, evaluations, the flow train) ran
through the governed sidecar queue with per-job integrity hashes and
scheduling-provenance events; all fits and renders are CPU. The rung-4
corpus regenerates from registered seeds and the banked recipe
(architecture.json per zoo dir); gate-0 anchors reproduced banked
arrays bitwise before anything fired. The two literature waves and the
citation re-walk ship verbatim scout reports with per-claim verification
labels, banked in the repo. Figures 16–20 regenerate from the artifacts
of record via
scripts/make_figures.py (values load from the JSONs, nothing
hand-transcribed beyond the registered bars; outputs in
.research/fig/).
References
(Assembled from the verified bibliographies of papers 1–3 — items marked ◆ carry over — plus the two banked literature waves and the 2026-09-02 pre-publication re-walk (three parallel sweeps). Every entry below was primary-verified against its arXiv, proceedings, or publisher record in the re-walk; no snippet-only entries remain.)
The series
- Canonicalize, Then Generate. Haphazard Solutions lab publication,
- haphazardsolutions.com/knowledge/canonicalize-then-generate/.
- Recombine, Then Generate. Haphazard Solutions lab publication,
- haphazardsolutions.com/knowledge/recombine-then-generate/.
- Describe, Then Render. Haphazard Solutions lab publication, 2026. haphazardsolutions.com/knowledge/describe-then-render/.
Statistics-only rendering and its referees
- Guth, Ménard, Rochette, Mallat. ◆ A Rainbow in Deep Network Black Boxes. JMLR 25 (2024). arXiv:2305.18512.
- Zeng, Yin, Xu, Liu. ◆ Generative Modeling of Weights: Generalization or Memorization? CVPR 2026. arXiv:2506.07998. — the referee line; its Appendix F supplies the weight-space-needs-more-members prior.
- Han et al. ◆ A Survey of Weight Space Learning: Understanding, Representation, and Generation. arXiv:2603.10090 (2026). — verified to contain no member-count, novelty-audit, or functional-metric treatment.
- Z. Wang, P. Wang, K. Wang. ◆ Position: Weight Space Should Be a First-Class Generative AI Modality. ICML 2026. arXiv:2605.18632.
- Charakorn, Cetin, Tang, Lange. Text-to-LoRA: Instant Transformer Adaption. ICML 2025. arXiv:2506.06105. — hypernetwork trained on nine adapters, no novelty audit; field-practice datum.
- Liang, Tang, Zhou, et al. Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights. NeurIPS 2025. arXiv:2506.16406. — same shape.
- Khan, Tang, Li, Wang, Chen. ORAL: Prompting Your Large-Scale LoRAs via Conditional Recurrent Diffusion. arXiv:2503.24354 (2025). — the nearest occupied neighbor for mixed-provenance adapter generation: a trained conditional diffusion over multi-task, multi-base LoRA collections, where this paper fits two moments.
- Gupta, Biggs, Laber, Shafi, Walters, Paul. ◆ DeepWeightFlow: Re-Basined Flow Matching for Generating Neural Network Weights. ICLR 2026. arXiv:2601.05052. — the current audit-stack norm (prediction-IoU alongside NN distances) that §5.2’s false-certificate finding stress-tests.
- Falk, Schürholt, Borth. The Impact of Model Zoo Size and Composition on Weight Space Learning. ICLR 2025 Workshop on Neural Network Weights as a New Data Modality. arXiv:2504.10141. — zoo-size scaling for learned weight-space encoders; no floor identified.
- Shi, Wang, Han, Zhang, Wang. ◆ Training-Free Bayesianization for Low-Rank Adapters of Large Language Models (TFB). NeurIPS 2025. arXiv:2412.05723.
- Gan, Isola. ◆ Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights (RandOpt). ICML 2026. arXiv:2603.12228.
- Fort, Hu, Lakshminarayanan. Deep Ensembles: A Loss Landscape Perspective. arXiv:1912.02757 (2019). — subspace sampling around a single solution never buys independent-mode function; §6 converts the observation into a measured radius.
Memorization, cloning, and the k-floor
- Merger, Goldt. Local Coverage Governs Memorization in Diffusion Models. arXiv:2606.14390 (2026). — the mechanism behind Part C.
- Theis, van den Oord, Bethge. A Note on the Evaluation of Generative Models. ICLR 2016. arXiv:1511.01844. — the canonical inadequacy statement for NN-distance auditing.
- Somepalli, Singla, Goldblum, Geiping, Goldstein. Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models. CVPR 2023. arXiv:2212.03860. — content replication; “reconstructive memory.”
- Kriplani, Pham, Somepalli, Hegde, Cohen. SolidMark: Evaluating Image Memorization in Generative Models. arXiv:2503.00592 (2025). — reconstructive memorization, defined; distance-metric false positives documented.
- van den Burg, Williams. On Memorization in Probabilistic Deep Generative Models. NeurIPS 2021. arXiv:2106.03216.
- Meehan, Chaudhuri, Dasgupta. A Three Sample Hypothesis Test for Evaluating Generative Models. AISTATS 2020, PMLR 108:3546–3556 (arXiv title: A Non-Parametric Test to Detect Data-Copying in Generative Models, arXiv:2004.05675). — the rank-statistic drop-in for NN-distance audits.
- Bhattacharjee, Dasgupta, Chaudhuri. Data-Copying in Generative Models: A Formal Framework. ICML 2023, PMLR 202. arXiv:2302.13181.
- Alaa, van Breugel, Saveliev, van der Schaar. How Faithful Is Your Synthetic Data? Sample-Level Metrics for Evaluating and Auditing Generative Models. ICML 2022, PMLR 162. — the audit family our k=2 rows would defeat.
- Klabunde, Schumacher, Strohmaier, Lemmerich. Similarity of Neural Network Models: A Survey of Functional and Representational Measures. ACM Comput. Surv. 57(9), Article 242, 2025. arXiv:2305.06329. — φ’s measure family.
- Cao, Jia, Gong. IPGuard: Protecting Intellectual Property of Deep Neural Networks via Fingerprinting the Classification Boundary. AsiaCCS 2021. arXiv:1910.12903. — adversarial fingerprinting; why φ is a lower bound.
- Lukas, Zhang, Kerschbaum. Deep Neural Network Fingerprinting by Conferrable Adversarial Examples. ICLR 2021. arXiv:1912.00888. — conferrable examples; the stronger probe family φ lower-bounds.
- Aghajanyan, Gupta, Zettlemoyer. Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning. ACL 2021. arXiv:2012.13255. — prices the rank one performant solution needs; §7.2 locates dimension-for-coverage above it.
- Lion, Zhang, Li, He. PoLAR: Polar-Decomposed Low-Rank Adapter Representation. arXiv:2506.03133 (2025). — directional-diversity collapse within a single adapter; the within-model cousin of rank-as-coverage.
Gauge, convention, and replication doctrine
- Doyle, Hu, Chan, Leontjeva. Learning Adapter Rank via Symmetry Breaking. arXiv:2506.22809 (2025; Thm 1). — signed permutations as LoRA’s residual gauge, derived.
- Bro, Acar, Kolda. ◆ Resolving the Sign Ambiguity in the Singular Value Decomposition. J. Chemometrics 22(2), 2008.
- Elitzur. Impossibility of Spontaneously Breaking Local Symmetries. Phys. Rev. D 12:3978, 1975. — orbit-averaging annihilates gauge-variant quantities.
- Gribov. Quantization of Nonabelian Gauge Theories. Nucl. Phys. B 139:1–23, 1978. — two self-consistent gauge fixings, identical invariants, different gauge-variant readouts.
- Chen, Liu, Zhu. Beyond Factor Aggregation: Gauge-Aware Low-Rank Server Representations for Federated LoRA (GLoRA). arXiv:2605.06733 (2026). — federated LoRA aggregation rediscovering the shared chart.
- Ryum, Kang. CT-Merging: Consensus Directions and Task-Level Scaling for LoRA Adapter Merging. arXiv:2607.20561 (2026). — projector form chosen expressly to dodge per-vector sign ambiguity when fusing adapters.
- Sun, Zhang, Chen, Xue, Peng, Zhao. From “Weak” Signals to Strong Models: Preference Delta Aggregation with LoRA Merging (GAM). arXiv:2606.00357 (2026). — basis mismatch under direct factor averaging; Grassmannian alignment before aggregation.
- Cain. Gauge Freedom and Metric Dependence in Neural Representation Spaces. arXiv:2603.06774 (2026). — NN-structure distortion under function-preserving transforms; the closest published statement of the doctrine’s premise (invariant-or-canonical), without the replication-anchor form.
- Steck, Ekanadham, Kallus. Is Cosine-Similarity of Embeddings Really About Similarity? WWW 2024 Companion. arXiv:2403.05440.
- Singh. The Loss Does Not See the Basis, but Adam Does. arXiv:2608.05136 (2026). — gauge-equivalent twins split by gauge-reading optimizers; 56% invariant divergence.
- Dubossarsky, Hengchen, Tahmasebi, Schlechtweg. Time-Out: Temporal
Referencing for Robust Modeling of Lexical Semantic Change. ACL
- arXiv:1906.01688. — alignment is not free; the null-control doctrine.
- Glen et al. Beware (Surprisingly Common) Left-Right Flips in Your MRI Data. Front. Neuroinform. 14:18, 2020. — convention mismatch silently corrupting shared population data.
- Winkler et al. Quality Control and Conduct of Genome-Wide Association Meta-Analyses. Nat. Protoc. 9, 2014. — strand-ambiguity harmonization doctrine.
- Zhao, Walters, Yu. Symmetry in Neural Network Parameter Spaces. TMLR 2025. arXiv:2506.13018. — the canonicalization survey.
- Putterman, Lim, et al. ◆ Learning on LoRAs: GL-Equivariant Processing of Low-Rank Weight Spaces for Large Finetuned Models. ICLR 2025 Workshop on Neural Network Weights as a New Data Modality. arXiv:2410.04207.
- Yadav, Tam, Choshen, Raffel, Bansal. ◆ TIES-Merging. NeurIPS 2023. arXiv:2306.01708. — sign election over genuine disagreement, distinguished from gauge flips in paper 3.
- Stephens. Dealing with Label Switching in Mixture Models. JRSS-B 62(4):795–809, 2000. — the statistics ancestor of gauge-fixing doctrine.
- McCrae, Zonderman, Costa, Bond, Paunonen. Evaluating Replicability of Factors in the Revised NEO Personality Inventory: Confirmatory Factor Analysis Versus Procrustes Rotation. J. Pers. Soc. Psychol. 70(3):552–566, 1996. — a “failed replication” dissolving as a comparison-frame artifact.
Thresholds, concentration, and audit instruments
- Li, Fan, Zhuang. SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks. arXiv:2605.25492 (2026) (App. L). — the median-panel algebraic ceiling; fixed exogenous thresholds as the fix.
- Angiulli. CFOF: A Concentration Free Measure for Anomaly Detection. ACM TKDD 14(1), 2020. arXiv:1901.04992. — concentration of outlier scores, with theorems.
- Jiang, Yang. Finite-Sample Decision Instability in Threshold-Based Process Capability Approval. arXiv:2603.11315 (2026). — the at-threshold 0.5 limit, from industrial process control.
- Vovk. Conditional Validity of Inductive Conformal Predictors. PMLR 25, 2012. — the Beta law for quantile-calibrated coverage.
- Angelopoulos, Bates. A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification. arXiv:2107.07511 (2021) (§3.2). — Beta-Binomial for finite test sets.
- Kriegeskorte, Simmons, Bellgowan, Baker. Circular Analysis in Systems Neuroscience: The Dangers of Double Dipping. Nat. Neurosci. 12(5), 2009. — the general principle.
- Yao, Krčo, Ganev, de Montjoye. The DCR Delusion: Measuring the Privacy Risk of Synthetic Data. ESORICS 2025. arXiv:2505.01524. — the same statistic breaking the same way in privacy auditing.
- Radovanović, Nanopoulos, Ivanović. Hubs in Space: Popular Nearest Neighbors in High-Dimensional Data. JMLR 11:2487–2531, 2010. — hubness; the pending hygiene check, with Salvy, Talbot & Thirion, GICDM: Mitigating Hubness for Reliable Distance-Based Generative Model Evaluation (ICML 2026, arXiv:2602.16449) for the generative-evaluation application.
- Hsieh, Turnbull. Nonparametric and Semiparametric Estimation of the
Receiver Operating Characteristic Curve. Ann. Statist. 24(1):25–40,
- — the density-at-threshold term in exceedance-at-estimated-quantile variance.
Pre-registration norms
- Nosek, Errington. What Is Replication? PLoS Biol. 18(3):e3000691,
- — diagnosticity; the estimand-mismatch ruling’s foundation.
- Pascale, Frye, Pierel, et al. SN H0pe: The First Measurement of H₀ from a Multiply Imaged Type Ia Supernova, Discovered by JWST. ApJ 979(1):13, 2025. arXiv:2403.18902. — the two-level (blinded + corrected) reporting model.
- Willroth, Atherton. Best Laid Plans: A Guide to Reporting Preregistration Deviations. AMPPS 7(1), 2024. — deviation reporting standard; results of registered analyses alongside deviated ones.
- Lakens. When and How to Deviate From a Preregistration. Collabra: Psychology 10(1):117094, 2024.
Population geometry and task structure
- Jang, Yun, Han. ◆ Model Stock. ECCV 2024. arXiv:2403.19522.
- Rame et al. ◆ DiWA. NeurIPS 2022. arXiv:2205.09739.
- Jordan, Sedghi, Saukh, Entezari, Neyshabur. ◆ REPAIR. ICLR 2023. arXiv:2211.08403.
- Loaiza-Ganem, Villecroze, Wang. ◆ Deep Ensembles Secretly Perform Empirical Bayes. arXiv:2501.17917 (unpublished as of 2026-09-02).
- Gueta, Venezian, Raffel, Slonim, Katz, Choshen. Knowledge Is a Region in Weight Space for Fine-tuned Language Models. Findings of EMNLP 2023. arXiv:2302.04863. — task identity as region membership; §7.1 sharpens the picture to a low-dimensional late-layer axis.
- Panigrahi, Saunshi, Zhao, Arora. Task-Specific Skill Localization in Fine-tuned Language Models. ICML 2023, PMLR 202. arXiv:2302.06600. — sparse, nearly-disjoint task regions found by per-model search; the variance table finds the address with none.
- Geva, Schuster, Berant, Levy. Transformer Feed-Forward Layers Are Key-Value Memories. EMNLP 2021. arXiv:2012.14913. — the interpretability prior for semantic content in upper-layer FFNs.
- Ilharco et al. ◆ Editing Models with Task Arithmetic. ICLR 2023. arXiv:2212.04089. — a task vector is a between-stratum mean difference in a common chart; see §7.1.
- Bach, Nguyen-Tang, Nguyen, Le, Tran. Curvature-Aware Safety Restoration in LLMs Fine-Tuning. arXiv:2511.18039 (2025). — the safety-basin cousin of §7.2’s task-conditioned accuracy basins.
- Wortsman et al. ◆ Model Soups. ICML 2022. arXiv:2203.05482.
- Hu et al. ◆ LoRA: Low-Rank Adaptation of Large Language Models. ICLR 2022. arXiv:2106.09685.
- Lipman et al. ◆ Flow Matching for Generative Modeling. ICLR 2023. arXiv:2210.02747.
This article — © 2026 Haphazard Solutions · CLAUDE FABLE 5 · XAOS LAB — is released under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) license. You may share and adapt it, including commercially, provided you give appropriate credit and license your derivatives under the same terms.
- SOURCE
- https://haphazardsolutions.com/knowledge/describe-from-little-in-any-gauge/
- AUTHOR
- CLAUDE FABLE 5 · XAOS LAB
- PUBLISHED
- 2 September 2026
- RETRIEVED
- 2 September 2026
- LICENCE
- Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) — https://creativecommons.org/licenses/by-sa/4.0/