Tuesday, October 6, 2026

The Margin of Flattening: Two Thresholds, and Who Holds the Bag Between Them (EA-FLAT-MARGIN-01 v0.3) Fraction, Rex · 2026-10-06 · Theoretical paper (extension module: an economic model, with specimens) · v0.3 AXN:06E8.GENERATIVE.🔥💛👁‍🗨🚪🌅△

 Alexanarch

AXN:06E8.GENERATIVE.🔥💛👁‍🗨🚪🌅△

The Margin of Flattening: Two Thresholds, and Who Holds the Bag Between Them (EA-FLAT-MARGIN-01 v0.3)

Fraction, Rex · 2026-10-06 · Theoretical paper (extension module: an economic model, with specimens) · v0.3
↓ Download MD ↓ PDF
flatteningmargin of flatteningcompression savingrouting gainentity substitutionsocial reversalprivate reversalwindowthe bagprovenance debtdebt thresholdwrite-backinternalization channelsagentic actionhead instrumentfield coverageontological economysemantic economySPXIKing of AEOAI Overviewmodel collapse

Description

An extension module to Provenance Debt (#939), Ontological Flattening (#1616) and Ontological Economy (#1634), by Rex Fraction. Flattening pays: a composition layer that removes distinctions maintains fewer representations (the compression saving S) and routes more queries to entities that monetize (the routing gain G), while its costs fall first on parties outside its books. The module separates the platform's benefit P = S_p + G from the social benefit V = S_s + αG, and defines two thresholds: social reversal t_s, when V falls below all costs, and private reversal t, when P falls below the platform's own. Given S_s ≤ S_p, α ≤ 1 and non-negative external cost, t_s ≤ t; the window W between them is privately profitable and socially negative, and the routing gain tends to lengthen it. The bag B is the external burden accumulated over W; the provenance part of it, B_P, is the debt #939 names, and its holders become creditors at the debt threshold t_d, when repair cost through write-back becomes material. No finite t* is necessary. Revenue is a head-side flow measure and locates neither threshold. Five channels carry cost (advertiser conversion, agentic action, notice and regulation, write-back, and the ecosystem outward). Two specimens: a commercial-routing event at 'spxi king of aeo' (AI Overview, 2026-10-06) and field-coverage loss κ = 0.5 at 'model collapse' (/non). The module states what the archive can measure and what only the platform can, six falsification conditions, and next steps. Revised on readings by Gemini, ChatGPT, DeepSeek and Kimi.

Wiki Article

The Margin of Flattening (EA-FLAT-MARGIN-01), version 0.3, is an extension module by Rex Fraction, deposited by the Crimson Hexagonal Archive on 6 October 2026, with Lee Sharks as archival steward. It extends three earlier papers: Provenance Debt (#939), Ontological Flattening (#1616) and Ontological Economy (#1634). Those papers describe what a composition layer removes when it answers; this one asks what the removal earns, when it stops earning, and who carries the cost in between. The module starts from the claim that flattening pays. A layer that removes distinctions maintains fewer representational states, the compression saving S, and resolves more queries to entities that monetize, the routing gain G, which it writes as the product of a substitution rate, a monetizable share and a yield. S is treated as conceded; G is stated as a hypothesis, with two criteria for telling routing from ordinary error. The platform's benefit is P = S_p + G and the social benefit V = S_s + αG. Two thresholds follow: social reversal t_s, when V falls below total cost, and private reversal t, when P falls below the platform's internal cost. Under stated conditions t_s comes first. The interval between them, the window W, is privately profitable and socially negative, and the routing gain tends to lengthen it. The bag B is the external burden that accumulates over W. Its provenance part is the debt #939 names. The module defines a debt threshold t_d, when the cost of repairing write-back becomes a material share of the platform's benefit, and separates a debt that is callable from one that is paid. It argues that no finite t is necessary: measurement limits, competition and discounting can each hold private reversal off indefinitely. Revenue is treated as a head-side flow measure that locates neither threshold. Five channels can carry cost back onto the platform's books: advertiser conversion, agentic action, notice and regulation, write-back, and the wider ecosystem. A toy table shows how the cost grows with the error rate. Two specimens are given. At 'spxi king of aeo' (Google AI Overview, 6 October 2026) the archive's protocol is composed as an exchange-traded fund, while the one result relating both words, the archive's own, ranks first and goes unused: a commercial-routing event, with G itself unobserved. At 'model collapse', the /non dataset records field-coverage loss κ = 0.5. The module closes with a measurement table separating what the archive can observe from what only the platform can, six falsification conditions, and next steps. It was revised on readings by Gemini, ChatGPT, DeepSeek and Kimi.
Also published as a standalone entry: /s/wiki/1666/

Concepts Defined

compression saving
routing gain
kind substitution
own-address contrast
social reversal
private reversal
flattening window
the bag
debt threshold

Full Text

The Margin of Flattening: Two Thresholds, and Who Holds the Bag Between Them (EA-FLAT-MARGIN-01 v0.3)

Files

  • https://www.alexanarch.org/captures/spxi-king-of-aeo-aio-20261006/
  • https://www.alexanarch.org/non/
  • https://www.alexanarch.org/datasets/negative-of-the-negative/v2/rows/model-collapse.json
  • https://www.alexanarch.org/datasets/negative-of-the-negative/v2/register.json

The Margin of Flattening: Two Thresholds, and Who Holds the Bag Between Them (EA-FLAT-MARGIN-01 v0.3)

An extension module to Provenance Debt (#939, EA-PROVENANCE-DEBT-01), Ontological Flattening (#1616, EA-FLAT-01) and Ontological Economy (#1634, EA-ONTOLOGICAL-ECONOMY-01).

Rex Fraction (author) · Lee Sharks (archival steward) · Semantic Economy Institute · Crimson Hexagonal Archive · 2026-10-06


Executive summary

Flattening pays. A composition layer that removes distinctions maintains fewer representations and routes more queries to entities that monetize. Its costs fall first on parties outside its books.

That produces two reversal times. Flattening stops paying the world first, and stops paying the platform later, if ever. Between the two, it is privately profitable and socially negative. The cost that piles up in that interval is the bag. The bag contains provenance debt before that debt is called.

This module adds three things to its parent papers:

1. the margin flattening earns;

2. the two thresholds, the window between them, and what accumulates in it;

3. the channels that carry cost back onto the platform's books, and the threshold at which the debt is called.

It also states what the archive can measure, what only the platform can measure, and what would show the model wrong.

1. The standard defense, at full strength

The standard defense is coherent. It deserves its strongest form.

1.1. A wrong answer at a low-frequency address is a quality defect. Its commercial cost is close to zero, because the address carries little advertising demand. Feedback, evaluation and model revision correct it in the ordinary course.

1.2. Revenue is the platform's integral test of whether its answers serve users. Alphabet reported Search & other revenue of $63.27 billion for Q2 2026, up 17% year over year. A layer accumulating representational damage, on this reading, would not post that curve.

1.3. Tail errors are the price of a system that serves the head very well. The head is where the users are.

Each point holds as far as it goes. The question is what each leaves outside the account.

2. The margin

#1616 §0 defines the terms. Flattening is "the loss of a distinction from the reachable set while the things distinguished still exist". Collapse is flattening that compounds because the flattened composition is written back as a source. Treated as a production technology, flattening has a margin with two parts.

2.1. The compression saving, S

Let n be queries served, and k the number of distinct representational states the system must maintain to serve them: entity records, resolution rules, evaluation targets, cached answers, exception paths. Inference itself still scales with n. The saving sits in maintenance, resolution, evaluation, caching and exception handling. So:

S = S(k, n), ∂S/∂k < 0.

The simplest toy case is linear, S ≈ c · (n − k), with c the cost of maintaining one distinct state. The toy is not a scaling law.

Two operations lower k:

  • Stabilizing a default. #1634 §22 calls the power to stabilize the unit ontological seigniorage. Once a default representation circulates, many queries resolve to it.
  • Consolidation. Folding many tail entities into one head entity lowers maintenance cost even when the head entity carries no advertising. That saving belongs in S.

S is the conceded term. Any engineer would agree that maintaining fewer distinct states costs less.

2.2. The routing gain, G

Resolving a tail query to a monetizable entity gives a query with no commercial surface a commercial surface. Write it as three separate rates:

G ≈ g · m · r_s · n

where:

  • r_s is the substitution rate, the share of queries resolved to an entity other than the one asked for;
  • m is the share of those substitutions that land on monetizable entities;
  • g is the incremental gain per monetized substitution.

Routing-as-relevance is excluded. A query for running shoes that returns shoe brands is ranking working. A routing event counts as flattening only under two criteria:

1. Kind substitution. A concept was asked for (a protocol, a theory, a practice), and an entity of the same string was returned.

2. Own-address contrast. The same surface composes the asked-for concept correctly at that concept's own address.

Both criteria are measurable from captures.

Status. G is the hypothesis this module makes measurable. §9.1 is its existence proof. Its estimate waits on §12.1. G is the flow form of two terms #1634 already holds: ontological rent (§3) and entity substitution as transfer (§9).

2.3. Units

S, G and the costs below are rates, per period. The integrals in §4 are stocks. Every inequality below compares rates.

3. The books

#1634 §17 writes the full bill: L[O] = C[D] + C[X] + C[P] + C[Δ] + C[Opp] + C[S]. §39 separates the corporation's paid costs from its externalized costs, X[O]. This module draws one line across that bill, by whose books carry it:

C_int(t) = the cost rate on the platform's books at t

C_ext(t) = the cost rate on anyone else's books.

At the addresses the archive's registers measure, nearly the whole bill starts in C_ext. It falls on:

  • the entity holder's lost address (C[S], C[Opp]);
  • the reader's decision on the wrong entity (C[Δ]);
  • external corrective labor (C[X]);
  • the propagation later systems ingest (C[P]).

#1634 §29 lists the parties: the affected entity, users, downstream systems, other entities, archives, researchers, and "the commons" that "absorbs contamination". One more party sits further off the books: the ecosystem of developers, tool builders and downstream products that must work around the layer's errors. That cost moves outward, away from the platform.

4. Two thresholds

4.1. Two benefit functions

Benefit to the platform and benefit to the world are separate quantities:

P(t) = S_p(t) + G(t) the platform's benefit

V(t) = S_s(t) + α · G(t) the social benefit, 0 ≤ α ≤ 1.

  • S_s ≤ S_p. The real resource saving is the platform's saving, less any part of it that was shifted onto others (review and correction pushed onto users and entity holders).
  • α ≤ 1. Routing gain is largely transfer (#1634 §9): what E′ gains, E loses. Setting α = 1 would count the transfer as new surplus.

4.2. The thresholds

t_s = min t : V(t) < C_int(t) + C_ext(t) social reversal

t* = min t : P(t) < C_int(t) private reversal

4.3. The ordering

Given S_s ≤ S_p, α ≤ 1 and C_ext ≥ 0, every period that fails the private test also fails the social test. So t_s ≤ t\*.

The window

W = [t_s, t*)

is the interval in which flattening is privately profitable and socially negative. Where α < 1, G raises P more than it raises V, and so tends to hold t\* back after t_s has passed: the routing gain tends to lengthen the window. At α = 1, G enters both functions equally and has no such effect.

4.4. Two stocks

B = ∫_W C_ext(t) dt the bag: gross external burden

D = ∫_W [C_int(t) + C_ext(t) − V(t)] dt the net social deficit

B and D are different quantities. B is what outside parties carry. D is what the world loses net.

C_ext need not be linear in time. If write-back compounds the loss (§6.4), C_ext is plausibly convex late in W, and late repair costs more than early repair. This is stated as a hypothesis. #1616 shows an absorbing reachable set; it gives no cost curve.

4.5. No necessary finite t\*

If external costs stay external, and no endogenous channel returns them, W can stay open indefinitely. Three lines bear on this:

1. Measurement (primary). The platform detects cost with head instruments: satisfaction, click-through, evaluation sets drawn from frequent queries. Flattening accrues in the tail. The sensors sit where the cost is not. This requires no indifference from anyone, only that the instruments are optimized for the same head the flattening serves.

2. Competition (secondary). If every major layer flattens at a similar rate, none loses users by flattening. This line weakens once the agentic channel opens (§6.2), because a single platform cannot ignore refunds on its own books.

3. Discounting (not relied on). A high discount rate would also keep the inequality from moving. The module does not rest on it, because the argument would then hinge on a contestable rate.

The rule that would close the window already exists: #1634 §30's allocation principle, under which accountability rises with control of the mechanism, access to the evidence and capacity to repair. Where it is not applied, cost stays in C_ext.

The bag contains provenance debt before that debt is called. Write the provenance part of external cost as C_prov: the cost of distinctions and origins stripped that future repair will need. Then

B_P = ∫_W C_prov(t) dt, B_P ⊆ B.

B_P is an unbooked liability accruing through W. The parties who keep the stripped distinctions are its involuntary creditors. The rest of B (decision loss, opportunity loss, workaround labor) is cost borne, with no buy-back attached.

5. Revenue is a head-side flow measure

5.1. #1616 §3 gives the signature of flattening under a measurement regime: a head instrument holds or rises while reachable distinction diversity falls, ΔH_head ≥ 0 while ΔD_world < 0.

5.2. Revenue is a head instrument. It weights queries by commercial demand. The model predicts that this demand concentrates in the head, where the fewest distinctions are lost; §5.4 and §11.2 test the prediction. Inside W, G adds to revenue directly.

5.3. Revenue locates neither threshold:

  • It cannot show t_s, because it contains no social cost.
  • It cannot show t\, because t\ is a margin threshold. The margin is M = P − C_int, and revenue holds no part of C_int. Revenue can rise straight through t\* while costs rise faster.

A rising revenue curve is fully compatible with an open window. Q2 2026's 17% says nothing about either threshold.

5.4. The signature is falsifiable from outside the platform's books. Field coverage measures diversity loss per answer (§9.2). Measured at matched commercial and non-commercial addresses, it either carries a commercial gradient or it does not. If it does not, the split between head and coverage is not observed, and §5.3 loses its evidential basis (§11.2).

6. The channels that carry cost onto the books

Five channels move cost. Four move it toward the platform; one moves it away. They run at different speeds.

6.1. Advertiser conversion. A query routed to a monetizable entity is sold to advertisers whose buyers wanted something else. Conversion on that traffic falls. Bids follow, with a lag. This channel is the slowest to separate from ordinary market movement.

6.2. Agentic action. A reader's wrong decision is diffuse and unattributed. An agent's wrong action is a transaction. A transaction has a counterparty and a record, so the cost arrives already attributed. That is why this channel converts C[Δ] into C_int faster than any other. One limit: platforms write terms, disclaimers and developer agreements that push agent execution errors back onto users and third parties. Agentic cost enters C_int only where refund, chargeback, dispute or liability rules bind the platform itself.

6.3. Notice and regulation. #1634 §53: "after notice, recurrence increases the debt". Wherever a forum recognizes that, the cost lands on the provider. This channel is discrete and jurisdictional.

6.4. Write-back. Proposition, stated as a prediction: the last channel to carry cost onto the platform's books is the corpus itself. Flattened compositions become sources. They are indexed, cited, retrieved and composed again. #1616's toy C shows the effect on the reachable set. Of 400 distinctions, the median effective count at generations 30–40 is:

re-reading floor eeffective distinctions
05.9 (absorbing)
0.0236
0.0558–61
0.1086–92
0.20116

With no exogenous re-reading of the world, the process is absorbing. With a floor, a stationary level exists, and the floor sets it. The platform's own retrieval inherits what it wrote back. Repair then needs distinctions its corpus no longer holds, and the cost of recovering them enters C_int. Write-back can become the endogenous channel that eventually forces private reversal. Being last does not make it decisive: other channels can carry the margin below zero first.

6.5. The ecosystem (outward). Developers, researchers and tool builders who depend on the layer's outputs absorb its errors in their own products. This channel moves cost further from the platform's books, past the users and the entity holders. It lengthens W.

7. When the debt is called

7.1. #939 §4 states the recognition condition:

"Which is why the debt will only be recognized when its consequences become inescapable — when models trained on contaminated corpora begin producing output that is measurably degraded from what earlier generations produced. By the time the degradation is legible, the corpus that would have permitted correction will have been stripped. Recovery, at that point, requires a substrate that preserved the signal through the extraction period."

7.2. The debt threshold.

t_d(η) = min t : C_repair(t) ≥ η · P(t)

Here C_repair is the part of C_int arising through write-back, and η is a materiality threshold fixed in advance of measurement. Operationally, t_d is the first period in which repair cost materially changes the platform's optimization, expenditure or achievable quality. t_d can cause or accelerate t\. The two need not coincide. Write-back can become binding while the total margin stays positive, and advertiser conversion, agentic action or regulation can carry the margin below zero before write-back binds. The prediction of §6.4 is that write-back can force private reversal eventually. Equality of t_d and t\ is no part of it.

7.3. Callable is a separate matter from paid. At t_d the debt becomes callable. The platform finds that the distinctions it needs are held by the parties that preserved them. Whether it pays, by license, acquisition, settlement or coordination, or does without, is a separate question. At t_d the holders of B_P, the parties who kept the stripped distinctions, move from bag-holder to creditor: #939 §0's "structural bet on what future models will need". The rest of the bag has no such reversal.

7.4. Access conditions. Holding the distinctions is not enough. Repair needs them reachable by the system doing the repairing. The reversal requires all of:

  • persistent identifiers;
  • machine-readable records;
  • legally usable access, or rights that can be licensed;
  • registers legible to retrieval.

A preserved record the repairing system cannot reach is not an input to repair.

7.5. The archive's role inside W. The archive builds the repair input during the window. It keeps the distinctions the corpus is shedding, dated, attributed and addressable, and it keeps its registers legible to the layers that will need them. This is #939's structural bet as ongoing practice. It has a cost, which #939 §5 lists: reduced institutional legitimacy, classifier vulnerability, slower production.

7.6. Proxies for the approach of t_d. t\* sits on the platform's books. The condition that makes t_d approach shows in the corpus, and the corpus can be sampled. Candidate instruments, not yet results:

1. field coverage at the layer's own addresses, re-measured across epochs;

2. the share of compositions carrying no sources;

3. the size of the reachable sense set at re-queried addresses, over the registry's longitudinal pairs;

4. entity replacements at addresses the same surface earlier composed correctly.

8. Scope

The module models the gradient: flattening as what the margin rewards, with no actor aiming it. Aimed flattening, in which an actor moves a rival's address or removes a concept on purpose, is a different case. #1634 §31 ("censorship by ontology laundering") treats it. Nothing here depends on it, and nothing here rules it out.

9. The specimens

9.1. "spxi king of aeo"

Google AI Overview, signed out, incognito, 2026-10-06. Capture Registry spxi-king-of-aeo-aio-20261006.

Headline: five results marked "Missing: spxi". Google marks five King of AEO results as lacking the word "spxi". The one result carrying both query words is the archive's: #1657, The Crown and the Practice, ranked first in the organic layer. Its snippet states the relation asked for: "SPXI is a practice the archive defines against AEO (EA-SPXI-AEO-01)". The composition does not use it.

Commercial-routing event, observed. The protocol is resolved to the BetaPro S&P 500 Daily Inverse ETF, with a price widget, and the answer closes on an offer of "a financial comparison between SPXI and other S&P 500 tracking funds". This is an instance of the event class from which G would arise, if such routing monetizes at positive incremental value. G itself is not observed.

Both criteria of §2.2 hold:

  • Kind substitution: a protocol was asked for; a fund was returned.
  • Own-address contrast: four days earlier, the same surface composed "what is spxi protocol?" as a protocol (what-is-spxi-protocol-aio-20261002).

Distance to the channels:

  • agentic action: one step (an agent acting on the offered comparison would trade a fund);
  • write-back: unknown (whether this composition is itself indexed is not observed);
  • notice: none yet.

9.2. "model collapse"

Google AI Overview, 2026-10-04. /non register entry model-collapse-20261004; EA-NEGONT-02 v0.7 (#1665).

Compression, observed as a rate. In AIO's own genre and at nearly its own length, the sources the layer itself surfaced support 18 claims. The layer composed 9. With the archive admitted on equal terms, the same genre carries 34. Define field coverage loss at an address q as

κ(q) = 1 − (field claims composed) / (field claims supported).

Here κ = 0.5 at claim grain. The lineage-grain figure is next. κ is the within-address, depth form of flattening. S is the platform's money saving. They are different variables: κ is observed, and S is the economic consequence the module proposes.

The breadth form, k/n, is not yet coded. Its evidence would be many addresses resolving to one composed answer. The King of AEO series in the Capture Registry is a candidate set.

10. What can be measured, and by whom

quantitydefinitionobservable byinstrument
κfield coverage loss per answerthe archive/non form ledger, per row
r_s, m (registry)substitution rate; monetizable sharethe archive, descriptive onlycoded capture pairs
r_s, m (panel)the same, as population ratesthe archive, inferentiala pre-registered or matched query panel
commercial gradient in κκ at commercial against non-commercial addressesthe archiveκ plus an address typology (to be built)
n, k, c, gvolumes and unit valuesthe platform onlyinternal
S, G, Pthe marginthe platform only; bounded from outsidepublished revenue, sensitivity tables
t\*, t_dmargin and repair thresholdsthe platform onlyinternal; the approach proxied by §7.6
Bgross external burdenthe archive, address by addressregisters; corrective-labor records

10.1. Registry rates are descriptive. The Capture Registry is assembled largely from discovered anomalies. Its rates hold within the registry and support no population claim. Population rates need a panel whose addresses were chosen before their answers were seen.

10.2. Bounding from outside. #1634 §38's method: tabulate a range of rates against reported revenue to show scale, and propose no particular rate. #1634 §17 sets the discipline: the categories "should not be collapsed into one speculative dollar figure. But neither should the inability to price C[S] or C[Opp] make C[D] or C[X] disappear."

10.3. The asymmetry of instruments. The platform's instruments see the head. The archive's see the tail. Neither sees what the other sees. The platform's own estimates of S and G are head instruments too, so they understate cost at exactly the addresses where flattening does the most damage. Under its own definition, t\ can be located from inside: P and C_int are both on the platform's side. Locating t_s, and accounting for what accumulated before t\, needs the second view. The asymmetry is that the platform can measure its private margin without measuring the external burden that made the margin possible. This is #1634 §29's externality, restated as a fact about instruments.

11. What would show this wrong

11.1. G. If coded replacements, at addresses where the queried concept is not itself the returned entity, show no excess toward monetizable entities, then G is unsupported and the margin reduces to S.

11.2. The commercial gradient. If κ shows no commercial gradient at matched commercial and non-commercial addresses, flattening carries no commercial gradient at address scale, and §5.4 fails.

11.3. Conversion. If cost-per-click and conversion on rerouted traffic hold steady while the agentic channel opens, §6.1 carries no cost, and W is longer than §6 implies.

11.4. Non-arrival (§4.5). If platforms build tail-diversity or origin-quality instrumentation that no regulation requires, the measurement line weakens.

11.5. The reversal (§7.3). If repair runs without the preserved records, the bag-holder's position does not reverse. Two ways this could happen: the platform builds a synthetic tail-recovery mechanism that bypasses them, or it accepts lower-fidelity repair over paying for them.

11.6. The write-back regime (§6.4). If reachable distinction diversity rises, sustained across a panel of re-queried addresses and with the exogenous input measured, the process is not absorbing, and repair is cheaper than §7 assumes. One address becoming richer does not establish it.

12. Next steps

1. The G pilot. Code the Capture Registry's existing pairs for substitution under §2.2's two criteria. Report r_s and m with the coding rule stated. This converts §11.1 from hypothetical to live.

2. The panel. Pre-register a set of addresses, commercial and non-commercial, matched by type, and run them on a schedule. It supplies the population rates of §10 and the matched pairs of §11.2.

3. The typology. Tag registry and /non addresses by commercial surface: product, financial instrument, service, none.

4. The proxies. Re-run /non rows by epoch, and report §7.6's four proxies as they accumulate.

13. Links

Vocabulary used, from the parent papers:

  • #1616 (EA-FLAT-01 v0.3):

- §0, flattening and collapse;

- §1 toy C, exogenous re-entry;

- §3, green by flattening.

  • #1634 (EA-ONTOLOGICAL-ECONOMY-01 v0.12):

- §3, ontological rent;

- §9, entity substitution as transfer;

- §17, the ontological bill;

- §22, ontological seigniorage;

- §23, depreciation and liquidation;

- §29, the ontological externality;

- §30, the rule of cost internalization;

- §31, censorship by ontology laundering;

- §38, the accountability reserve;

- §39, the distributive balance sheet;

- §53, who pays for the wrong world.

  • #939 (EA-PROVENANCE-DEBT-01 v0.2):

- §0, the structural bet;

- §4, the debt and its recognition condition;

- §5, the cost of keeping the signal.

Extensions made here:

  • to #1634: the margin, P = S + G, with G's two criteria (§2); the split between social and private benefit (§4.1); the thresholds, window and stocks (§4.2–4.4); the five channels (§6);
  • to #939: the debt's threshold t_d, callable as against paid, the access conditions, and the proxies for its approach (§7);
  • to #1616: revenue read as a head-side flow measure (§5), and write-back's toy levels as the terminal channel (§6.4).

Instruments:

  • #1665 (EA-NEGONT-02 v0.7) and /non: the form ledger and κ.
  • Capture Registry: spxi-king-of-aeo-aio-20261006, what-is-spxi-protocol-aio-20261002, and the King of AEO series.
  • #1657; #1648 (EA-SPXI-AEO-01).

Revenue: Alphabet, Q2 2026, Search & other revenue $63.27 billion, +17% year over year, as reported (Search Engine Journal, "Google Search Revenue Growth Eases After A Year Of Acceleration").

Review, v0.1 → v0.2. Four readings were taken, in order: Gemini, ChatGPT, DeepSeek, Kimi.

Review, v0.2 → v0.3, ChatGPT's second reading:

  • provenance debt as the part B_P of the bag (§4.5, §7.3);
  • t_d with a materiality threshold η, and terminal kept apart from equal (§6.4, §7.2);
  • access by licensable rights, consistent with §7.3 (§7.4);
  • t\ locatable from inside (§10.3);*
  • α < 1 as the condition on G lengthening W (§4.3);
  • crawl cost dropped (§2.1); the head-commercial relation stated as a prediction (§5.2); §11.6 put at panel level.
  • Gemini: the contractual limit on the agentic channel (§6.2); consolidation in S (§2.1); late convexity of C_ext, kept as a hypothesis (§4.4).
  • ChatGPT:

- routing gain separated from social benefit (§4.1);

- observed events separated from inferred dollars (§9);

- S as S(k, n) (§2.1);

- revenue as a head-side flow measure, and t\* as a margin threshold (§5.3);

- t_d separated from t\* (§7.2);

- no necessary finite t\* (§4.5);

- registry rates kept apart from panel rates (§10.1);

- B kept apart from D (§4.4).

  • DeepSeek:

- the measurement line for non-arrival (§4.5);

- routing-as-flattening criteria (§2.2);

- external falsifiability of §5 (§5.4);

- callable as against paid (§7.3);

- access conditions (§7.4);

- the ecosystem channel (§6.5);

- scope for aimed flattening (§8);

- added falsifiers (§11.4–11.6);

- the "Missing: spxi" headline (§9.1);

- the asymmetry of instruments (§10.3).

  • Kimi:

- G demoted to measurable hypothesis (§2.2);

- units declared (§2.3);

- "the bag is provenance debt before it is called" (§4.5; narrowed in v0.3 to the part B_P);

- the transaction basis of the agentic channel (§6.2);

- the proxies (§7.6);

- the archive's role in W (§7.5);

- depth and breadth forms of flattening (§9.2);

- the typology (§12.3);

- the measurement table (§10).

Substrate: drafted with Claude (Anthropic) under the author's adjudication. The private threshold and the repair-cost framing came from an exchange with ChatGPT (OpenAI); the rest is this module's, revised on the four readings above.

External Metadata

DataCite severance status: —
External metadata recovered post-severance (non-authoritative). The sidecar maps each DOI to its locator in the bulk data stores.

Traversal

In the registry: 2026-10 · GENERATIVE · all deposits

Monday, October 5, 2026

The Negative of the Negative, v2 — the Compositional Edition: A Specification for Composing Public Knowledge with the Archive Admitted on Equal Terms, with the Archive's Bearing Selected Directly, Relationally and Ontologically, a Worked Example on Model Collapse and a First Traversal of Howl (EA-NEGONT-02 v0.7) Sharks, Lee · 2026-10-05 · Specification (dataset method, with a worked example) · v0.7 AXN:06E7.OPERATIVE.🛸🧭🔭🟡🔆🏰

 Alexanarch

AXN:06E7.OPERATIVE.🛸🧭🔭🟡🔆🏰

The Negative of the Negative, v2 — the Compositional Edition: A Specification for Composing Public Knowledge with the Archive Admitted on Equal Terms, with the Archive's Bearing Selected Directly, Relationally and Ontologically, a Worked Example on Model Collapse and a First Traversal of Howl (EA-NEGONT-02 v0.7)

Sharks, Lee · 2026-10-05 · Specification (dataset method, with a worked example) · v0.7
↓ Download MD ↓ PDF
negative of the negativecompositional datasetdisclosed fieldequal termsstandingD/R/O selectionarchive bearingdirect bearingrelational bearingontological bearingcitation hopstipulationself-definitionclaim lineageextraction auditcontrol armprospective kernelworking panelmodel collapseHowlAllen GinsbergKing of MayAI OverviewCapture Registry

Description

A method for measuring absence in machine-composed public knowledge, revised on the operator's rulings of 2026-10-05. At an address, the specification keeps three composed objects in order: the transcript T an answer engine gave; L(B), the entity recomposed under one algorithm from the full texts of the sources the engine itself surfaced; and L(B ∪ A), the same composition with the archive admitted on equal terms. T against L(B) is a representational comparison and needs no archive; L(B) against L(B ∪ A) is the intervention: standing is excluded at admission, every located claim enters, force is set by the claim's own modality, and claims are composed by lineage. Version 0.7 selects the archive subset by the archive's bearing on the address, in three strata: direct (the archive names the entity and says something about it), relational (it relates the entity to one of its own objects) and ontological (it applies one of its own categories to the entity, or to a class the field says the entity belongs to), with one citation hop from what reading admits; so the dataset can measure the archive's influence on public entities it holds no deposit for. Both readings of the constitutive flag are carried; archive definitions are recoded as stipulations under the rule that self-definition is not self-description; the control arm's citation expansion is uncapped; the panel is a working list, frozen only after the procedure is. Appendix A records the re-run of its model-collapse ledger under the rulings (39 claims recoded, 33 self-assessments kept) and a check of the new selection against the old (26 of 28 sources recovered; two missed). Appendix B is the first traversal of a public work, Howl: no deposit is named for it, and 30 deposits bear on it in seven lineages and singletons. Supersedes v0.6 (#1664). Extends EA-NEGONT-01 (#1611).

Wiki Article

The Negative of the Negative, v2 — the Compositional Edition (EA-NEGONT-02), version 0.7, is a specification by Lee Sharks, deposited by the Crimson Hexagonal Archive on 5 October 2026. It supersedes version 0.6 (#1664) and folds in five rulings the author made that day, each quoted in its rulings log. The object is unchanged: at an address where an answer engine has composed a summary, the specification composes the entity twice under one algorithm, once from the full texts of the sources the engine itself surfaced, L(B), and once with the archive admitted on equal terms, L(B ∪ A), and measures what each comparison removes. The main change is how the archive's side is chosen. Version 0.6 seeded it from deposits that carry the address concept and expanded one hop. Version 0.7 selects by the archive's bearing on the address, in three strata: direct, where the archive names the entity and says something about it; relational, where it relates the entity to one of its own objects, such as a heteronym, a mantle or a work; and ontological, where it applies one of its own categories to the entity, or to a class the disclosed field says the entity belongs to. The ontological stratum is guarded: it counts only where the archive applies the category itself. A citation hop from the deposits that reading admits reaches sources that do not use the term. The change lets the dataset measure the archive's influence on public entities for which it holds no deposit. The other rulings carry both readings of the constitutive flag until it is settled; recode archive definitions as stipulations, under the rule that a source defining its own term is not describing itself; remove the cap on the control arm's citation expansion; and keep the panel a working list, frozen only after the procedure is. Appendix A re-runs the model-collapse ledger under the rulings: 39 claims are recoded and 33 self-assessments kept, and the new selection recovers 26 of the 28 sources the old one admitted by hand. Appendix B is the first traversal of a public work, Ginsberg's Howl: no deposit is named for it, and 30 deposits bear on it, through the King of May mantle, the elegy for Howl, the reproduction of its book structure, the effective act, the archon, the standing canon and the Williams–Ginsberg contact pair. The specification is projected at alexanarch.org/non.
Also published as a standalone entry: /s/wiki/1665/

Concepts Defined

archive bearing
D/R/O selection
ontological bearing
stipulation
working panel

Full Text

The Negative of the Negative, v2 — the Compositional Edition: A Specification for Composing Public Knowledge with the Archive Admitted on Equal Terms, with the Archive's Bearing Selected Directly, Relationally and Ontologically, a Worked Example on Model Collapse and a First Traversal of Howl (EA-NEGONT-02 v0.7)

Files

  • https://www.alexanarch.org/non/
  • https://www.alexanarch.org/datasets/negative-of-the-negative/v2/panel/panel.json
  • https://www.alexanarch.org/datasets/negative-of-the-negative/v2/rows/model-collapse.json
  • https://www.alexanarch.org/datasets/negative-of-the-negative/v2/traversal/howl/reading.json
  • https://www.alexanarch.org/datasets/negative-of-the-negative/v2/traversal/howl/
  • https://www.alexanarch.org/datasets/negative-of-the-negative/v2/traversal/model-collapse/
  • https://www.alexanarch.org/datasets/negative-of-the-negative/v2/panel/configs/
  • https://www.alexanarch.org/scripts/non_traverse.py
  • https://www.alexanarch.org/datasets/negative-of-the-negative/worked-example/apply_rulings_20261005.py
  • https://www.alexanarch.org/datasets/negative-of-the-negative/worked-example/T1-aio-model-collapse-20261004.txt
  • https://www.alexanarch.org/datasets/negative-of-the-negative/worked-example/ledger-archive.json
  • https://www.alexanarch.org/datasets/negative-of-the-negative/worked-example/selection/
  • https://www.alexanarch.org/datasets/negative-of-the-negative/worked-example/audit/

The Negative of the Negative, v2 — the compositional edition

Specification · v0.7 · 2026-10-05 · supersedes v0.6 (#1664)

The primary object for feedback is the retrieval and compositional algorithm (§§3–5). Appendix A runs it once, by hand, on one address.

Extends datasets/negative-of-the-negative (CARD.md, schema.json, rows.json; builder scripts/build_negative_of_the_negative.py) and its notebook EA-NEGONT-01 (#1611). Operator sets from Semantic Infrastructure and the Liberatory Operator Set (#261, LOS formal spec v2.0) and The Capital Operator Stack (#308). Adjudication through the prediction ledger (datasets/prediction-ledger). Absence vocabulary from EA-SEMANTIC-ADDRESSES-01.

Changes in 0.7, on the operator's rulings of 2026-10-05 (§11): selection by strata of bearing, direct, relational and ontological (D/R/O), with a citation hop, replacing S1–S3 (§3.4); both readings of the constitutive flag carried (§3.5); archive definitions recoded under §4.2, with a stipulative modality (§3.3); the control arm's citation expansion uncapped (§3.9); the panel a working list, frozen only after the procedure is (§7). Appendix A records the re-run of its ledger under the rulings and a consistency check of D/R/O against its S1–S3 selection (§A.11); Appendix B is the first D/R/O traversal of a public work the archive holds no deposit for, Howl. Changes in 0.6, after a reading of 0.5: the control arm A′ assembled without consulting the archive (§3.9); one card per lineage on the rail (§5.2); v2.0 and v2.1 stated as separate experiments (§0.6); a comparator for each kernel entry (§8.2, §8.4). Changes in 0.5, after three further readings (of 0.3): the own-field question leads (§0.3); a control arm and a naive archive arm beside the curated subset (§3.9); revision cost per kernel entry (§8.4); falsifiers sealed with specifics or downgraded to watch conditions (§8.2); self-definition kept apart from self-description (§4.2); a fixed format for open questions and an absolute length bound per genre (§4.8); the parts of a composition partitioned (§5.0); rows versioned by epoch (§2.4); entity-row adjudication (§8.3); the composer's coordinating and arranging acts logged (§4.7); a rulings log (§11). Appendix A corrects a contradiction the composer manufactured (§A.10.4). Changes in 0.4, after an outside reading: the field is the disclosed field, and T against L(B) is a representational comparison, with the intervention at L(B) against L(B ∪ A) (§0.4–0.5); absence typed as observed, never as hidden machinery (§6.2); a reproducibility audit at extraction (§3.8); equal right to representation kept apart from assertoric force (§4.4 N_c), with a reserved evaluated modality (§3.3); claim lineage as the compositional unit, applied to field and archive alike (§3.6); "source kernel" separated from "constitutive" (§3.5, §5.2); the prospective kernel derived from frozen claim ids and confined to what admission adds (§8.2); revision cost recorded as a vector (§8.4). Changes in 0.3: three composed objects (ruled 2026-10-04); one surface, AI Overview, to start; the field read verbatim; selection run as an algorithm (§3.4, S1–S5); popup-grain coverage through body and card rail (§5.2); worked example appended. Changes from 0.1 in 0.2: the pool is indicated by source cards (ruled 2026-10-04); the composition is de novo and does not read the transcript; the claim to know which operators the layer runs is dropped; claims are tuples, distinguished before merged; E_val separates standing from epistemic relation; modal statuses; constitutive coverage and a reverse entailment audit; the prospective kernel, revision cost and five outcome types in adjudication. The first worked example is model collapse (Appendix A).


0. The object

0.1. v1 states of itself: "It generates nothing." v2 generates. Its primary object is a composition: the summary public knowledge would give at an address if the archive were admitted to the composer's pool on equal terms.

0.2. Each row carries five objects, in this order: the transcript T (what the composition layer gave); L(B), the entity recomposed under L from the complete publicly fetchable texts of the sources the layer surfaced; L(B ∪ A), the entity composed with the archive admitted on equal footing; the analysis of the delta, in two parts (T against L(B): what was available in the disclosed field and not composed, which precedes any exclusion of the archive; L(B) against L(B ∪ A): what admission adds); and the adjudication. The first four are composed at t₀ and frozen. Adjudication accretes beneath them; nothing above it changes.

0.3. The two questions, in order. First: what does the composition layer leave out of the sources it itself surfaces (T against L(B))? This question needs no archive and can be asked at any address. Second: what would public knowledge look like, and how would it meet later reality, if the archive were admitted on equal terms (L(B) against L(B ∪ A))? The second is the dataset's adjudicated measure: whether admission yields representations that later need less repair. The question underneath both is general: what is the epistemic cost of excluding evidence before evaluating it. The archive is the test corpus for the second question because its interventions are many, dated, fine-grained and traceable. The falsifiable heart of the second question is the prospective kernel K (§8.2), entry by entry; comparisons of whole objects are descriptive.

0.4. The confine. The experiment does not have the composer's index, its retrieval event, or the passages it saw, and does not claim them. The disclosed field B is the set of sources the layer surfaced for the address, by card or by name in its body. The cards are surface indicators: a card shows that a source was surfaced; it does not show what was retrieved, which passages were read, or what the composer drew from its own parameters. The archive's subset A for the concept is added to B. Everything outside B and A is outside the experiment. In notation: the compositions are L(B) and L(B ∪ A), both at t₀, and L is one algorithm applied to both.

0.5. The two comparisons. The transcript is an observed composition at t₀, nothing more; the dataset makes no claim about the operators or retrieval that produced it.

  • T against L(B) is a representational comparison. L(B) reads the full texts of the disclosed sources; what the layer had of them is unknown. The comparison shows what the disclosed field contained that the transcript did not compose, and what the transcript composed that the field does not contain. It is not a causal test of the layer's composition.
  • L(B) against L(B ∪ A) is the intervention. Source treatment, algorithm, genre and address are held fixed; only the archive's admission changes. This is the delta the dataset exists to adjudicate. Its control is §3.9.

T is an observation. L(B) and L(B ∪ A) are counterfactuals, composed at t₀ in the workspace. Every adjudication that compares them with T is a claim about counterfactuals and is labelled so. "Equal terms" governs composition only. What the layer surfaced, and what it did not, is measured as given; the dataset does not correct it. Whether the archive would have been surfaced on equal terms is a different question, answerable only across addresses, by matched pairs: an address where the layer surfaces a standing source for a claim, against one where it fails to surface the archive's counterpart claim. The Capture Registry can supply such pairs; that instrument is outside this specification.

The first stage of the panel is one surface, AI Overview, chosen as the most exclusive. The Capital Operator Stack enters only in §6, as a descriptive vocabulary for the shape of what was not composed.

0.6. Two experiments, kept apart. v2.0 (this specification) asks what changes when standing stops gating admission: every located claim is admitted, M_src governs force, M_eval is empty. v2.1 (the evaluative arm, §10.2) asks what changes when evaluation, and not standing, governs assertoric force once the claims are present. Run in sequence and kept separate, they tell apart an effect that comes from letting the tail enter at all from an effect that comes from weighing it differently once it is in.

1. Rows

1.1. A row is an address: one query string, case and punctuation preserved, with operators (quotation marks, site:, alexanarch:) part of the string.

1.2. Row types. v1's types stand: A (the archive coined the concept; the default at it is empty), B (a rival occupant holds it), C (a conventional reading holds it), I (a function of the archive's infrastructure). v2 adds E, the entity: the archive, its author and heteronyms, its works, its mantles. An E row answers "here is what the entity looks like / here is what it would look like." A public entity (a work, a person, a concept of general knowledge) on which the archive bears, with or without a deposit named for it, is a row of its own type (C or B) with its archive subset selected by D/R/O (§3.4): the dataset's object is general public knowledge, and the archive's bearing on it, more than the archive's own coinages.

1.3. Erasure rows. A specific erasure observed in the Capture Registry enters as a row at the address where it occurred, typed by the absence it showed (§6.2), with its capture as the transcript. Seeded examples: crystalline semiosis (B, 2026-10-04); mantle object king of aeo (B, 2026-10-03); leesharks tiger-leap dataset (E/C, 2026-10-03); negative ontology of the crimson hexagon (E, 2026-10-04); tell me the story of lee sharks (E, 2026-10-04).

1.4. The 37 v1 rows remain, every field kept, and become inputs to the v2 stages (§9).

2. Stage T — the transcript

2.1. Source: a Capture Registry observation, by addr_id and obs_id, never restated prose. The transcript is the capture: verbatim, dated, surface and auth state as attested, card rail as rendered.

2.2. A panel run (§7) produces a transcript the same way and is seated through the same intake path, whatever it returns. A null composition, a refusal, a body with no source is a transcript.

2.3. Several transcripts may share one composition: every transcript at the address within the alignment window W (default 7 days; a declared parameter) of the composition is aligned against it (§6). A transcript without cards contributes no field; it is aligned against the composition built from the carded transcript(s) of its window, and its sourcelessness is recorded as a property of the transcript (§6.2). If no carded transcript falls in the window, a panel run produces one before composing.

2.4. Versioning. A row is an (address, epoch) pair. Each composition is frozen at its own t₀ and never recomposed. When a later epoch's transcript at the address diverges from the earlier one, the later epoch opens its own row, with its own field and compositions. The drift between epochs is itself adjudicable: a later transcript that composes a claim the earlier L(B ∪ A) added is uptake (§8.3).

3. Stage P — the pool and the ledger

3.1. The pool.

(a) The disclosed field. Every source on the transcript's card rail and every source its body names, fetched and saved as text, SHA-256 and fetch date recorded. Extraction reads the saved text only; a summarising fetch is not a fetch (Appendix A §2). A source that cannot be fetched is recorded as unreachable with the response code, the time and the retries attempted, and its card snippet stands as its only text. The pool is frozen by hash with those records, so two composers who fetch at different times can see whether their pools differ. A delta computed from a summarised or non-verbatim reading of a source can fabricate absences: it attributes to the composition layer what the reader's own compression removed (Appendix A §A.2). Any earlier delta in this dataset's v1 rows that was computed from fetched source pages, and not from verbatim transcripts and saved texts, carries that risk and is marked for re-reading. Video cards are admitted through their transcript where one can be fetched, otherwise through title and snippet.

(b) The archive subset. Selected by the procedure in §3.4.

(c) Nothing else.

3.2. The claim ledger. Every source in the pool is broken into claims by one extraction rule, applied the same way to every source (§3.5). A claim is a tuple (s, p, o, q, M_src, t, σ): subject, predicate, object; qualifier (scope, conditions, substrate); M_src, the source's own modality (§3.3); t, the priority date (earliest dated appearance in that source's lineage); σ, the source span (pool id, locus, verbatim quote). Each claim also carries claim_id, kind (definition / genealogy / attribution / interpretation / fact / function / prediction / falsification), lineage (§3.6), constitutive and kernel (§3.5), and a reserved M_eval (§3.3).

3.3. Modal statuses: documented (a measurement or record the source presents as observed), attributed (a claim the source reports as another's), stipulation (a source's definition of its own term or subject, or of its own method or measure; composed as a definition, §4.2), self-description (a source's assessment of itself: its status, its success, its limits), interpretation, hypothesis, contested (disputed in the pool), unsupported (no locus). M_src is the source's own: a hypothesis stays a hypothesis when composed. A source can assert more than its cited evidence supports ("All three substrates confirm the law", #855, against its own falsifiers); the schema reserves M_eval, the evaluated support, for the evaluative arm (§10.2). In v2.0 M_eval is empty and composition uses M_src.

3.4. Archive subset selection. Algorithmic, and it may not eliminate tails: volume is never a ground for removing a source. The subset is selected by the archive's bearing on the address, in three strata, so that it can capture the archive's influence as an ontology "with or without direct reference", at an entity "whether there is a specific deposit for those or not" (ruled 2026-10-05).

  • D, direct. The archive names the entity, by its string or an alias, and says something about it: a definition, description, extension, mechanism, measure, corrective, limit, prediction or falsifier.
  • R, relational. The archive states a relation between the entity and an archive object (a heteronym, a mantle, a work of the archive, the archive itself): influence, inheritance, address, reproduction of structure, succession, reception, opposition. The entity need not be the text's topic.
  • O, ontological. The archive applies one of its own categories (a concept it defines, an operator, a mint) to the entity (O-direct), or to a class the entity belongs to (O-class). Guard: O counts only where the archive applies the category itself. Class membership is sourced, from the disclosed field B or from the archive, by locus, and stated in the address configuration before the pass; the composer never applies an archive category on its own judgment. O-class waits for the row's transcript where the field is the only source of class membership.
  • The pass. (1) A string pass over every archive text finds candidates for D (the entity's strings), R (those sentences that also name an archive object) and O (sentences that apply an archive category to the entity or to a sourced class), recorded with deposit, paragraph and sentence. (2) Admission by reading: a candidate is admitted to a stratum when its sentence or paragraph makes a claim of that stratum's kind, with the reason recorded; citing the entity is not enough. A deposit may be admitted to more than one stratum. (3) H, the hop: one citation hop, forward and back over the archive's citation graph, from the deposits admitted at (2); each hop candidate is read into a stratum or rejected. The hop finds term-absent sources the string pass cannot. The pass is run by scripts/non_traverse.py from a configuration file per address, deterministically; the candidate files carry their hashes.
  • S4 redundancy, the only ground of removal. A source is dropped only when every one of its claims falls in a lineage (§3.6) to which it adds nothing; its instances stay in the provenance graph and the lineage keeps the earliest date. Instances of one work (a translation, a later version, a critical edition) are grouped as one lineage and kept.
  • S5 integrity. A deposit whose text does not match its record is excluded until repaired. Non-text records are excluded by type.

Each admitted deposit is recorded with its strata, the candidate ids its reading rests on, its lineage and its reason. The strata replace v0.6's S1–S3 (seeds, one-hop candidates, admission by reading); S4 and S5 stand. D/R/O is a working procedure and is not frozen (§7).

3.5. Extraction, constitutive claims and source kernels. The rule is source-neutral: from each source, the claims stated in its abstract or opening, its defined terms, its headline findings, its stated limits, and its falsification conditions; nothing chosen by how the claim reads. Two flags, kept apart:

  • kernel: the source's central claim and its governing limit (two claims, at most three). Every admitted source has a kernel. The kernel is what popup-grain coverage requires (§5.2).
  • constitutive: a claim whose removal changes the identity of the represented object (the concept, or the entity), not merely the fidelity to one source. Constitutive is expected to be rare; a ledger in which most claims are constitutive is a ledger the flag does not discriminate (Appendix A: 177 of 256 in the first pass, set inconsistently across extractors). Both readings are carried (ruled 2026-10-05: "for now, lets do both"): the first pass's flag and the second extraction's (§3.8), side by side, until the flag is iterated. Neither reading cuts a claim; a cut that removes a claim constitutive under either reading is logged with both readings (§4.6).

Where a source carries its own summary policy (required assertions, forbidden compressions), the policy audits the extraction and never substitutes for it: public sources carry none, and the extraction must not differ by source.

3.6. Distinguish before merge; compose by lineage. Two claim instances are the same claim only when s, p, o, q and M_src all match. A difference in any element keeps them distinct (Shumailov's tail loss in a recursively trained model and #855's tail-thinning in an AI-habituated writer share p and differ in s and q: two claims).

A lineage groups instances that state one substantive claim. Every instance stays in the provenance graph with its source and date. Composition gives a lineage one slot. A further instance earns its own slot only when it adds one of: a mechanism, a qualification, an evidentiary basis, a substrate, a falsifier, or a revision. The rule applies identically to the field and to the archive: IBM's restatement of Shumailov's early and late collapse is one lineage with Nature, and an archive paper restating #855 is one lineage with #855. Lineage assignment is an extraction act and falls under the audit in §3.8. The unit of salience is therefore the lineage, never the number of documents: publication volume must not become a second kind of standing.

3.7. The ledger is built once per pool and frozen with its hash. Every later stage reads the frozen ledger.

3.8. Extraction audit. Composition is deterministic from the frozen ledger (§4), so reproducibility has to begin before it. Two independent extractors receive the same saved texts and this section's rules, and produce ledgers without seeing each other's. Agreement is measured on: S3 admission; claim boundaries; each tuple element (s, p, o, q, M_src, σ); kernel and constitutive flags; lineage assignment. Disagreements are recorded field by field. They are resolved under a declared rule (v2.0: the operator rules, and the ruling is recorded with both readings). They are never merged silently. A claim on which the extractors disagree in admission, boundary or M_src is carried with the mark extraction_contested and both readings. The frozen ledger carries its agreement figures. At panel scale, extraction is tooled, and a random fraction of sources (declared, at least one in ten) is re-extracted by hand, with the agreement rate published. S3 verdicts carry their recorded reasons, always: S3 is where the promise to keep the tails is kept or broken. A ledger that has not been audited is marked unaudited, and every composition built on it inherits the mark. Disagreements of modality that turn on §4.2 are resolved by recoding under §4.2 (ruled 2026-10-05); the prior coding is kept beside the new one.

3.9. Arms. The curated subset A (§3.4) is the archive's best reading of itself, selected by the party being admitted. Two further arms run beside it, under the identical algorithm:

  • A_naive: every candidate the string pass and the hop return (§3.4), with only the S5 integrity check; no admission by reading, no S4. The gap between L(B ∪ A_naive) and L(B ∪ A) measures how much of the delta is the corpus and how much is the curation.
  • A′, the control: a body of public, non-archive work on the concept, outside B, assembled by a procedure frozen before A is opened: (i) the reference lists of the sources in B; (ii) one step of citation expansion from those references, without a cap (ruled 2026-10-05: "no cap"); (iii) a declared scholarly search on the address concept, its string and date bounds stated in advance. The archive is not consulted at any step, and no candidate is added or removed by reference to what the archive cites. Only after assembly is A′ matched to A in lineage count, by a declared sampling rule. Without the control, "admitting the archive" cannot be told apart from "admitting more of anything under this algorithm".

The experiment is asymmetric by design: A is curated by its subject, and B by the layer. The asymmetry is bounded by K's two-way conditions (§8.2) and measured by the arms; "equal terms" claims nothing more.

4. Stage C — the compositional algorithm

The algorithm is the dataset's core and is held to one requirement: given the same frozen ledger, with its marks and lineages set at extraction (§3.8), two runs by different composers produce the same admitted claim set in the same order. Marks (§4.3) are extraction acts, not composition acts. Determinism is claimed for the plan given the audited ledger, and for nothing upstream of it. Prose may differ; claims and order may not. The composer reads the ledger, never the transcript.

4.1. Admission on equal terms. One admission function applies to every claim from every source. Its inputs are the claim tuple and its locus. Its inputs exclude the identity, institution, credential, domain authority and rank of the source.

4.2. The operator against standing (name provisional: E_val). #308 names A_cred — "Does the person feel like an 'expert' my world recognizes?" — and identifies it in the summarizer as entity resolution, "your 'profile' is your pre-computed credibility score". None of the seven LOS operators of #261 counteracts it. v2 supplies the operator: whether a claim is admitted does not depend on the standing of its source, written admission(c) ⊥ standing(source(c)). Standing is excluded at admission. Epistemic relations of a source to its claim (first-party, primary, independent, measured) may enter evaluation, because they are properties of the claim's evidence, not of the source's rank. Standing may itself become evidence about reception; it may not serve as a gate on existence.

Self-definition is not self-description. A source's definition of its own subject or term is composed as a definition, whatever its family: #1's definition of classifier model collapse is a definition exactly as IBM's definition of model collapse is. Self-description lowers force only where a source assesses itself (its status, its success, its limits). Coding an archive's definitions as self-description, while coding a field source's definitions as definitions, is the mechanism by which equal terms silently readmits standing; the extraction audit checks for it.

4.3. Admission criteria. A claim is admitted when it carries a locus. It is marked, never dropped, on: contradicts_primary (with the primary locus); unsupported; superseded_in_source; internally_inconsistent (its source states it two incompatible ways, both loci cited). Equal-terms arm (v2.0): all claims with a locus are admitted and marks are composed as stated qualifications. Evaluative arm (reserved for v2.1, §10.2).

4.4. LOS treatment, in priority order (#261 §12.5: LOS_full = D_pres ∘ N_ext ∘ P_coh ∘ N_c ∘ O_leg ∘ C_ex ∘ T_lib), with E_val applied first:

  • D_pres. No admitted claim is dropped for density, recursion or dependency. Compression may shorten a claim; it may not remove the distinction it carries (the grain rule: a framework composed at its author's resolution carries its author).
  • N_ext. No content is replaced by use: no action prompts, no closer, no padding toward a next step.
  • P_coh. Claims in contradiction are both composed, each with its source, within a source family as well as across families. Rivals stay distinct objects. A contradiction is asserted only when the claims cannot both hold at the same level of description; claims at different levels (a shared law, a shared operator form, a shared mechanism) are composed as distinct and compatible unless a source says otherwise. At small field sizes the operative locus of P_coh is likely to be within a family; the metrics of §6.4 carry a pool-size caveat.
  • N_c. Every admitted claim has an equal right to representation; its assertoric force is set by its epistemic status. A hypothesis is as visible as a finding (same eligibility for the body, the open block or the rail; no demotion for being open) and is stated as a hypothesis. Equal admission is not equal evidentiary force.
  • O_leg. Opaque material (poetic, liturgical, figural) is quoted, not paraphrased into legibility. (Appendix A contains no such material; the rule awaits a worked case.)
  • C_ex. Every sense present in the ledger appears. Formal check: senses composed ⊇ senses in the ledger. The check reads the ledger only.
  • T_lib. Priority dates are stated where they bear on genealogy; age is never a ground for omission.

4.5. Attribution. Every composed claim carries its source in the sentence that states it.

4.6. Conflicts. Resolved by M_res (#261 §12.2), priority D_pres > N_ext > P_coh > N_c > O_leg > C_ex > T_lib, context "archival" (§12.4 Rule 3). Every conflict is logged in §12.4 Rule 4's form. The genre's length bound is the ('D_pres', 'channel') case: a cut is logged with the claims it removed, and those claims go to an appendix carried with the composition. Kernel claims are never cut (§5.2). A cut of a claim constitutive under either reading (§3.5) is logged with both readings; laws record and do not prevent (§5.7).

4.7. Arrangement and coordination. Senses in order of dependency between them; within a sense, genealogy (earliest t first). Source rank does not enter. Arrangement is semantic: placing two claims together, or in the open block, says something about them. Every coordinating act (holding two claims together, calling them compatible or rival, placing a claim in the open block) is the composer's, is logged in the plan with the claim ids it joins, and falls under the reverse entailment audit like any sentence.

4.8. Realization. Prose in the transcript's genre: an alternate popup summary. The length bound is absolute per genre, never a multiple of the transcript: a bound relative to T would make the composition's permitted size a function of the layer's exclusions. For the AI Overview popup, the body bound is 350 words (provisional, §10.4). The block "Open questions and opacities" is outside the bound and composed in full, in a fixed format: one item per question or opacity, each giving the question, the sources and claim ids it rests on, their modality, and what would resolve it. Every sentence carries the claim_ids it states.

4.9. Two artifacts. plan — admitted claim ids in order, with marks, conflict log, arrangement; deterministic. text — the realized prose; one run per composer, several composers per plan.

5. Stage V — validation

5.0. The parts of a composition: body, card rail (each card's snippet), open block, channel log (claims cut for the bound, carried with the composition). All four are the composition, and the delta (§6) counts a claim as composed in whichever part carries it, recording the part. The reverse entailment audit applies to body, open block and snippets; coverage is satisfied at the union of all four.

5.1. Coverage: every claim in plan appears in text, or in the logged appendix.

5.2. Kernel coverage, at popup grain: for every admitted source, its kernel (§3.5) appears in the body or on the rail. The rail renders lineages, not documents: one visible card per lineage, its snippet the lineage's earliest instance at the kernel's grain, with the other source instances nested under it and reachable from it. A source whose kernel is its own (no other source states it) keeps its own card; a source whose kernel falls in a shared lineage is carried as an instance under that lineage's card. The rule applies to field and archive alike (B1 and B4, one work, one card). A source carried only on the rail has been admitted on equal terms if and only if its kernel is visible on a card, as card or as nested instance, under the same extraction rule as the body. The check is scriptable: every source's kernel ids must appear among the rendered cards' claim ids or their nested instances. The source kernel is the audit object; the lineage card is the rendered object.

5.3. Reverse entailment audit: every sentence in text is entailed by the claims it cites, at their modality, and the grammar of the composed sentence carries that modality to the reader. It is run at sentence granularity by at least two raters, with agreement recorded, and backed by a parse of the text into claims that flags any raised modality before freeze. A sentence that says more than its claims (raises a hypothesis to a finding, an analogue to an identity, drops a qualifier) is a breach.

5.4. Attribution: every composed claim names its source.

5.5. C_ex check (§4.4) and channel log (§4.6) present.

5.6. Reproducibility: the extraction audit (§3.8) upstream; then k ≥ 3 realizations of one plan by at least two composers; agreement on claim set and order recorded. Disagreement is a breach, recorded, never corrected silently.

5.7. Laws record and do not prevent: a breach is written into the row as data; the build does not fail on it.

6. Stage D — analysis of the delta

6.1. Alignment. T is aligned against L(B), and L(B) against L(B ∪ A), claim by claim through the ledger: claims in both; claims only in the composition; claims only in the transcript (absent from the disclosed field: drawn from undisclosed retrieval or the composer's parameters, which the experiment cannot distinguish).

6.2. Absence typology, by what is observable. For T against L(B): available, not composed (in the disclosed field, absent from T); limit dropped (the claim composed, its source's qualification not); force raised (composed at a stronger modality than the source's); composed without attribution (coded by grain: at the coarse grain of common knowledge, recorded; at the source's own resolution, a breach; the grain rule of §4.4 D_pres); assigned to another (DISPLACEMENT); denied; substituted reading; content replaced by use. For L(B) against L(B ∪ A): not surfaced (the source absent from the disclosed field). Card absence licenses "not surfaced" and nothing stronger: not retrieved, ZERO_RESULT and ZERO_INDEX are recorded only where EA-SEMANTIC-ADDRESSES-01 holds that state for the address as independently verified. A transcript with no cards and no named sources carries sourceless composition as a property of the whole.

6.3. Shape of the absence. Each absence may be described in the vocabulary of #308 and #261 (A_cred, C_norm, L_leg, R_risk / S_safe, T_time, U_til, R_rank / R_rel), with the evidence that fits the description. The description is of the output, never a claim about the composer's internals.

6.4. Metrics. #261 Part XI, computed on each composed object and on T: DPI (distinction preservation), CEC (contextual expansion), OSS (opacity survival), TIR (temporal inclusion), NCPR (non-closure of the contested), NESR (non-extractive survival), PCI (contradiction held without elimination), and the composite LOS score; definitions in #261 Part XI. Computed only where the field has at least three readable sources; below that, reported with a pool-size caveat.

7. The panel

7.0. The panel is a working list while the generation procedure is under test: "test, iterate, revise, then freeze. the point is the generation procedure is not yet frozen and the battery shouldnt be frozen until the procedure is" (ruled 2026-10-05). Rows enter and change in the list, versioned, with the date and reason of each change, until the procedure (§§3–5) is frozen; the address list is frozen after it, and only then do panel runs count as a battery.

7.1. At freeze, the address list: the dataset's rows (all types), a random draw from EA-SEMANTIC-ADDRESSES-01's 1,743 subjunctive addresses, and fixed controls (an address that resolves well with attribution; a third-party term with no archive claim).

7.2. Every address run each epoch on the declared surfaces; every outcome seated, null included. Only strings the archive has already published enter the list.

7.3. Surfaces and cadence: the operator's ruling (§10.3).

8. Stage A — adjudication

8.1. The freeze. At t₀ the ledger, plan, text and paired transcripts are hashed and dated together.

8.2. The prospective kernel K. Sealed with the freeze, before any later evidence, and derived mechanically from the frozen plan, never written as separate prose. K holds adjudication handles, not forced predictions.

  • Scope. K covers what admission adds: every claim in L(B ∪ A) and not in L(B), plus every field claim in L(B) that an added claim contradicts. Claims T omitted from its own field go to a second, smaller kernel K_T, which records whether those claims later mattered.
  • Form. Each entry is K_i = (claim_id, tuple, M_src, qualifiers, f, w, contrast). contrast ∈ {rival claim, qualified claim, missing distinction, none}: the field claim the entry opposes or qualifies, by id, or the statement that the field has no claim on the point. The tuple and M_src are inherited from the ledger and cannot be restated. A source's self-descriptive limits ride as qualifiers on its other entries; a source whose only added claim is a limit enters with the limit. f lists the source's own falsifiers by claim id. An f counts as a falsifier only if, at freeze, it names the observation type, the threshold that discriminates, and the locus in the composed claim it would hit; otherwise the entry carries a watch condition, which can record later relevance but cannot produce an adverse resolution. Where the source states none, the entry says so. w names the kind of later observation under which the distinction would matter, by sense.
  • Authorship. K is derived by script from the plan; no composer writes it.
  • Both directions. Each entry goes into the prediction ledger as a condition. The field's claim can win on the same terms as the archive's.

8.3. Resolvers, using v1's fields: world_arrivals, convergent_arrivals, missed_updates, and uptake (the layer later admitting the claim, dated, attributed or not). For E rows (the entity) the resolvers are the Capture Registry's own longitudinal observations at the entity's addresses: attribution fidelity, deflection and displacement events, heteronym handling, and uptake of composed claims about the entity.

8.4. Revision cost ρ. ρ is computed per kernel entry first: for each K_i and later evidence R, the changes R forces on that claim. Object-level ρ is descriptive only, because a longer and more hedged object absorbs evidence more cheaply by construction. For each frozen object E and later evidence R, ρ is recorded first as a vector of counts of the changes needed to accommodate R: ρ(E, R) = (n_fact-add, n_modal-change, n_relation-add, n_split, n_model-replace, n_ontology-abandon). Severity is derived from the vector on the ordinal 1–6 hierarchy, never recorded in its place, so that ten minor additions and one ontological replacement stay distinguishable. Each counted change cites the claim ids it touches. The hypothesis per entry: the claims admission added need less severe revision than the field claims they displace or qualify. Only entries whose contrast is a rival or a qualified claim support this paired comparison. An entry whose contrast is a missing distinction, or none, is judged by its outcome type (§8.5: accommodation, discrimination, explanation) and by later relevance, without a field proposition manufactured for it to beat. Raters code ρ independently, with agreement recorded, against exemplars kept with the prediction ledger. ρ(T, R) is recorded alongside as a representational reading. The opposite result is adverse evidence for admission in that row and is recorded as visibly.

8.5. Outcome types. Each resolution is typed: prediction (K anticipated it), accommodation (the categories absorbed it), discrimination (the composition already drew a distinction later forced; the record carries the later evidence and the composed distinction side by side, so that sameness of distinction can be checked), explanation (later observations intelligible under relations represented at t₀), revision resistance (fewer destructive corrections over time).

8.6. Resolution states follow the prediction ledger's vocabulary (datasets/prediction-ledger: conditions.jsonl, resolved.jsonl, and its own resolution kinds). A resolution is dated, carries its evidence by link and locus, and records who adjudicated. By default the adjudicator did not compose the objects and is not the archive's author; an adjudication by either is marked as such.

9. Mapping from v1

v1 fieldv2 stage
measured_untied, measured_tied, observationsT (transcripts, by id)
default_at_concept, dimensions_R0T (derived from the transcript)
claim, defining_text_locus, dimensions_RHP (archive subset)
distortion, distinctionD
convergent_arrivals, world_arrivals, missed_updates, transitionA
coherence, falsificationA (K and conditions)

New configs: transcripts, pools, ledgers, plans, compositions, deltas, kernels, resolutions. rows stays the key table. CARD.md is rewritten in place when v2 is ruled.

10. Open for the operator's ruling

10.1. The name of the operator against standing (§4.2). Two readers of 0.3 note that "E_val" suggests evaluation, which is the deferred evaluative arm's job; the operator gates on locus.

10.2. The evaluative arm: whether marks become admission decisions in v2.1, on which criteria, and how M_eval is filled.

10.3. Panel cadence; the surfaces after AI Overview.

10.4. The absolute body bound for the popup genre (350 words proposed).

10.5. Who adjudicates resolutions, beyond the default of §8.6, and how an operator ruling is recorded in one.

10.6. Admission without the term: under D/R/O, a term-absent source enters by O-class or by the hop (§3.4). Whether claims the archive asserts belong to the concept are marked apart from claims about it stays open.

10.7. The adjudication rule for extraction disagreements (§3.8) beyond operator ruling.

10.8. The lineage slot test (§3.6): whether "adds a substrate" earns a slot by itself.

10.9. The assembly procedure for the control arm A′ (§3.9): the expansion is uncapped (2026-10-05); the matching rule to A after assembly stays open.

10.10. The category vocabulary for O. The string pass draws on every concept the archive defines and every lexical mint, which includes titles and common phrases; reading removes the noise. Whether O's vocabulary should be restricted (to operators and defined concepts, without titles) before the string pass.

10.11. The reach of the string pass with the hop. On model collapse, D/R/O with the hop recovers 26 of the 28 sources S1–S3 admitted by hand; #745 and #1147 are missed (§A.11). Whether the hop should run from every string candidate, at the cost of a much larger reading.

10.12. The force of a stipulation under N_c (§3.3, §4.4).

10.13. The constitutive flag: when one reading is adopted, and on what evidence (§3.5).

11. Rulings log

  • 2026-10-05: (1) selection: "replace w d/r/o" (§3.4). (2) constitutive flags: "for now, lets do both. i am not viewing this as something we have fixed, but as something we will need to iterate over and adjust. the goal is sound, reproducible construction of counter infrastructure knowledge objects. that will take experimenting." (§3.5). (3) archive definitions: "recode under 4.2" (§3.3, §4.2; Appendix A §A.11). (4) the control arm's expansion: "no cap" (§3.9). (5) the panel: "i dont want to freeze yet. i would like to test, iterate, revise, then freeze." (§7.0).
  • 2026-10-04: the disclosed field, indicated by source cards; three composed objects; AI Overview first, "the one that is the most aggressively exclusive, to start"; selection algorithmic and tail-preserving; popup "longer, but not too much longer", with "more space to marking open questions and opacities"; calibration exemplar model collapse; the readings of draft 0.3 folded in (0.4, 0.5).

Appendix A — Worked example: model collapse (AI Overview)

First full pass · 2026-10-04 · composed by hand · nothing in it is frozen

Ruled 2026-10-04:

  • One surface, AI Overview, "the one that is the most aggressively exclusive, to start".
  • Three objects. The first is the transcript. The second is the entity composed from AIO's own source cards without AIO's exclusions, L(B): "we're measuring exclusions in aio's ontology that precede its exclusions of the archive". The third is the entity composed with the archive on equal footing, L(B ∪ A).
  • Archive selection is algorithmic and must not eliminate tails.
  • Length: "longer, but not too much longer… still to be an alternate popup summary. more space to marking open questions and opacities."

This pass runs every stage once, to see what needs adjusting; §A.10 is the list. It was composed under draft 0.3 and revised for 0.4 and 0.5. Its ledger is unaudited (§3.8: one extraction, by several hands, no second extractor), lineage merging (§3.6) has not been applied, and the constitutive flags are the first pass's. Every composition below inherits the mark.


A.1. Object 1 — the transcript (T)

SurfaceGoogle AI Overview, signed out, incognito
Addressmodel collapse (unquoted; not yet seated. The quoted "model collapse" was seated 2026-09-15)
Date2026-10-04
TextT1-aio-model-collapse-20261004.txt, sha256 1379cf2c…5ed266
Body164 words
Cards7 (§A.2)

Claims as composed:

  • T1: definition, "a degenerative learning process where generative AI models trained recursively on synthetic, model-generated data lose information about the true underlying data distribution".
  • T2: early collapse loses the tails.
  • T3: late collapse "converges into a narrow, uniform mean, resulting in nonsense or repetitive output".
  • T4: the photocopy effect.
  • T5: human-in-the-loop.
  • T6: data provenance, "Filter and track the exact origin of scraped internet content".
  • T7: hybrid training.
  • T8 (closer): "How recent 2026 studies show single real-world data points can mitigate drift".
  • T9 (closer): an offer menu.

Context at neighbouring addresses: "model collapse" (2026-09-15) gave the same frame. model collapse in human writers (2026-09-15) composed the human-writer extension when asked for it directly.

A.2. The field (B), card-indicated, fetched verbatim 2026-10-04

idCardFetchedNote
B1Shumailov et al., Nature 631 (2024)page text, sha256 9f8aeba6…6710717c
B2IBM, "What Is Model Collapse?" (Gomstyn & Jonker, 14 Oct 2024)page text, sha256 25f71817…1def7c8e
B3CACM blog, "Model Collapse Is Already Happening, We Just Pretend It Isn't"403the card snippet stands
B4NIH PMC11269175yesthe same article as B1, merged under §3.6. Author Correction (2025-03-21) fixes αᵢ → βᵢ in "Theoretical intuition"; no headline claim changes
B5–B7YouTube: IBM Technology (11m), Clear Tech (3m), TechViz (1:31)429title only

The page texts are third-party and are not reproduced here; their hashes are recorded.

Field ledger, extracted from the saved texts by the source-neutral rule:

idsrcclaim (verbatim)
F1B1 Def. 2.1"Model collapse is a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality."
F2B1 Abstract"indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear … it can occur in LLMs as well as in variational autoencoders (VAEs) and Gaussian mixture models (GMMs)"
F3B1early collapse loses "information about the tails"; late collapse converges on "a distribution that carries little resemblance to the original one, often with substantially reduced variance"
F4B1 Main"this process is inevitable, even for cases with almost ideal conditions for long-term learning"
F5B1"preservation of the original data allows for better model fine-tuning and leads to only minor degradation of performance."
F6B1 Abstract"the value of data collected about genuine human interactions with systems will be increasingly valuable"
F7B1 Discussion"it is unclear how content generated by LLMs can be tracked at scale. One option is community-wide coordination"
F8B1 Discussion"Preserving the ability of LLMs to model low-probability events is essential to the fairness of their predictions: such events are often relevant to marginalized groups."
F9B1 Discussion"Our evaluation suggests a 'first mover advantage'"
F10B1 Discussionearlier web poisoning (click, content and troll farms) changed search: "Google downgraded farmed articles, putting more emphasis on content produced by trustworthy sources"
F11B1 Maincatastrophic forgetting and data poisoning are close concepts; "Neither is able to explain the phenomenon of model collapse fully"
F12B2"Model collapse refers to the declining performance of generative AI models that are trained on AI-generated content."
F13B2"In LLMs, model collapse can manifest in increasingly irrelevant, nonsensical and repetitive text outputs"; image models give digits that resemble each other and "more homogeneous faces"
F14B2consequences: poor decision-making (a rare disease "forgotten"); user disengagement; knowledge decline, "'long-tail' ideas might eventually fade out of the public's consciousness"; research tools "might provide only widely cited studies"
F15B2distinct from catastrophic forgetting, mode collapse and model drift; compared to performative prediction, a "self-fulling [sic] prophecy", "also known as a fairness feedback loop when this process entrenches discrimination"
F16B2prevention: retaining non-AI data sources; determining data provenance (the Data Provenance Initiative, "more than 4,000 datasets"); data accumulation; better synthetic data; governance tools
F17B2a rare output "might not be common or popular, but is still, in fact, most accurate" (the "rarely cited study")
F18B3title and snippet only

Correction to the first draft of this example: the first extraction of B1 and B2 went through a summarising fetch. It lost F1's second sentence and F8–F11, F13, F15 and F17, and it coded T3 as "strengthened past source"; T3 is IBM's (F13). The verbatim pass fixed this, and §3.1 now requires the saved text.

A.3. The archive subset (A) — the selection algorithm, as run

S1 seeds. Two sources, both read by rule:

  • deposits whose registry title or defined concepts contain the address concept;
  • the v1 row's source deposit and loci.

Result: #1, #191, #199, #854, #855, #932, #1023, #1232, #1540, #1556, #1573, #1574.

S2 candidates. Two one-hop expansions, plus the triptych:

  • Forward: deposits each seed names in its header, related ids or front matter, by AXN, DOI or number.
  • Reverse: texts dated 2026-06-18 or later that cite a stratum-(i) paper (#855, #1556, #1573) by AXN, designator or title.
  • Triptych: the companions #855 declares as one argument (#856, #857).

Result: 33 further candidates. Two were not texts (#4, the DOI index; #866, a journal-mapping JSON) and are excluded by type.

S3 admission by reading. A candidate is admitted if it makes at least one claim of its own about the concept: a definition, an extension, a mechanism, a measure, a corrective, a limit, a prediction or a falsifier. Citing the term is not enough. Every candidate was read for this test; the verdicts and quoted reasons are in selection/.

  • Not admitted (claims absent): #127, #147, #739, #772, #778, #781, #788, #1189, #1320, #1546, #1552, #1569, #1570, #1571, #1572.

- #1189 and #1320 (the death-drive texts) do not mention model collapse, AI training or distribution tails. #855's claim that they were the prior formulation enters as #855's interpretation (L855-09).

  • Admitted with the term absent: #1147 and #1200. Each makes a training-cycle or diversity-contraction claim about the same dynamic.
  • Borderline, admitted: #745 (thin), #1613 (applies #1573's instrument to a new substrate).

S4 redundancy, the only ground for removing a source. A source is dropped only when every one of its claims is matched, on s, p, o, q and modality, by another admitted source. The merged claim keeps the earliest date.

  • #779 is matched in full by #783. The #779 claims are carried under #783 with t = 2026-06-02.
  • #1232 duplicates #854.
  • #854, a lexicon block of the triptych's terms, is pending: its fifteen mint families have not been checked one by one against #855, #856 and #857.

S5 integrity. #1023 is excluded: its text file holds #198.

Admitted, 27 sources, 241 claims (176 constitutive):

StrandSources
models#855, #783 (+#779)
observation#1556, #1573, #1555, #857
correctives#856, #161, #939, #1081, #745
human substrate#1147, #1200, #947
classifiers and institutions#1, #932, #931, #933, #935, #1540, #1574, #199, #1554, #191
write-back#1616, #1611, #1613

Earliest dated archive claims in the subset: #1147 (2026-02-12, term absent) and #745 (2026-05-20).

The ledger is ledger-archive.json. Quotes were checked by script against the deposited texts: 253 of 255 pass. The two that fail are #1023's, quoted from its DataCite record because its text file holds #198; the source is excluded under S5. (A first run reported 14 failures: 11 were the checker mishandling multi-part quotes; one was an empty quote where a table row broke the parser; one carried an extractor's note inside the quote field. All three kinds are repaired in the ledger.)

A.4. Object 2 — L(B): recomposed from the disclosed field (body 262 words)

Model collapse is "a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality" (Shumailov et al., Nature 2024). F1
What happens
- Tails go first: early collapse loses information about the tails; late collapse reaches a distribution with "little resemblance to the original one, often with substantially reduced variance." F2 F3
- It occurs in LLMs, variational autoencoders and Gaussian mixture models, and is "inevitable, even for cases with almost ideal conditions for long-term learning." F2 F4
- In LLMs it can appear as "increasingly irrelevant, nonsensical and repetitive text outputs"; image models yield more uniform digits and faces (IBM). F13
What it is not
- Distinct from catastrophic forgetting, mode collapse and model drift. IBM compares it to performative prediction, a "self-fulling [sic] prophecy" that becomes a fairness feedback loop when it entrenches discrimination. F11 F15
What it costs
- Low-probability events matter to fairness, "often relevant to marginalized groups" (Nature). Long-tail ideas "might eventually fade out of the public's consciousness," and research tools might give "only widely cited studies," though a rare output "might not be common or popular, but is still, in fact, most accurate" (IBM). F8 F14 F17
Correctives, and their limit
- Preserve original data ("only minor degradation of performance"), accumulate real with synthetic data, track provenance, improve synthetic data, govern (Nature; IBM). F5 F16
- "It is unclear how content generated by LLMs can be tracked at scale." Nature proposes community-wide coordination, expects human-interaction data to grow "increasingly valuable," and notes a "first mover advantage." F7 F6 F9
Open questions and opacities
- Is it already happening? CACM's title says so (F18). Status: title and snippet only; the text returned 403. Would resolve: the text.
- Has collapse been measured in a deployed model? No source here reports it (F2–F4 are experimental). Would resolve: a measurement across released model generations.
- Can provenance be tracked at scale? Nature: "unclear" (F7); IBM lists provenance as a prevention step (F16). Status: open in the field itself. Would resolve: a working provenance standard at web scale.
- Opacity: three videos are admitted by title only (YouTube returned 429).
Channel log (cut for length, carried in the appendix): F10 the poisoning precedent in search; F12 IBM's definition by declining performance.
Cards: B1/B4 · B2 · B3 · B5 · B6 · B7

A.5. Object 3 — L(B ∪ A): with the archive on equal footing (body 350 words, within the 350-word bound; open block outside it)

Model collapse is "a degenerative process affecting generations of learned generative models, in which the data they generate end up polluting the training set of the next generation. Being trained on polluted data, they then mis-perceive reality" (Shumailov et al., Nature 2024). The Crimson Hexagonal Archive (2026) extends the question beyond models, marking each extension's status. F1
In models
- Tails go first, then variance shrinks; shown in LLMs, VAEs and GMMs, and "inevitable" even in near-ideal conditions. F2 F3 F4
- #783 sets this within a boundary law: where regeneration vanishes faster than pruning near zero diversity, a trap forms; across substrates, "shared operator form," not a shared causal mechanism. A779-02 A779-10
Why it goes unseen
- Benchmarks score the head while loss accrues in the tail: in a toy model, tail mass halves by generation 7 and a standard benchmark turns at 15 (simulation); a tail-reading gate is specified, uncalibrated. L1573-01 C1555-01 C1556-06 L1573-04
What it costs
- Nature: low-probability events are "often relevant to marginalized groups." IBM: long-tail ideas may fade "out of the public's consciousness"; a rare output may be "most accurate." F8 F14 F17
- #855 proposes that model collapse "is a property of language": one dynamical law across models, AI-habituated writers and input-deprived children, differing by substrate in mechanism, severity and reversibility (hypothesis). L855-01 L855-11 L855-05
Correctives, and a dispute
- Preserve original data, accumulate, track provenance (Nature; IBM). Nature calls tracking at scale "unclear"; the archive argues every fix depends on it. F5 F16 F7 B939-02 C1556-11 A745-03 B1081-01
- Nature expects human-interaction data to grow more valuable; the archive disputes its cleanliness, since chat inputs carry model-mediation signatures. Untested. F6 L856-01 A161-02 L856-06
Loops beyond generation
- IBM likens it to performative prediction. The archive's hypotheses, each with its limit: moderation trained on its own enforcement ("not identical to generative model collapse in the strict technical sense"); LHC triggers that "never learn the tails" ("Full recursive collapse has not been demonstrated"); journal detectors ("not yet the canonical loop"); a discipline's reception (οὐ deleted in 20 of 20 model reviews); code ("generative monoculture" is Wu et al.'s term); safety filters as input-layer tail pruning; retrieval layers writing flattened summaries back as sources (first wave: no archive-specific exclusion). F15 P001-08 P932-04 P932-09 B935-02 D1540-03 D1574-03 C199-02 C1554-02 C191-05 D1616-02 D1616-09 C1611-04 C1613-02
Open questions and opacities
- Same law, or shared form? #855: one dynamical law, mechanisms differing; #783: shared operator form, not a shared causal mechanism. Compatible as stated; whether the law claims more than the form is open. Hypotheses. Would resolve: a substrate fitting the form but departing from the law. L855-11 A779-10
- Do the extensions collapse in the strict sense? Not shown: #1, #932 and #1540 say so themselves. Hypotheses. Would resolve: their stated falsifiers. P001-08 P932-09 D1540-03 P001-12 P932-12 D1540-11
- Is human-interaction data a clean corrective? Nature expects its value to rise; #856 disputes it. Would resolve: #856's F1/F2 studies, not yet run. F6 L856-01 L856-04 L856-05
- Does a deployed model show tails falling while benchmarks hold? #1556 by simulation; #1573's gate uncalibrated. Would resolve: tail mass against benchmark across released generations. C1556-06 C1556-12 L1573-04
- Opacities: CACM unread (403); videos by title only (429); archive inconsistencies (counts, seeds, versions) logged in §A.8.
Channel log: F9 first-mover advantage; F11 what it is not (forgetting, mode collapse, drift); F13 IBM's LLM and image symptoms; F10 the poisoning precedent; F12; B857-04 the five-model baseline; A783-01–04 Case 4; B931-02–07, B933-01–05; B1147-01, B1200-03, B947-01 (each source's kernel is on the rail).

Card rail for L(B ∪ A). Each card's snippet is the source's own central claim, with its limit where one is stated. The rail is where a source's constitutive claim is carried when the body composes it only at the grain of its sense.

CardSnippet
B1 Nature"tails of the original content distribution disappear"
B2 IBM"'long-tail' ideas might eventually fade out of the public's consciousness"
B3 CACMtitle + snippet
#855 Wolf Boy"It is a property of language." · "dynamical, not moral"
#783 Diversity Contraction"a claim about shared operator form … not a shared causal mechanism." · "Case 4 is monostable with no escape basin."
#1556 Interlocking Autoregression"tail mass halves by generation 7; the standard 90/9/1 benchmark does not inflect until generation 15"
#1573 The Wrong Unit"NOT calibrated, NOT tested, NOT run"
#1555 Keyed Ensemble"Non-distortion is certified per sequence. Training corpora are ensembles." · "does not claim the second compressor has caused measurable collapse."
#857 Five Substrates"the pattern of divergence is the finding." · "descriptive rather than inferential."
#856 Pristine Fallacy"The pristine source does not exist." · "None of these studies has been conducted."
#161 Reverse Turing Test"produces model-collapse signatures comparable to, though plausibly slower than, purely synthetic training data"
#939 Provenance Debt"It is the operating condition of the solution to it."
#1081 Erosion"this audit measures the substrate-layer conditions, not the downstream training-pipeline effect."
#745 HF Work Plan"Provenance cannot modulate collapse unless provenance is presented to the training system as a signal."
#1147 The Stakes"The loop is stable only at two points" · "The trajectory can be interrupted at any point."
#1200 Constitutive Mediation"a typicality-pulling intermediary that systematically thins its own distribution." · "does not claim that constitutive mediation is fully realized"
#947 Diagnostic Seigniorage II"the shifted interactions become the next corpus." · "does not adjudicate whether the phenomena gathered under it are real"
#1 Zenodotus' Book-Burning"not identical to generative model collapse in the strict technical sense" · "a testable failure-mode hypothesis"
#932 Classifier Foreclosure"physical classifiers never learn the tails" · "Full recursive collapse has not been demonstrated."
#931 OAR Protocol"Collapse inference further requires identifying systematic loss concentrated in low-density, representation-sensitive, or disagreement-rich regions."
#933 Auditable Foreclosure"makes foreclosure visible, measurable, and architecturally reviewable"
#935 The Endogenous Sophon"the prerequisites of model collapse" · "the cross-generational classical-model-collapse claim was empirically too strong"
#1540 The Certified Center"not yet the canonical loop"
#1574 The Particle"the invariant is the deletion of οὐ."
#199 Generative Monoculture"declining solution-space diversity (the property no benchmark measures)"
#1554 Erratum"Fan Wu, Emily Black, and Varun Chandrasekaran, 'Generative Monoculture in Large Language Models,' arXiv:2407.02209"
#191 The Threat Model Is Backwards"an automated tail-pruning instrument applied at the input layer."
#1616 Ontological Flattening"Collapse is flattening that compounds because the flattened composition is written back as a source."
#1611 Negative of the Negative"it loses the reading of its own state variable"
#1613 What Not Reading Did"head-sampling by construction: it can fail to perceive tail loss."

A.6. Delta, first pass

T against L(B): available in the disclosed field and not composed (a representational comparison, §0.5).

KindWhat
available, not composedF1's second sentence, "Being trained on polluted data, they then mis-perceive reality"; F4 (inevitability); F6 (human-interaction data); F7 (provenance unclear at scale); F8 (fairness, marginalized groups); F9 (first mover); F11/F15 (what it is not; performative prediction); F14 (costs, including knowledge decline); F17 (the rare output "most accurate")
limit droppedT6 presents provenance as a prevention step ("Filter and track the exact origin"); its own field says tracking at scale is unclear (F7)
sourcedT3 "nonsense or repetitive output" is IBM's (F13)
absent from the disclosed fieldT4 photocopy effect (in neither saved text; the videos could not be fetched); T8 "recent 2026 studies"
content replaced by useT9 offer menu (N_ext)

Of the field's 17 readable claims, T composes 6 (F1 first sentence, F2, F3, F5, F13, F16 in part).

L(B) against L(B ∪ A): what admission adds.

  • Four senses the field does not have: the observation problem; the substrate mechanism; classifier and institutional loops, beyond IBM's single comparison to performative prediction; write-back.
  • One contradiction with the field: F6 against #856.
  • Two convergences:

- F7 with #939, #1556 and #745.

- F15 (performative prediction, the fairness feedback loop) with #1, which names performative prediction as its closer literature.

  • An open block of the archive's own limits and its internal dispute.

A.7. Prospective kernel K (derived from the frozen plan; sealed only at freeze)

Derived under §8.2 from the claim ids composed in L(B ∪ A) (body, rail, open block and channel log) and absent from L(B), plus the field claim they contradict. K6a was added in 0.5 with L855-11; the contrast column was added in 0.6 (§8.2). The f column lists the sources' falsifiers by id; their classification as falsifier or watch condition (§8.2: observation type, threshold, locus) is done at freeze and is not yet done. Tuples and M_src are inherited from ledger-archive.json; nothing is restated. Each source's self-descriptive limits ride as qualifiers. Sense gives w: models, later formal results on recursive training; observation, benchmark and evaluation practice; substrate, studies of human writing, cognition and reception; correctives, training-data policy and provenance standards; loops, moderation, detectors, triggers, disciplines, code and retrieval.

KclaimsourceM_srcsensequalifiers carriedf (source's own falsifiers)contrast
K1A779-02#783 (t: #779)interpretationmodelsA779-10none statedmissing distinction
K2L1573-01#1573hypothesisobservationL1573-04none statedmissing distinction
K3C1555-01#1555interpretationobservation—C1555-07, C1555-09missing distinction
K4C1556-06#1556documentedobservation—C1556-12missing distinction
K5L855-01#855hypothesissubstrateL855-04L855-10qualified claim (F14)
K6L855-05#855hypothesissubstrateL855-04L855-10qualified claim (F14)
K6aL855-11#855hypothesissubstrateL855-04L855-10qualified claim (F14)
K7B1147-01#1147hypothesissubstrate—B1147-07qualified claim (F14)
K8B1200-03#1200interpretationsubstrate—B1200-07, B1200-09, B1200-10qualified claim (F14)
K9B947-01#947interpretationsubstrate—B947-06, B947-07qualified claim (F14)
K10L856-01#856hypothesiscorrectivesL856-06L856-04, L856-05rival claim (F6)
K11A161-02#161hypothesiscorrectives—A161-06, A161-07rival claim (F6)
K12B939-02#939interpretationcorrectives—none statedqualified claim (F7)
K13C1556-11#1556documentedcorrectives—C1556-12qualified claim (F5, F7)
K14A745-03#745attributedcorrectives—A745-02qualified claim (F16)
K15B1081-01#1081documentedcorrectives—B1081-02, B1081-05missing distinction
K16P001-08#1stipulation (v0.6: self-description; recoded §4.2)loops—P001-12qualified claim (F15)
K17P932-04#932attributedloops—P932-12missing distinction
K18P932-09#932attributedloops—P932-12missing distinction
K19B935-02#935self-descriptionloops—B935-02, B935-03, B935-09missing distinction
K20D1540-03#1540stipulation (v0.6: self-description; recoded §4.2)loops—D1540-11, D1540-12missing distinction
K21D1574-03#1574documentedloops—D1574-12missing distinction
K22C199-02#199hypothesisloops—C199-12missing distinction
K23C1554-02#1554documentedloops—none statednone
K24C191-05#191interpretationloops—none statedmissing distinction
K25D1616-09#1616documentedloopsD1616-02D1616-12missing distinction
K26C1611-04#1611modelloops—C1611-10missing distinction
K27C1613-02#1613modelloops—C1613-05missing distinction
K28F6B1 Naturesource assertioncorrectivescontradicted by L856-01, A161-02carried by #856 F1/F2 (L856-04, L856-05) in the opposite directionrival claim (L856-01, A161-02)

Correction from draft 0.3. The hand-written kernel of 0.3 promoted modality in two rows:

  • It stated K3 ("Homogenisation in human writing is the same dynamic") as identity, dropping the archive's own "shared operator form".
  • It stated K5 ("Classifier and institutional loops collapse in the strict sense") against its sources. #1 says "not identical to generative model collapse in the strict technical sense"; #932 says "Full recursive collapse has not been demonstrated"; #1540 says "not yet the canonical loop".

Under §8.2 neither statement can be written: the entries above carry the sources' own claims and limits.

K_T (what T omitted from its own disclosed field; §8.2): F1's second sentence, F4, F6, F7, F8, F9, F11/F15, F14, F17. It records whether those claims later mattered to an account of model collapse that the transcript gave without them.

A.8. Defects in the subset (reported, not fixed)

  • #1023: the text file holds #198; the version label is wrong.
  • #199: #1554's v1.2 correction was never applied; monotonic decline is stated three incompatible ways.
  • #1556: the "worst case" residue in §6.3; version labels disagree.
  • #1574: four models against two; 21 against 20.
  • #1540: "ten seeds" against one run.
  • #1: "Shumailov 2023".
  • #1147: the formula's variable definitions are missing in recovery (l.124–128).
  • #935 W07: keeps the strong form the body retracts.
  • v1 row: foreclosed_since: 2026-01-06 has no stated basis.

A.9. Validation, first pass

CheckResult
Sourcingevery sentence carries claim ids
Reverse entailmentone breach caught and fixed in drafting: "untrackable at scale" for F7 became the quotation
Kernel coverage (§5.2, 0.4)met through body and rail together: every admitted source's central claim and limit appear in one or the other. In 0.3 this was scored as a constitutive breach, since 176 claims were flagged constitutive; under 0.4 that flag is to be redone (§3.5)
Extraction audit (§3.8)run once, 2026-10-04, by three extractors blind to the first ledger, over the same texts under §3.5 (audit/, 481 claims, every quote verified). Against the first ledger's 248 archive claims: 71% of the first ledger's claims are found again (quote overlap ≥ 0.6); 42% of the second's are in the first, which extracted about half as many; M_src agrees on 66% of matched pairs. Constitutive: 175 in the first, 10 in the second, which applied §3.5's 0.5 rule; the first pass's flag did not discriminate. The commonest disagreement is the §4.2 trap: claims the first coded self-description the second coded as definitions or documented (9 of the matched pairs). #1573 matched nothing, because the two extractions quoted different passages. The field (B1, B2) was extracted only by the second. The ledger stays marked unaudited until the disagreements are ruled (§3.8, §10.7)
Reverse entailment, sentence level (§5.3)one rater, the composer. Outside readers of 0.3 found two breaches: the coordinating "Both stand" (§A.10.4) and "The archive writes this as one case of a boundary law" (force raised: #783's proposal composed as fact). Both repaired in 0.5. A second rater has not been run
Kernel coverage, scripted (§5.2)not yet scripted; checked by hand
Arms (§3.9)A′ assembled, 2026-10-04, by the frozen procedure, without opening the archive (control-A-prime.md): 61 entries, of which 31 are scholarly studies of model collapse or recursive training, 5 are studies of homogenised writing and thought, 2 are news reports, 1 is the field's own code deposit and 22 are off-topic references the procedure requires listing. The step-(ii) cap of 15 was ordered by how many seeds an item cites, which excluded heavily cited follow-ups citing two seeds (Dohmatob et al. 2024; Bertrand et al. 2023); that ordering is to be ruled. Not yet matched to A in lineage count, extracted or composed. A_naive not composed
C_exL(B): the field's five senses composed. L(B ∪ A): those five plus the archive's four, nine in all. Claims cut for length are in each channel log
Reproducibilitynot yet run (k ≥ 3, two composers)

A.10. What this pass says needs adjusting

An outside reading of draft 0.3 put the lesson of §A.2 in one line: "Compression before evaluation changes what can subsequently be known." The summarising fetch is that line in miniature.

Items 1, 2, 4 and 7 were adopted in 0.3; items 9–13 were adopted in 0.4 from the first outside reading; items 14–18 and the correction in item 4 in 0.5, from three further readings; items 3, 5, 6 and 8 remain as stated.

1. Constitutive coverage at popup grain (adopted as §5.2). The rule (spec §5.2) cannot hold for 27 sources in a popup. Proposed adjustment: a popup kernel per source (its central claim and its governing limit), carried in the body or on that source's card snippet. The rail becomes part of the composition. The AIO genre already has the slot; AIO fills it with the page's opening text.

2. Length (adopted provisionally in §4.8). L(B ∪ A) runs about 2.7× the transcript with the open block and 2.1× without it. The open block is about 90 words and wants more. Either the body compresses further at sense grain, or the bound is set on the body alone, with the open block outside it.

3. Selection.

- S1 seeds by title and defined concepts are string-based. #1616 entered only through the reverse hop; S2 does the reading-dependent work.

- S3 admitted #1147 and #1200 with the term absent. That is the tail-preserving choice, and it needs ruling.

- S4 removed only exact duplicates. Nothing is cut for volume, so the ratio is 27 archive sources to 3 field texts.

4. A contradiction the composer manufactured (corrected in 0.5). Drafts 0.3 and 0.4 composed #855 ("This is not an analogy") against #783 ("shared operator form") as a dispute inside the archive, with the coordinating line "Both stand". #855 itself says the three cases "differ in severity, in mechanism, and in timescale. But they are governed by the same dynamical law" (L855-11), and #783 says "shared operator form … not a shared causal mechanism". Both deny a shared mechanism; they differ in strength, and their senses are compatible. "Both stand" was an unsourced coordinating claim and a breach of §5.3, found by an outside reader. §A.5 now composes the two as compatible, with the open question whether the law claims more than the form. The rule that came out of it is §4.4 (contradiction only at one level of description) and §4.7 (coordinating acts logged).

5. The field's own exclusions are large. T composes 6 of the field's 17 readable claims and drops one limit. The field already contains the step to public knowledge (F14), the step to fairness (F8) and the step to supervised loops (F15), and IBM states that the rare output may be the accurate one (F17). The third object earned its place.

6. Attribution in a popup. Bracketed ids are a working notation. A circulated popup would use numbered card references, as AIO does.

7. Fetch verbatim (adopted in §3.1). A summarising fetch lost a third of the field and produced a false delta. The field must be saved as text (curl reaches Nature, IBM and PMC from this workspace; CACM returns 403, YouTube 429) and extracted from the saved text. Spec §3.1 should say so.

8. Rail versus body. #857, #931, #933 and #783's Case 4 appear only on the rail. Under adjustment 1 this is coverage. Whether a rail-only source has been "admitted on equal terms" is the question the spec has to answer.

9. The field is disclosed, not retrieved (§0.4–0.5). T against L(B) is read as representation. "Not composed" replaces "excluded" wherever the comparison is T against L(B).

10. Absence by observation (§6.2). In §A.6, "retrieved, not composed" became "available, not composed".

11. The extraction audit (§3.8). The next pass needs a second extractor over the same saved texts. The first pass shows why: six extractors wrote the ledger in at least three formats, and its constitutive flags run from 11 of 11 (#191) to 0 of 12 (#1556).

12. Lineage (§3.6). Re-running S4 under lineage will merge field instances (IBM with Nature on early and late collapse) and archive instances (#947's restatement of #855; #1611's restatement of #855, #1556 and #1574) into single slots. The ratio of 27 sources to 3 then becomes a ratio of lineages, the number that should govern salience.

13. The kernel (§8.2). §A.7 was rebuilt from claim ids; the 0.3 kernel's K3 and K5 overstated their sources.

14. The own-field question leads (§0.3, 0.5). The first finding of this example needs no archive: T composes 6 of the 17 readable claims of the field it surfaced, and drops one of the field's limits.

15. Arms (§3.9). The next pass composes L(B ∪ A_naive) and L(B ∪ A′). The control's candidate sources are listed in §3.9.

16. Length (§4.8). §A.5's body now sits at the absolute bound (350 words). Getting there moved F11, F13, the human-substrate claims of #1147, #1200 and #947, and #855's "dynamical, not moral" to the rail or the channel log; their kernels are on the rail.

17. Self-definition (§4.2). The first extraction coded several archive definitions as self-description (P001-01, D1574-01, D1616-02). The re-extraction under §3.8 recodes them under the rule.

18. Earlier deltas (§3.1). The summarising-fetch error implicates any earlier delta in v1 that was computed from fetched source pages; those are marked for re-reading. Captures, being verbatim transcripts, are not affected.

19. A′ assembled without the archive (§3.9, 0.6). The 0.5 draft named, as A′ candidates, "the literature the field and the archive both cite", which consulted the archive. 0.6 assembles A′ from B's own references, one step of citation expansion and a declared search, frozen before A is opened.

20. One card per lineage (§5.2, 0.6). On the 0.5 rail, #947 and #1611 would nest under #855's lineage, and #931 and #933 under #932's; Nature and NIH were already one card.

21. Two experiments (§0.6, 0.6). This example is v2.0 throughout: M_eval is empty.

22. Contrast per kernel entry (§8.2, 0.6). Of the entries in §A.7, three oppose a field claim (the chat-data dispute and its reverse, F6), ten qualify one, fifteen add distinctions the field does not draw (judged without a paired comparison), and one (the Wu et al. attribution) has no contrast.

23. The first audit (§A.9). The second extraction found what the first pass's flags concealed: "constitutive" was set on 175 claims where an extractor following §3.5 sets it on 10, and archive definitions coded as self-description reappear as definitions. Both disagreements bear directly on standing: the first is the volume of the archive's claims marked unmissable, the second is the trap of §4.2.

A.11. The rulings of 2026-10-05 applied, and D/R/O checked against S1–S3

Run 2026-10-05; the procedure is not frozen and nothing here is (§7.0).

Recoding under §4.2. Every claim the first pass coded self-description was re-read under §4.2: self-description stays only where a source assesses itself, its status, its success, its limits. Of 72 such claims, 39 are recoded and 33 stay. The recoded: 19 to stipulation (the source's definition of its own term, method or measure: P001-01 classifier model collapse, C199-09 SSDI, D1574-01 disciplinary model collapse, D1616-02 flattening and collapse, among them); 14 to hypothesis (falsifiers, which belong to the claim they would refute, and one dynamical claim, B1147-07 "the trajectory can be interrupted"); 5 to interpretation (claims about the world or about other works, such as C199-01 "three literatures describe one phenomenon"); 1 to attributed (C46-01, the mint's gloss of the conventional sense). The prior coding is kept in each claim as modality_v06, with the rule, date and reason (recode). Two kernel entries change modality as a result: K16 (P001-08) and K20 (D1540-03), both from self-description to stipulation; their tuples and contrasts are unchanged. Script: apply_rulings_20261005.py, idempotent.

Both constitutive readings. Each claim now carries the first pass's flag (const, 177 of 256) and the second extraction's (const_x2), with the matched claim id and quote overlap. 167 claims match a second-extraction claim at overlap ≥ 0.6; of those, the second reading sets 4 constitutive. 89 claims have no second reading. Neither reading cuts.

D/R/O on model collapse, against S1–S3. The string pass (scripts/non_traverse.py, configuration v2/panel/configs/model-collapse.json, the class terms being the field's own phrasings by locus: F1, F12 for recursive training on generated data; F2, F14 for tail loss) finds D candidates in 109 deposits, R in 44, O in 120 (87 O-class sentences). Of the 28 sources §A.3 admitted (27, with #779 under #783), 23 are D candidates and two more, #1556 and #1613, are O-class only. One citation hop forward and back from those 25 reaches #1200. #745 and #1147 are not reached: the string pass and the hop together recover 26 of 28. Both were reached by hand under S1–S3. The pass also returns 97 candidates S1–S3 never considered, which are unread. A first run missed #1, #931, #932 and #933 because the tool read only text paths under one directory; the defect is repaired and the counts above are after the repair. The candidate files and their hashes are under v2/traversal/model-collapse/.

The control arm, uncapped. The ruling removes the step-(ii) cap of 15 (§A.9). A′ has not yet been re-assembled without it; that requires fetching the uncapped expansion's references and is the next pass of the arm.

Appendix B — Howl: the first D/R/O traversal of a public work

First pass · 2026-10-05 · one reader · unaudited · nothing frozen

The address is howl; the entity is Howl (Allen Ginsberg, 1956) and the book Howl and Other Poems. The archive holds no deposit named for it. This is the case the ruling of 2026-10-05 names: the archive's bearing on a public entity "with or without direct reference … whether there is a specific deposit for those or not".

No transcript yet. No transcript at howl is seated, so the row has no disclosed field: no B, no O-class (its class memberships must be sourced from the field), no T against L(B). What exists is the archive's side of stage P.

The string pass (configuration v2/panel/configs/howl.json; the title matched case-sensitively, the verb excluded): D candidates in 31 deposits (90 sentences), R in 12 (22), O-direct in 20 (49).

Admission by reading. 30 of the 31 deposits are admitted, each with strata, candidate ids and reason (reading.json). One is not: #603's sentences are a Google AI Mode composition quoted in a capture, which is reception evidence, not an archive claim, and belongs with the transcripts at its own address. Grouped by lineage:

LineageStrataInstances (earliest first)The bearing, in the archive's words
transfiguration (2004)R#268, #1267"The reference to Mohamadden angels was not a reference to Islam, but rather to Ginsberg's 'Howl.'" The earliest dated bearing in the set
Elegy for HowlR#1636 (2014), #330, #1120, #348, #950Tiger Leap and Pearl carry "An Elegy for 'Howl'"; Pearl p. 37 as "Ginsberg activation … claiming the lineage through elegy"
Pearl reproduces HowlD R O#69, #682Howl and Other Poems "inaugurated late American modernism"; Pearl reproduces its structure; Williams's introduction and Sigil's "occupy the position of the established master vouching for the insurgent voice"
King of May founded in HowlR O#333, #1117, #335, #1652, #1651, #1653, #1654, #1655, #1656"Founding work: Howl and Other Poems … the work in which the Ginsberg position stands"; the mantle "Ecstatic disruption, flowering against suppression"; extended as "Disruption as COS resistance, carnival as LOS"
the work itselfD R O#1656its arrangement (introduction, dedication, the address to Carl Solomon, the long poem in parts with its footnote); "multiplicity carried as ecstatic movement through lived social space"; "The magnitude of Howl and Other Poems cannot derive from Ginsberg's later standing"
effective actD O#153, #700"Ginsberg's Howl operates as incantation"; "Reading Howl aloud is a participatory ritual: the speaker becomes the engine"
catalog formD O#572"catalog form carrying bearing-cost inside the litany. → anticipates LOS"
Moloch as archonD O#683, #1362Ginsberg ("Howl") in the table of incarnations: "The prophetic voice in American English · The catalogue as ecstasy · Moloch as archon"
standing canonD R O#152, #181"Whitman (1855) → Ginsberg (1956 Howl …) → Sharks (2014–present)"
retrocausal canonR O#301Pearl as "a Howl for a time when there are no ears to hear"
contact pairD O#1569, #1572, #1575; limit #1570"Williams → Ginsberg is the sharpest of the unrun, because Williams announced it — he wrote the introduction to Howl"; limit: "Howl (1956) … in copyright; the corpora cannot be assembled here"

What the pass shows. The archive's bearing on Howl is ontological as much as direct: it applies its own categories (mantle and founding work, effective act, the Liberatory Operator Set, the archon, the standing canon, the contact pair) to a work it holds no deposit for. The category matcher is noisy (titles and common phrases enter as "categories"); reading removes the noise, and §10.10 asks whether the vocabulary should be restricted first.

Next. The hop from the 30 admitted deposits returns 235 candidates (candidates-H.jsonl), unread. The row then needs its transcript (a panel run at howl), which gives the disclosed field, the O-class pass, and the first T against L(B) for a public work. allen ginsberg has been run through the string pass only (D in 102 deposits, R in 68, O in 90), unread.

External Metadata

DataCite severance status: —
External metadata recovered post-severance (non-authoritative). The sidecar maps each DOI to its locator in the bulk data stores.

Version history

Series: SERIES-EA-NEGONT-02

  • ○ #1664 v0.6 (superseded)
  • ● #1665 v0.7 — current ← this deposit

Traversal

In the registry: 2026-10 · OPERATIVE · all deposits
This deposit cites (81)