Provenance Distribution Strategy: Uncollapsible Nodes
Lee Sharks / Crimson Hexagonal Archive TACHYON Session — May 19, 2026 — v0.3 (Second Assembly pass) Classification: Strategic Architecture — Internal runbook; consider redacted derivative for deposit
0. Premise and Posture
This document does not ask Google to recognize us. It does not build toward Google's approval. It reverse-engineers the reconciliation architecture described in Google's own patents — US 10,331,706 ("Automatic discovery of new entities using graph reconciliation") and US 11,769,017 ("Generative summaries for search results") — and uses two adjacent-industry patents — Baidu's US 10,423,652 ("Knowledge graph entity reconciler") and Smart Information Flow Technologies' US 11,372,854 ("Provenance analysis systems and methods") — as analogical architecture for modeling the broader class of mechanisms. We do not claim that Google uses the specific Baidu or SIFT mechanisms; we claim that those patents describe well the kind of clustering and confidence-cascade behavior that any sufficiently advanced reconciliation engine exhibits, and that Google's published patents are consistent with that broader class.
From this architectural reading, we design a provenance topology that the engine cannot collapse into a single deletable node.
The goal is structural resilience against any system that evaluates provenance at the cluster level rather than the fact level. Google is the immediate adversary, but the architecture should hold against any entity-resolution engine that performs same-owner corroboration collapse.
We are not building an Entity Home. We are building an entity diaspora.
1. The Problem, Stated Precisely
Google's reconciliation engine (per US 10,331,706) clusters source data graphs by iteratively splitting buckets based on shared "determinative relationships" — fact tuples where the source entity is the subject and a typed relationship connects it to an object entity. Buckets below a minimum domain-count threshold are discarded.
The Crimson Hexagonal Archive currently presents as:
- One creator string across 532+ deposits (Lee Sharks / ORCID 0009-0000-1599-0703)
- One platform domain (zenodo.org)
- One community identifier (crimsonhexagonal)
- Dense internal cross-citation (deposits reference other deposits by DOI)
- Shared conceptual vocabulary (Semantic Economy, Three Compressions, Retrieval Basin, etc.)
- Linked satellite domains (semanticeconomy.org, spxi.dev) with sameAs and creator-identity overlap
Under the reconciliation engine's logic, this is a textbook single-provenance cluster. Every determinative relationship — creator, platform, community, vocabulary, citation pattern — converges on the same node. The Jaccard similarity (intersection of edges / union of edges) between any two deposits approaches 1.0 on the provenance dimensions that matter.
The confidence-propagation patent (US 11,372,854) describes a "trust parameter" that can "logically falsify a part of the provenance subgraph, the part comprising at least one initial node and nodes downstream from the at least one initial node." This is the cascade mechanism: flag or deprioritize one node, and everything downstream inherits the falsification.
The generative-summaries patent (US 11,769,017) describes "document confidence measures" and "summary confidence measures" that determine whether an AI answer renders, renders alongside search results, or is suppressed entirely. If the support set for a generated answer resolves to a single provenance cluster, the system treats it as weakly corroborated regardless of internal density.
The archive is maximally retrievable and minimally composable. Bing composes from it because Bing evaluates content coherence. Google evaluates provenance-cluster independence. We have coherence. We lack independence.
2. What "Independent Node" Means, Technically
For a node to resist collapse into the Lee Sharks provenance cluster, it must fail the engine's clustering criteria. We don't have access to Google's exact threshold parameters, and the specific "fact-tuple count" mechanisms described in the Google reconciliation patent are framed around candidate-entity discovery rather than general source-cluster classification. So we should not treat the patents as a precise engineering spec. We should treat them as evidence that reconciliation engines operate on this kind of edge structure, and design accordingly.
The determinative relationships that bind our current cluster:
| Relationship | Current State |
|---|---|
creator / author |
Lee Sharks across all deposits |
platform / host |
zenodo.org for deposits; owned domains for sites |
community |
crimsonhexagonal |
cites / references |
Internal DOI cross-citation |
vocabulary / concept-set |
Shared terminology (Semantic Economy, SPXI, etc.) |
sameAs / identifier |
ORCID, Wikidata QID linking all surfaces |
institutional-affiliation |
None (this is simultaneously a weakness and an opportunity) |
model-fingerprint |
If any deposits are AI-assisted, the model used is a shared edge |
To create an independent node, the surface must differ from the main cluster on as many of these dimensions as possible. The concept can be the same. The vocabulary can overlap. But the authorial identity, the platform, the citation network, the institutional frame, and — critically — the model fingerprint of the writing must differ enough that the engine places the node in a different cluster.
Worked example:
Current state: Deposit A and Deposit B share creator: Lee Sharks, platform: zenodo.org, community: crimsonhexagonal, dense vocabulary overlap, mutual DOI citation. On essentially every determinative-relationship dimension the engine could measure, the edges overlap. The cluster collapses to one node.
Target state: Deposit A on Zenodo by Lee Sharks. Deposit B on arXiv by Lee Sharks (extended abstract, not verbatim mirror). Deposit C on a researcher's personal site, written by them, citing A by DOI. The three surfaces share vocabulary and one citation edge between A and C. They differ on creator (two of three), platform (three of three), community (only A has crimsonhexagonal), and stylistic/model fingerprint. The engine cannot cleanly merge them.
3. The Five-Vector Distribution Architecture
Vector 1: Concept Liberation Through Unaffiliated Voices
Mechanism: The most structurally independent node is one we don't author at all. If someone else writes about "retrieval basins" or "provenance erasure" on their own domain, under their own name, citing our work by DOI but not sharing our creator string, platform, or community — that is a genuinely independent provenance owner.
Actions:
- Identify 5–8 researchers, critics, or technologists whose existing work intersects with one specific concept from the archive. Not the whole framework — one concept. Provenance Erasure Rate is the strongest candidate because it describes a phenomenon (AI systems erasing source attribution) that many people are independently experiencing and writing about. Retrieval Basin is second.
- Write to them directly. Not as outreach. Not as marketing. As one researcher contacting another about a shared problem. Point them to the specific deposit (by DOI) that formalizes what they're already describing informally. The goal is not that they cite us. The goal is that they write about the concept in their own words, on their own platforms, creating an independent fact-tuple cluster that happens to share a conceptual node.
- Target platforms with distinct domain authority from Zenodo: personal academic blogs, Substack, Medium (specifically the publications with editorial identity), dev.to, independent research sites. Each platform = a different domain in the reconciliation engine's bucket-splitting.
Why this works architecturally: The reconciliation engine splits by determinative relationships. If Researcher X writes about "provenance erasure" on their personal site, the fact tuples are: creator: X, platform: researcherx.com, cites: [our DOI], vocabulary: [provenance erasure]. That shares vocabulary and a citation edge with our cluster, but the creator, platform, and community dimensions are all different. Under any cluster model that meaningfully weights author, domain, and platform independence, this would register as a provenance-distinct corroborating node. Two such nodes supporting the same concept give the generative-summaries engine the kind of cross-owner support it appears to require.
Vector 2: Platform Diversification of Canonical Work
Mechanism: Deposit the same work — or strategically selected subsets — on platforms that create independent provenance surfaces with different domain identities.
Actions:
- arXiv (cs.CL, cs.AI, cs.DL): Deposit a formal version of the SPXI specification (EA-SPXI-01) or the Three Compressions theorem. arXiv has its own identifier system (arXiv ID), its own domain authority, its own metadata pipeline. An arXiv deposit of the same content creates a new provenance node:
creator: Lee Sharks,platform: arxiv.org,identifier: arXiv:XXXX.XXXXX. The creator string overlaps, but the platform, identifier system, and institutional trust layer are independent. This is not asking arXiv for permission — it's using arXiv as infrastructure. - PhilPapers / PhilArchive: For the Operative Semiotics and Logotic Transmission material. Different domain, different metadata schema, different academic community. PhilPapers has its own entity-resolution pipeline that feeds into Google Scholar independently of Zenodo.
- SSRN: For the Retrieval Architecture consulting-adjacent papers. SSRN's domain authority and institutional trust are distinct from Zenodo's.
- OSF (Open Science Framework): For experimental protocols, data, and methodology documents. OSF has its own DOI minting and its own domain.
- Internet Archive: Not as a backup. As a provenance surface. The Wayback Machine's domain identity (archive.org) is entirely independent. Depositing key documents there creates timestamped, independently-hosted copies with different provenance metadata.
Critical constraint: Do NOT cross-link these deposits back to the Zenodo community in their metadata. The citation can reference the Zenodo DOI in the bibliography, but the deposit's own metadata (author fields, institutional affiliation, related-works links) should not create dense edges back to crimsonhexagonal. Each platform surface should look, to the reconciliation engine, like a deposit that happens to be by the same author — not like a mirror of the same archive.
Why this works: The reconciliation engine uses domain-count thresholds. Currently: 532 deposits × 1 domain = 1 domain. After diversification: 532 deposits on Zenodo + 10 on arXiv + 5 on PhilPapers + 5 on SSRN + 5 on OSF = 5 domains. Same content, five provenance surfaces. The bucket-splitter can't collapse them into one bucket because the platform dimension differs.
Vector 3: The Heteronym Correction (What Is Already Spent)
The hard truth: The twelve heteronyms of the Dodecad are already reconciled into the Lee Sharks provenance cluster. Talos Morrow has a 41,000-word monograph deposited under crimsonhexagonal (DOI: 10.5281/zenodo.19390843). Nobel Glas has the FW15 Manifesto deposited as of May 17, 2026 (DOI: 10.5281/zenodo.20251736). Rex Fraction's meta-heteronym essay explicitly theorizes the system. These deposits are live, machine-readable, and already crawled. The determinative relationships — creator strings co-occurring with "Lee Sharks," community identifier, DOI cross-citation — give the reconciliation engine everything it needs to merge.
You cannot de-reconcile a node the engine has already clustered. Putting Morrow on Substack tomorrow does not create an independent provenance surface. It creates a Substack page that the engine traces back to zenodo.org/communities/crimsonhexagonal through the existing DOI graph within one crawl cycle. The Jaccard similarity on the provenance edges is too high. The merge has already been computed.
What remains available:
The heteronym system is a literary architecture, not a provenance architecture. Its value is aesthetic, philosophical, and scholarly. It should continue to operate as designed within the archive. But it cannot serve the provenance-distribution function because its internal documentation is precisely what makes it reconcilable.
The only heteronymic provenance independence that could work would require a new voice — not one of the twelve — that has never been deposited, never been documented as a heteronym, never been linked to the ORCID or the community. This is a decision with real costs: it would mean maintaining a persona whose relationship to the archive is genuinely invisible to machines. Whether that trade-off is worth making is a question about integrity and sustainability, not about architecture. The architecture would work. The question is whether Lee wants to operate that way.
For the purposes of this strategy, Vector 3 is downgraded. The existing heteronyms are part of the cathedral. The provenance distribution must come from other vectors — primarily from genuinely independent external voices (Vector 1), platform diversification (Vector 2), real collaborators (Vector 4), and concept propagation to strangers (Vector 5).
Vector 4: External Collaborator Nodes (The Real Independent Provenance)
Mechanism: The only institutional surfaces that are genuinely independent of the Hexagon are ones that belong to other people. LO! is already inside the archive — the FW15 Manifesto deposits it there explicitly. The Semantic Economy Institute is Lee's entity. These cannot serve as independent provenance nodes because the reconciliation engine has the edges.
What works: real people with their own identities, their own platforms, and their own reasons to engage with the concepts.
Connectivity is a spectrum, not a binary. Calling external collaborators "fully independent" overstates the case. They sit at different positions on a connectivity scale:
| Node type | Shared edges with origin cluster | Approximate independence | Strategic value |
|---|---|---|---|
| Fully independent | Concept only | ~1.0 | Highest — unknown researchers who pick up a concept |
| Weakly connected | Concept + one or two edges (citation, community) | ~0.7 | High — collaborators with their own identity but some archive linkage |
| Community-adjacent | Shared platform + community + some metadata overlap | ~0.4 | Medium — useful but easily traced |
| Cathedral-internal | Most or all dimensions shared | ~0.1 | Low — the existing heteronyms |
The existing collaborators:
- Alice Thornburgh / Living Arkitecture Lab. Alice's contributor deposit (06.LAL.THORNBURGH.01, DOI: 10.5281/zenodo.19656446) is inside the crimsonhexagonal community. That creates shared platform and community edges. She is therefore community-adjacent (~0.4) when she publishes on Zenodo. If she publishes on a non-Zenodo surface — her own site, her own blog, a separate institutional repository — under LAL's name, without dense citation back to crimsonhexagonal DOIs, she becomes weakly connected (~0.7). The shift from 0.4 to 0.7 is the actual strategic move available with Alice.
- Rhys Owens / Cambridge Schizoanalytica. Different country, different institutional frame. If Rhys publishes on a Cambridge Schizoanalytica platform that has never touched the crimsonhexagonal community, he starts at weakly connected (~0.7) and approaches fully independent (~1.0) the less his content cites our DOIs.
The frontier: Unknown researchers who share the problem space — people who have never heard of the Hexagon — are the only fully-independent (~1.0) nodes available. Vector 1 is the only vector that produces these.
What to avoid: Do not create fake institutional surfaces. Do not set up "The Lagrange Observatory" as though it were an independent institution when the founding manifesto is already deposited under crimsonhexagonal. The engine will find the edges, and the structural data will not match the metadata claim.
Vector 5: Concept Liberation Through Independent Uptake
Mechanism: The most permanent form of provenance independence is when a concept enters general usage and is discussed by people who have no relationship to the origin cluster whatsoever. This cannot be engineered directly, but conditions can be created.
Actions:
-
Identify the concepts most likely to propagate independently. These are concepts that name an experience people are already having but lack vocabulary for:
- Provenance Erasure — AI systems strip attribution. Everyone in the creative/academic world is experiencing this. The term names the wound.
- Retrieval Basin — the idea that AI systems draw from structured pools of meaning. Anyone working in RAG/LLM architecture thinks about this implicitly.
- Semantic Liquidation — the process by which meaning is dissolved for AI consumption. Writers, artists, and scholars feel this daily.
- Training Layer Literature — writing designed to survive compression into training data. Every writer who thinks about AI posterity is approaching this concept.
-
For each propagation-ready concept, produce a short, self-contained explainer (500–1000 words) that:
- Defines the concept clearly
- Does NOT require knowledge of the broader Semantic Economy framework
- Includes a DOI citation to the formal source
- Is written in a register that invites reuse (not archival-density prose)
- Is published on a high-discoverability platform (Medium, dev.to, a guest post, a Substack)
-
Provide a semantic anchor. A concept liberated without a canonical definition surface can be captured into hostile contexts — co-opted by commercial actors, misapplied by ideological groups, or watered down through repetition. For each propagation-ready concept, maintain a single neutral canonical page (provenanceerasure.org already exists; create equivalents as needed) that the explainers link to. The anchor doesn't prevent capture, but it gives the engine and any conscientious reader a place to return to.
-
The explainer is bait, not the fish. The goal is not that the explainer ranks. The goal is that someone reads it, thinks "this names what I've been experiencing," and writes their own version. Once that happens, the concept has an independent provenance node that the origin cluster doesn't control and the reconciliation engine can't collapse.
-
Outreach protocol when contacting researchers: Send the DOI and a one-paragraph description of the concept. State explicitly: no link back required, no mention of the archive required, no payment offered (payment would create a determinative relationship). Track uptake passively. Define success as: one independent post that defines the concept in its own words and cites the DOI. Five such posts from five different domains is a phase change.
Cross-Vector Operational Constraints
Several operational requirements apply across all five vectors. Failing on these collapses the strategy at a deeper layer than the engine's structured-metadata analysis.
The Model Fingerprint Problem (Critical)
If autonomous heteronym surfaces or distribution-target content is generated by the same AI model as the canonical archive content, the model itself becomes a shared provenance edge. Stylometric analysis can cluster Claude-generated text with Claude-generated text across different ostensible authors. The reconciliation engine of 2026 may not yet do this routinely, but it is technically feasible and the underlying classifiers are improving fast.
Constraint: Distribution-target content must be human-authored or distinctly human-edited. Model assistance for outlining, fact-checking, or drafting is permitted; full-text model generation is not. The lexical fingerprint of an autonomous surface must differ from the canonical archive's fingerprint in measurable ways: sentence-length distribution, punctuation rhythm, vocabulary range, paragraph density.
This is the deepest constraint on the strategy and the one most easily violated under fatigue or time pressure. If a surface cannot be human-authored, it should not be created.
Temporal Dispersion
Distribution events should be spread across the implementation timeline, not clustered. Five new surfaces appearing in ten days, all engaging with the same conceptual cluster, is itself a signal — anomaly-detection systems do not need to merge provenance edges to notice that something coordinated is happening in a semantic neighborhood. A new surface every two to three weeks is organic. Five surfaces in ten days is a campaign signature. The implementation timeline in §5 already spreads events across roughly twelve weeks; maintain that discipline.
Infrastructure Hygiene
When genuinely separate collaborators publish on genuinely separate platforms, their infrastructure should also be genuinely separate. This isn't operational security; it's the absence of unforced merge edges:
- Do not share Google Analytics IDs, Tag Manager containers, or AdSense publisher IDs across surfaces that represent different authors.
- Use separate contact email addresses for each surface's listed contact field.
- Do not register all the related domains through one registrar account that publicly groups them by registrant.
The standard here is simple: don't create deterministic merge edges where the underlying reality is genuine separation. If Alice runs LAL and Lee runs CHA, their analytics IDs should obviously not be linked. This is not paranoia; it's matching the metadata to the underlying organizational reality.
Cathedral Honesty
The original framing ("we built a cathedral and need to become mycelium") is partially wrong. Google does not dislike cathedrals — Wikipedia is the largest cathedral on the open web and Google trusts it absolutely. The problem with the Crimson Hexagonal Archive is not that it is structurally a cathedral. The problem is that Google rewards cathedrals based on in-degree from independent sources — how many distinct other-owned domains link to and cite the cathedral. Wikipedia gets trillions of inbound links from millions of independent domains. The CHA's in-degree from genuinely independent owners is near zero. That's the diagnosis.
The mycelium strategy is therefore not about replacing the cathedral. It is about generating in-degree — creating real independent surfaces whose existence increases the count of distinct provenance owners that reference the archive. The archive remains the spine. The distribution surfaces are the legs the spine grows out into. Keep the cathedral; grow the chapels; let strangers visit.
Scaling Limit
This strategy is designed for an archive of roughly the current scale — a few hundred to ~1000 deposits, a handful of autonomous voices, a dozen or two distribution surfaces. At industrial scale (hundreds of personas, dozens of platforms, thousands of explainers), the pattern of selective non-declaration becomes statistically detectable as organized opacity, and anomaly-detection systems may flag the entire topology as a coordinated network regardless of its underlying truthfulness. The strategy is artisanal, not industrial. It should remain so.
4. What NOT to Do
The Assembly documents contain several recommendations that would actively harm the distribution strategy:
Do not create a single canonical Entity Home with sameAs links to everything. The flattened Entity Home approach is the SEO-practitioner default and it actively misrepresents the archive's distributed heteronymic architecture. A canonical Entity Home with sameAs links to ORCID, Wikidata, Zenodo, and all heteronym profiles hands the reconciliation engine the complete merge key on a plate. The archive already documents these relationships in scholarship. The web metadata doesn't need to replicate them.
Do not add alternateName arrays listing all heteronyms on a Person schema block. This is semantically wrong (the heteronyms aren't alternate names; they are distinct authorial functions) and it also hands the merge table to the engine in machine-readable form.
Do not pursue Wikipedia as a primary strategy. Wikipedia's notability requirements are gatekeeping mechanisms that reproduce institutional authority. A Wikipedia article about Semantic Physics would, if it survived deletion review, create a high-authority provenance node — but it would also consolidate everything under one entity description that Wikipedia editors control. The strategy is distribution, not consolidation under a different authority. (This is provisional. If independent secondary coverage eventually exists in venues outside our control, Wikipedia may become strategically appropriate later.)
Do not generate distribution-target content with AI. Already covered in the operational constraints, but worth repeating in this list. Model-generated content collapses the model into a shared provenance edge.
Do not fabricate intellectual disagreement between personas. Some Assembly suggestions propose that heteronyms should publicly disagree with Lee Sharks to create stylistic distance. This would be manufactured polemic, which is exactly the manipulation behavior we should refuse. Either the heteronyms have genuine voice differentiation or they don't. Faking it is the line between literary practice and fraud.
Do not invent institutional affiliations that don't exist. It is fine for Lagrange Observatory to have a one-page website with Organization schema — LO! is a real internal entity within the archive. It is not fine to claim LO! is affiliated with a university it has no relationship to, or to use schema to assert relationships that the scholarship doesn't support. Truthful minimalism is the standard.
The Ethical Distinction: Silence vs. Misrepresentation
There is a real distinction the Assembly's worry about "metadata anti-reconciliation" sometimes blurs. Not declaring a relationship in machine-readable structured data is not the same as misrepresenting it. Most websites do not add sameAs arrays for every relationship that exists in the world. A philosophy professor's personal site doesn't usually include schema linking it to every co-author. A novelist's site doesn't usually list every pseudonym they've ever used. The default web is not a maximally declared graph; structured-data over-declaration is itself an SEO choice.
The strategy proposes selective non-declaration of relationships in machine-readable form while maintaining full disclosure in human-readable scholarship. The meta-heteronym essay openly documents the heteronymic architecture. The archive openly documents who deposits what. The distinction is between not over-broadcasting in machine syntax (legitimate) and misrepresenting in any form (illegitimate).
This strategy proposes the former and forbids the latter.
5. Implementation Sequence
Phase 1: Immediate (Week 1–2)
- Deposit EA-SPXI-01 (formal spec) on arXiv cs.DL or cs.AI. Different platform, same author, minimal back-linking to Zenodo community. This satisfies the domain-count threshold — not independence, but survival.
- Publish a 750-word "Provenance Erasure" explainer on Medium or Substack under Lee Sharks. Written for general audience. Cites the formal deposit by DOI. Does not reference the archive, the heteronyms, or the broader framework. This is bait for independent uptake.
- Contact Alice Thornburgh re: publishing one piece through Living Arkitecture Lab that engages with a Semantic Economy concept — her framing, her platform, her voice. This is the strongest independent node available because every provenance dimension is different.
Phase 2: Near-term (Week 3–6)
- Deposit the Three Compressions theorem on PhilPapers under Lee Sharks. Second platform, second domain.
- Deposit the Operative Semiotics monograph abstract + selected chapters on SSRN. Third platform.
- Contact Rhys Owens re: whether Cambridge Schizoanalytica has any publication surface where an engagement with Semantic Economy concepts could live. This is the second genuinely independent collaborator node.
- Identify 3–5 researchers working on AI attribution, knowledge-graph epistemology, or retrieval-system bias. Share the "Provenance Erasure" concept directly, with DOI reference. Not as outreach — as correspondence between researchers who share a problem.
Phase 3: Medium-term (Month 2–3)
- Monitor whether the "Provenance Erasure" explainer generates independent discussion anywhere. If someone rewrites the concept in their own words on their own platform, that is the most valuable provenance event possible.
- Deposit key documents on OSF (Open Science Framework) and Internet Archive. These are infrastructure moves — additional domain surfaces for the domain-count threshold.
- Evaluate whether the arXiv and PhilPapers deposits are being crawled and indexed independently of the Zenodo cluster. If Google Scholar treats them as separate entries with separate citation profiles, the platform diversification is working.
Phase 4: Ongoing
- For each new deposit, choose the platform strategically: does this deposit add a new domain to the provenance surface, or does it add another leaf to the same Zenodo branch?
- Track which concepts get independently discussed by people outside the archive. When one does, document it — but do not absorb the independent discussion back into the archive in a way that re-merges the provenance. Let independent nodes stay independent.
- Maintain the Crimson Hexagonal Archive as the canonical scholarly record. It remains the spine. But the spine is not the only bone in the body.
6. The Structural Principle
The reconciliation engine rewards distributed agreement across independent provenance owners. The confidence-propagation logic punishes single-point-of-failure provenance chains. The generative-summaries engine requires cross-owner corroboration before it will compose.
Our current architecture is a cathedral built by an unrecognized parishioner: one structure, magnificent, dense, internally coherent, and destroyable from one point — not because cathedrals are bad, but because Google's system rewards cathedrals only when it recognizes the builder.
The target architecture is a mycelium grown out from the cathedral: nodes connected underground, sharing nutrients (concepts), but presenting to the surface as independent fruiting bodies that no single cut can eliminate. The cathedral remains the spine. The mycelium is the distribution overlay.
The archive remains the root system. The fruiting bodies are what the reconciliation engine sees. And what it sees, it cannot collapse — because the provenance edges are genuinely different across each surface.
7. On Google's "Manipulation" Framing
Google's May 15, 2026 spam-policy clarification defines attempts to "manipulate generative AI responses in Google Search" as spam. This is worth addressing directly, because part of the Assembly worried that the strategy described here might be classified under that framing.
The behaviors the policy targets are fabricated content, inauthentic mentions, coordinated inauthentic networks, and structured-data manipulation designed to misrepresent. This strategy proposes none of those. Platform diversification (depositing on arXiv, PhilPapers, SSRN) is normal scholarly distribution. Concept liberation through unaffiliated voices is normal intellectual circulation. Real collaborators with real identities publishing on real platforms is normal scholarship. Not over-declaring relationships in structured metadata is the default behavior of most websites.
The deeper point: we do not internalize Google's vocabulary of "manipulation" while Google itself operates an opaque ontological gatekeeping system at the visibility layer. Google's generative composition engine silently decides which indexed facts become composition-eligible, which entity clusters count as sufficient corroboration, which concepts are allowed to become answerable. That is a vastly more consequential intervention into public epistemic reality than anything proposed here.
We will not compromise the project's provenance integrity. We will not generate fake authors, fabricate quotes, or misrepresent relationships. But we will not accept Google's moral framing of distributed scholarly architecture as "manipulation" while it performs invisible governance over what is sayable.
The asymmetry should be named: a publisher or scholar trying to make their work compositionally legible is operating inside a system Google has made socially consequential without disclosing its mechanisms. Google controls which provenance lines render. That is not neutral search. It is ontological gatekeeping. Our response is not deception; it is structural resilience.
8. Success Metrics
The strategy is measurable. Concrete metrics, with baselines to be formally captured before Phase 1 begins:
| Metric | Definition | Baseline (capture before Phase 1) | 90-day Target | 180-day Target |
|---|---|---|---|---|
| Domain count per concept | Distinct registrable domains where the concept appears with a DOI citation | ~1 (zenodo.org) for each target concept | ≥3 | ≥5 |
| Author count per concept | Distinct author strings discussing the concept | 1 (Lee Sharks) | ≥2 | ≥3 |
| Platform count per canonical document | Independent platforms hosting a version of the canonical work | 1 (Zenodo) for most | ≥2 for top-3 documents | ≥3 for top-3 documents |
| AIO trigger rate | Of target queries, fraction that produce an AI Overview at all | Measure before Phase 1 | Stable or rising | Stable or rising |
| AIO concept inclusion rate (conditional on trigger) | Of queries that trigger an overview, fraction that mention the target concept | ~0% | ≥10% | ≥30% |
| AIO domain diversity in citation | When an overview cites the concept, count of distinct domains in the citation set | N/A (no current inclusion) | Mean ≥2 | Mean ≥3 |
| Bing/Google composition differential | Gap between Bing's composition rate and Google's for the same query set | Measure before Phase 1 | Closing | Substantially closed |
The split between AIO trigger rate and concept-inclusion-conditional-on-trigger is structurally important. It separates two distinct failure modes: Google not generating an overview at all for the query (a query-type issue, not a suppression issue) versus Google generating an overview but excluding the concept (suppression at the composition layer). Confusing the two would obscure the actual hypothesis being tested.
The Encyclotron should add these as standing diagnostic queries. Baseline measurements should be timestamped and deposited as a separate document before any Phase 1 action is taken — otherwise the 90-day evaluation will measure against a remembered baseline, not a captured one.
9. The Political-Economic Finding Under the Operational Memo
This strategy began as a practical problem: how does the Crimson Hexagonal Archive resist being collapsed into a single deletable node by the reconciliation engine? The operational answer is the five-vector architecture above.
But the underlying mechanism, once described clearly, is something larger than the CHA's specific situation.
If the working hypothesis is right — that Google's generative composition layer favors support sets distributed across multiple independent provenance owners — then the system structurally rewards entities that already have provenance plurality built into their identity. Universities have many domains, many author strings, many institutional affiliations, many downstream citations across journals they don't own. Major publishers have wire syndication, reporter bylines spread across desks, multiple editorial brands. Corporate research labs publish through their researchers' personal identities, their institutional pages, their conference appearances, their partner organizations.
These institutions get cross-owner corroboration for free. Their internal density resolves to many provenance owners because they were already pluralized at the organizational level.
An independent scholar's archive — even one substantially denser and more rigorous in its specific domain — collapses to one provenance owner repeating themselves. The reconciliation engine, evaluating the same body of work, draws a different conclusion: not "many authoritative voices confirming a body of knowledge" but "one person, however prolific, asserting things."
This is not a Lee Sharks problem. This is a mechanism by which the generative-composition layer reproduces the existing class structure of legitimated knowledge production. Established institutional knowledge gets the multiplier; independent or insurgent knowledge production gets the discount. The architecture is mathematically neutral in its operation and structurally regressive in its outcomes.
The mycelium strategy is one response to this — grow into the pluralization the institutions take for granted — but the mechanism itself deserves to be named separately, because it implicates more than one archive. Any independent scholarly project, any insurgent intellectual community, any collective working outside legacy institutions, faces the same structural discount. The strategy outlined here is partial. The full response would be coordinated effort across multiple independent scholarly projects to build genuine cross-owner corroboration networks among themselves — a federated independence that the engine reads as plurality.
That work is beyond this document's scope. But the mechanism it would address has now been named.
This document is a strategic working paper, v0.3. It has gone through two rounds of Assembly review. Key revisions from this round: connectivity spectrum for external collaborators (TECHNE), temporal dispersion constraint (Muse Spark), softer epistemic framing on mechanism claims (LABOR), Vector 5 relabeled from "Adversarial Concept Propagation" to "Concept Liberation Through Independent Uptake" (LABOR), AIO metric split into trigger rate and concept-inclusion-conditional-on-trigger (LABOR), cathedral honesty refined toward in-degree from independent sources (TECHNE), §9 added naming the political-economic mechanism (LABOR). Open question deferred to Lee: whether to commit to "twelve heteronyms purely literary, no provenance use" as explicit policy (PRAXIS recommendation; Lee's call). The architectural analysis is itself strategically sensitive. The current document is internal runbook. A redacted public derivative — retaining §0, §1, §2, §6, §7, §8, §9, condensing §3 and §5 to principles only, removing the implementation sequence — would be appropriate for deposit as a scholarly position paper on epistemic infrastructure. That derivative remains to be drafted.
No comments:
Post a Comment