Tuesday, August 11, 2026

alexanarch datasets tab 8-11-2026

 Alexanarch

Datasets

Every machine-readable dataset in the Alexanarch corpus — sizes, record counts, what each carries, and which UI surfaces consume it. The canonical machine-readable catalog is /api/index.json; this page is its human-readable companion. Counts in teal are fetched live from the index on page load.

Primary registries

Canonical sources of truth. Editing these is the only way to change what the site shows.

data/blog-index.json10.3 MB · 2,809 posts
Complete index of the authorial publication surface, mindcontrolpoems.blogspot.com, built from the Blogger feed: every post with its title, URL, publication date, length and first 4,000 characters of text. Built 2026-08-08. It should have existed before the first restoration pass. Every pass before it derived a candidate URL from a deposit title and fetched that one URL; when the guess was wrong the work was recorded unrecoverable, and when the guess was right the gate could still reject it. Two measured cases: a record whose truth title carried a DOI its source predates by five weeks, scored 0.43 against a 0.75 gate; and a record whose correct URL sat unread in the queue's own field while the matcher scored six others. The index takes four minutes to build.
data/blog-deposit-map.json0.8 MB · 994 provisional pairings
Blog post ↔ deposit mapping for posts published 2026-01-01 or later, the Zenodo-deposit era. 1,075 posts in scope; 994 provisionally mapped; 81 with no deposit found. Every entry carries status: provisional and read: false, plus a field naming what would settle it — a shared DOI, document ID or hex coordinate in both bodies, because a title alone does not. The generator validates itself against eight hand-checked pairs and refuses to build below 8/8. Confidence bands: 971 strong, 12 good, 11 weak — read the weak ones first.
data/registry.json5.97 MB · 1,444 deposits
The canonical deposit registry. Each entry: bibliographic metadata, canonical v2 AXN, content hash, full-text path, entities[] (subject/predicate/object/type/evidence_status triples — the graph's source), wiki_article, Phase C references_concepts[] + references_concept_count, legacy AXN aliases, and glyphic_canary.
data/entity-index.json4.93 MB · 7,173 concepts
Curated concept layer. Each: term, definition, defined_in (founder deposit, on 7,097 of 7,173), entity_triples[], type taxonomy (specification/extracted/structural/empirical/theoretical/formal/genre/foundational/method), engagement type, and Phase C referenced_in[] + reference_count on 2,120.
Refined from:data/lexical-minting-registry.json
Bridged with:data/semantic-addresses.json (348 of 7,173 concepts also targeted by canonical queries)
Consumed by:/wiki//graph/
data/lexical-minting-registry.json3.52 MB · 12,032 raw terms
Broader pre-curation surface. Every term minted, coined, or formally extracted across the corpus — before noise-filtering and curation into the entity-index. 7,045 terms overlap with entity-index, 4,987 are LMR-only (raw), 128 are entity-index-only.
Consumed by:/lexical/
data/citation-graph.json1.59 MB · 4,866 edges
Inter-deposit citation edges. After Phase B extraction (commit ee1a1db) edges include: doi_resolution (4,311 — legacy), deposit_number_reference (346 — #N), ea_id_reference (158 — EA-* sovereign IDs), axn_hex_reference (21), axn_reference (12 — full canonical), plus 11 hand-curated types. 696 citing deposits, 445 cited.
Generated by:scripts/citation_extractor.py
Consumed by:/citations/
data/semantic-addresses.json1.28 MB · 2,061 addresses
Query addresses — canonical queries posed to the composition layer (AI Overview, Google Search) with their observation status. Conceptually a field-set on lexical entities: 100% of the 348 unique refers_to targets exact-match into entity-index. Each address: canonical_query, is_quoted, refers_to[], type, battery_membership[], sources[], observations[], observation_class.
Class counts:subjunctive1,750· observed291· verified-non20· unrated83
Reconciled with:data/EA-WG-CAPTURES-01-v8.11.json (observations[] entries are capture references)
Consumed by:no dedicated UI surface yet — planned at /addresses/
data/EA-WG-CAPTURES-01.json282 KB · 204 captures · v9.6
AI Overview Capture Registry — captured queries with match status (mt): EXACT MATCH / BROAD MATCH / ADOPTION / ZERO RESULT / null. Each entry: section, slug, query, date, source-format, status, description, image refs. Reconciled into semantic-addresses via slug. Versionless alias tracks the current mirror; individual versioned artifacts (e.g. v8.11, v9.6) are preserved as immutable deposit snapshots.
Consumed by:/captures/· deposit #3's wiki entry
data/doi-resolution-index.json1,937 mappings · 3.5 · 1,360 live redirects / 577 quarantined
Legacy Zenodo DOI → Alexanarch AXN resolution table with per-entry evidence envelopes (identifier validity × archive membership × relationship × evidence × quarantine reason). A redirect is emitted only for verified identifiers with confirmed membership and a same-work relationship; everything else resolves to inspection. Counts on this card are hydrated live from /data/resolver-status.json — the computed status artifact — and are never hand-written.
Consumed by:/resolve/(inspection) and/go/(redirect); immutable snapshots atzenodo-datacite-batch
data/datacite-full-backup.json9.06 MB
Full DataCite metadata snapshot of all 1,817 CHA-minted DOIs. The empirical foundation for the audit in EA-MPAI-DOI-IMPERMANENCE-01 v2.0 (#868). Methodology replicable via DataCite API at https://api.datacite.org/dois/{doi}.
Paginated copies:page 1·page 2·page 3·page 4
Consumed by:Reference dataset; no live UI projection

Derived surfaces

Regenerated from primary registries. Edits to these are overwritten on next regeneration. Generated by scripts/regenerate_surfaces.py.

data/browse-index.json443 KB · 1,448 entries
Compact deposit list (no full-text bodies). Used by the static no-JS browse page.
data/chunks/registry/9 chunks · ~1 MB each
Registry split into ~1 MB chunks for human-loadable browsing. Each chunk covers a contiguous deposit-number range; _index.json catalogs the chunks.
/sitemap.xml103 KB
XML sitemap. Every static record page enumerated for crawler indexing.
/SHA256SUMS.txt160 KB · 1,448 lines
Content-addressable checksums for every deposit. Lets any mirror verify byte-exactness of the corpus.

Protocols & schemas

Machine-enforced definitions. Hand-editing these produces drift detected by bootstrap_familiarization.py. Use protocol_update.py to amend protocols atomically.

api/index.jsoncentral catalog
Single source of truth. Lists all protocols, schemas, registries, derived surfaces, scripts — each with content_sha256, canonical_path, and referenced_by. New instances run bootstrap_familiarization.py to verify nothing has drifted before any work.
api/deposit-protocol.jsonv1 (alexanarch-deposit-protocol/v1)
Deposit validation rules. Rule families: PV (protocol version) / REQ (required fields) / AXN (identifier) / CONS (consistency) / SUR (surface) / IDX (index integrity). Enforced by scripts/validate_deposit.py and CI.
api/axn-protocol.jsonv2 (axn/v2)
Alexanarch Identifier protocol. Format: AXN:<HEX>.<FAMILY>.<6 EMOJI> where the six-emoji suffix is derived from the first 6 bytes of SHA-256 of canonical bytes, mapped through 256 curated emoji. v1 (4-emoji) aliases preserved in deposit legacy_axn / axn_history.
api/enrichment-protocol.jsonv1 (enrichment/v1)
Citation extraction + concept backlink protocol. Defines via-types for citation edges (axn_reference, axn_hex_reference, deposit_number_reference, ea_id_reference, doi_resolution, etc.) and the Phase C bidirectional concept↔deposit indexing.
api/deposit-schema.json + api/schemas/deposit-entry.schema.jsonJSON Schema
Submission schema (form-facing) and registry entry schema (storage-facing).

Supporting datasets

Reference data not (yet) consumed by any UI surface. Available for inspection, citation, and downstream tooling.

data/external-source-registry.json84 KB
Catalog of external reception sources (sites, archives, mirrors) referenced across the corpus.
data/heteronym-doi-sift.json44 KB
Cross-reference of Dodecad heteronym attribution across the legacy Zenodo DOIs.
data/JOURNAL-MAPPING-PRELIMINARY.json227 KB
Preliminary mapping of deposits to journal-style groupings.
data/batch-axn-assignment.json974 KB
Historical record of batch AXN assignments during the CHA → Alexanarch migration.
data/restoration-batch-plan.json20 KB
Restoration plan ledger.
data/zenodo-link-scan.json217 KB
Survey of inbound and outbound Zenodo links across the corpus.
/datasets/new-human-primary/343 curated · 218 canon-declared · v0.1
Curated inventory of New Human Project primary texts within the archive: poetry, creative works, and short works by Lee Sharks and the Dodecad heteronyms. Packaged alongside the classical-translation corpora (Perseus, Gutenberg) as the archive's own primary-text substrate. v0.1 placeholder — full curation follows.
Consumed by:landing UI
/datasets/negshape-deletion-bibliography/13.3 MB · 20 files · v1.0
The Negative Shape of the Work — the 2026-06-19 bulk deletion constituted as a formal bibliographic container (sovereign identifier CHA-DELETION-CORPUS-20260619, the publisher having issued none): 1,834 identifier entries, 1,621 membership-confirmed, 6,484 formal citations (MLA 9 / Chicago 17 / APA 7 / BibTeX) with rendering withheld for non-confirmed entries; bibliographic matrix at 99.98% core-field coverage over the confirmed corpus, per-cell metadata provenance plus per-row membership provenance; rejected-candidate ledger preserving 68 name-string collisions including the Jack E. Feist heteronym collision. Generators included; reproducible from the DataCite backup and DOI Resolution Index.
/datasets/tombstone-mirror/9.48 MB · 6 files · v1.0
The Tombstone Mirror — complete machine-readable mirror of what Zenodo serves for the 2026-06-19 deletion batch: full API-tombstone census of all 2,027 deleted objects (872-second sweep chronology, removed_by total enumeration), 57 rich pre-shrink metadata captures with file checksums and terminal statistics, the kill-ledger cohort extraction, and the stripped-page corpus documenting the surface’s live reduction on 2026-07-12. SHA-256 provenance per file.
/datasets/zenodo-datacite-batch/22.3 MB · 9 files · v0.1
The DOI Resolution Index v3.7.2 packaged alongside the DataCite metadata backup captured before the 2026-06-19 Zenodo severance. Together these instruments constitute the empirical foundation for EA-EROSION-01 (population census) and EA-REMEDIATION-01 (outward-facing whitepaper). SHA-256 provenance per file.

Archive substrates

Named-identity registries and mirrored blog materials. Data-rhizome is the source of truth; alexanarch and downstream surfaces (leesharks.com, Mandala Oracle) all consume the same JSONL files.

datasets/heteronyms/84 MB · 12 Dodecad positions · 13 total · 3 frames · crosswalk + 4 derived fields
Twelve Dodecad heteronym positions + Jack Feist / LOGOS* standing outside the twelve + Lee Sharks (MANUS) outside as human bearer + Viola Arquette Adjacent. Twelve portraits from the September 21, 2025 Logotic Science voicecast series on mindcontrolpoems.blogspot.com, reconciled against the canonical v1.1 Constitutional Registry (AXN:03EE, 2026-07-03). Rotating-outside architecture: different frames put different figures outside the twelve — Nobel Glas in Sep 2025 Logotic Science, Lee Sharks in v1.1 Constitutional, Feist as *LOGOS in the archive-default frame. frames.jsonl documents each framing with the outside figure + twelve inside positions. Avatar assets (240px + 120px WebP) for the Mandala Oracle message-thread avatar feature.

Added 2026-08-11: crosswalk.json joins this dataset to the lexical minting registry, the concept map, the wiki and the citation graph on the creator field — 9,453 deposit-level citation edges collapsed to 1,910 position-level edges. With four derived figures, generated from the corpus rather than composed from a reading of it: the Dodecad graph, the citation field (9,136 chords), the lexical field (12,073 terms), and the capture field (261 reception events). Deposited under tether #1448.
Substrate:data-rhizome/datasets/heteronyms/ (heteronyms.jsonl · frames.jsonl · schema.json · MANIFEST.json · avatars/{240,120,src}/ · mindcontrolpoems-2025-09/ raw preservation)
Purpose:canonical heteronym reference for cross-substrate rendering; Mandala Oracle avatar lookup; leesharks.com heteronym-tab source

External corpora

Public-domain literary corpora mirrored for the L1/L2 transform pipeline. Source-of-truth bytes preserved in data-rhizome; alexanarch surfaces the indexes and browseable landing pages.

datasets/perseus-classical/804 KB · 1,161 works · 764 aligned
Mirror of PerseusDL canonical TEI-XML: canonical-greekLit (100 authors, 826 works, 648 with English translation), canonical-latinLit (54 authors, 334 works, 115 English), canonical-farsiLit (Hafez Divan with Persian/English/German). Every work addressable by canonical URN — e.g. Homer's Iliad is urn:cts:greekLit:tlg0012.tlg001.perseus-grc2 (Greek) with Murray 1924 (perseus-eng3) and Butler 1898 (perseus-eng4) in the same directory.
Substrate:data-rhizome/datasets/corpora/perseus/ (474 MB, snapshot 2026-07-06, tarball SHAs in MANIFEST.json)
Purpose:aligned substrate for L1 (source→clarity translation) and L2 (voice restoration) transform pipeline
datasets/gutenberg-classical/~800 KB · 972 identified · 206 mirrored
Project Gutenberg catalog snapshot (90,252 rows) plus a pilot of 206 canonical public-domain classical translations (Chapman & Butler Homer, Dryden Vergil, Longfellow Dante, Jowett Plato/Thucydides, Rawlinson Herodotus, Ormsby & Motteux Cervantes, Gummere & Hall Beowulf, Marcus Aurelius, and others. 972 additional target texts identified in the catalog but not yet fetched.
Substrate:data-rhizome/datasets/corpora/gutenberg/ (~97 MB, snapshot 2026-07-06, per-file SHAs in classical/MANIFEST.json)
Purpose:coverage beyond Perseus — alternative translations and canonical works Perseus doesn't index (Beowulf, Cervantes, Marcus Aurelius, etc.)

Per-deposit text bodies

Every deposit's canonical text is stored at data/texts/AXN-<HEX>-text.md. The hex maps via the registry's hex field to a canonical deposit_number. Body SHA-256 anchors each text into the AXN.

data/texts/AXN-*-text.md~25.5 MB total · 1,448 files
One Markdown file per deposit. The full text body. Source for citation extraction, concept backlinks, and wiki article generation.

Scripts

Canonical operational scripts. The full catalog with descriptions lives in api/index.json → scripts.

scripts/bootstrap_familiarization.pyrequired-first-read
Verifies every protocol/schema content hash matches what api/index.json claims. New instances run this with --strict at session start. Receipt appended to data/instance-familiarization.log.
scripts/protocol_update.pyatomic protocol amendment
The only supported path for modifying a protocol JSON. Recomputes hash, updates index, appends to change_log atomically. Direct hand-editing produces drift.
scripts/axn_lib.pycanonical AXN derivation
256-entry AXN_GLYPHS table + cluster catalog. Derives v2 6-emoji suffix from first 6 bytes of SHA-256.
scripts/regenerate_surfaces.pyidempotent
Brings every derived surface (browse, browse-index, chunks, sitemap, SHA256SUMS) into agreement with data/registry.json. Run after every registry change.
scripts/validate_deposit.pyCI-enforced
Validates the registry against the deposit protocol. Rule families: PV/REQ/AXN/CONS/SUR/IDX. Runs on every commit via .github/workflows/validate-registry.yml.
scripts/citation_extractor.pyenrichment
Scans deposit texts for AXN refs, EA-* IDs, #N references, and DOIs. Writes new edges to data/citation-graph.json.
scripts/concept_backlink.pyenrichment
Scans every deposit text for every entity-index concept. Writes referenced_in[] and reference_count onto entity-index concepts and references_concepts[] onto registry deposits.
scripts/backfill_axn_compliance.pyhistorical
One-time migration. Backfilled 13 pre-v2 AXNs from 4-emoji to 6-emoji canonical, preserving v1 forms in legacy_axn and axn_history.
datasets/axnidentifiers/0.0 MB · 6 files
The AXN identifier system as data.
The capture registry as data: dated reception events with query, section, attribution state and finding.
datasets/dataflow-atlas/2.9 MB · 18 files
Eleven versions of the archive's own map — where data enters, what transforms it, where it is published, and which instruments watch the map. v1.1 is the head; superseded versions are kept because the map's history is evidence.
A fixture for testing whether a system's deletion behaviour conforms to what it claims.
Which DOI resolves to which work, graded by truth source. Built after 1,817 Zenodo DOIs were tombstoned.
Programmed bibliographic suppression, documented at commit level.
A provenance-erasure case file.
datasets/registry-audit/18.4 MB · 48 files
The archive auditing itself. Batch findings across the full deposit range, the content-fingerprint scan, the DW-verified register, and RECORD-SHAPE-AND-PROPAGATION — the specification of the nine declaration sites and the seven-step propagator pipeline. This is where the analysis lives.
datasets/venues/0.0 MB · 2 files
Publication venues.

No comments:

Post a Comment