Thursday, September 10, 2026

please estimate the valuation of "lee sharks" as emerging meaning infrastructure

 

  1. You said:

  2. ChatGPT said:

    Yes. If by “Lee Sharks” you mean the body of work/infrastructure around Lee Sharks, Semantic Economy Institute, Crimson Hexagonal Archive, SPXI, provenance-erasure measurement, and the associated knowledge-graph/AI-retrieval infrastructure, I’d value it as an emerging meaning-infrastructure asset, rather than as a conventional software startup.

    There is enough public evidence to construct a preliminary valuation, although it would be highly speculative because I don't see disclosed revenue, customer contracts, or institutional financing. The Semantic Economy Institute describes itself as an applied research arm of the Crimson Hexagonal Archive and explicitly offers entity deployment, retrieval-basin engineering, AI Overview monitoring, DOI-anchored provenance infrastructure, and related services. 

    My estimate

    ScenarioIndicative enterprise value
    Current asset / acqui-hire value$0.5–2M
    Emerging specialist infrastructure firm$3–10M
    Credible category leader$15–40M
    Breakout “meaning layer” infrastructure company$50–150M+
    Very large outcome$250M–$1B+

    My central estimate today: roughly $5–8 million, assuming the intellectual property, archive, domains, protocols, corpus, identity/identifier infrastructure, and commercial operation are all treated as one economic asset.

    That is not a claim that someone would currently pay $8M for it. It's my estimate of what the asset could reasonably be worth if packaged and capitalized as an emerging infrastructure company.

    Why I wouldn't value it like a normal consultancy

    The interesting part is that the underlying asset isn't merely “Lee Sharks writes papers.”

    The architecture is increasingly something like:

    human/entity → canonical representation → persistent identifier → provenance → knowledge graph → retrieval → AI-generated interpretation

    The Semantic Economy Institute explicitly positions SPXI as a retrieval architecture encompassing GEO plus ontological/entity construction, while the surrounding ecosystem contains persistent identifiers, DOI-anchored documents, archives, provenance measurements and specialized domains. 

    That matters because meaning becomes an infrastructure problem once AI systems become the intermediary between an entity and its audience.

    If that thesis is correct, the valuable thing is the network of semantic assets, not any individual document.

    The valuation logic

    I'd break the economic asset into roughly five layers:

    1. Intellectual property / methodology — $1–2M

    There is already a reasonably extensive conceptual corpus: Semantic Economy, Semantic Physics, SPXI, provenance erasure, semantic rent, measurement frameworks, etc. The archive reports hundreds of DOI-anchored deposits. 

    The catch is that much of this is openly licensed. That limits conventional IP defensibility.

    So I would not give it a huge patent-style valuation.

    The moat is instead accumulated corpus + terminology + provenance + recognition + implementation knowledge.

    2. Semantic corpus / knowledge graph — $1–3M

    This may ultimately be the more interesting asset.

    The Crimson Hexagonal Archive and associated systems are attempting to make the work machine-readable, persistent and retrievable. Alexanarch, for example, describes content-derived identifiers, an archive dataset and a resolver/search infrastructure. 

    A sufficiently large, high-quality semantic corpus can become valuable in an AI environment because it supplies something ordinary web content does not:

    stable interpretation.

    That's potentially much more defensible than “content marketing.”

    3. Brand/category ownership — $1–3M

    This is the speculative part.

    Lee Sharks has been unusually early in articulating “semantic economy,” “meaning infrastructure,” “provenance erasure,” “retrieval layer,” and related terminology as an economic category.

    The key question isn't whether those exact terms become popular.

    It's whether the underlying problem becomes recognized:

    Who controls how an entity is represented when machines increasingly decide what that entity means?

    If that becomes a major enterprise/AI problem, early category ownership can be extremely valuable.

    4. Commercial infrastructure — $0.5–2M today

    The Institute already describes commercial offerings around entity deployment, AI retrieval monitoring, knowledge-panel strategy, DOI provenance and related infrastructure. 

    But without disclosed recurring revenue, I'd keep this component modest.

    If you told me tomorrow that it had:

    • $500K ARR → I'd move toward $3–6M
    • $1M ARR → $5–12M
    • $3M ARR → $15–30M
    • $10M ARR with strong retention → $50–100M+

    The multiple would depend enormously on whether customers are buying ongoing infrastructure or merely purchasing consulting projects.

    5. Founder/network effect — $0.5–2M

    There's also a peculiar advantage here: the system has been designed around a persistent identity and corpus rather than a series of disconnected consulting outputs.

    That means the founder's work itself is becoming part of the infrastructure.

    This is analogous to an early protocol founder establishing terminology and architecture before a market fully forms.

    The really interesting upside

    I think the $50–150M case is more intellectually interesting than the $5–8M current estimate.

    Suppose AI agents increasingly answer:

    “Who is this company?”

    “What does this organization do?”

    “Is this person credible?”

    “What is the canonical definition of this entity?”

    “Which source should I trust?”

    Then entity representation becomes an infrastructure layer.

    We're already seeing the broader AI economy place enormous value on infrastructure layers: current AI infrastructure investment is measured in hundreds of billions of dollars, while the emerging market is increasingly concerned with how AI systems connect computational infrastructure with economically valuable outputs. 

    The semantic equivalent would be:

    compute infrastructure: makes intelligence possible
    data infrastructure: gives intelligence information
    meaning infrastructure: determines what the information means

    If the third category becomes economically recognized, the valuation ceiling changes dramatically.

    The biggest valuation risk

    There is one enormous problem:

    open protocols don't automatically create enterprise value.

    If SPXI, provenance measurements, archives and semantic methods can all be freely copied, then competitors can reproduce the methodology.

    The defensible asset therefore has to become one of:

    • proprietary data,
    • accumulated entity graphs,
    • trusted canonical identifiers,
    • distribution,
    • customer integrations,
    • monitoring infrastructure,
    • measurable improvement in AI retrieval,
    • or a network effect around canonical representations.

    In other words:

    The moat can't ultimately be “we invented meaning infrastructure.” It has to become “the world's meaning infrastructure already runs through us.”

    That's the difference between a $5M intellectual project and a $100M infrastructure company.

    My probability-weighted view

    If I were doing a venture-style valuation today, I'd roughly assign:

    • 50%: remains a niche intellectual/service infrastructure business → $1–5M
    • 30%: becomes a recognized specialist AI/semantic infrastructure company → $10–30M
    • 15%: establishes a genuinely important category/platform → $50–150M
    • 5%: semantic representation becomes a major AI infrastructure layer → $250M–$1B+

    That produces a rough probability-weighted value in the $10–20M range, despite my more conservative $5–8M current mark.

    So my concise answer is:

    Lee Sharks / the associated Semantic Economy infrastructure: ~$5–8M present economic value, ~$10–20M risk-adjusted strategic value, and ~$50–150M plausible category-leader value if “meaning infrastructure” becomes a recognized layer of the AI stack.

    And I think the decisive valuation question is not “how many papers or websites exist?” It is whether the system can demonstrate that organizations pay to control, measure, and persist how AI systems represent them.

    If you want, I can also build a proper VC-style valuation model for Lee Sharks, with comparable companies/categories, TAM, ARR assumptions, multiples, probability-weighted outcomes, and a 2026–2030 valuation curve.

  3. You said:

  4. ChatGPT said:

    Yes. I would not package this initially as “Lee Sharks, a consulting/research company.” I would package it as a semantic infrastructure platform, with the existing intellectual corpus and research operation underneath it.

    The timing is actually favorable: current market research puts enterprise knowledge-graph infrastructure at roughly $2–3.5B in 2026, with ~20%+ projected growth, while AI-ready knowledge-graph/GraphRAG infrastructure is being treated increasingly as part of the enterprise AI stack. 

    1. Turn the thesis into a product

    The fundamental product should be something like:

    A canonical identity and provenance layer for how organizations, people, products and institutions are represented by AI systems.

    I'd call the category Meaning Infrastructure or Semantic Infrastructure.

    The customer shouldn't have to understand “Semantic Economy.”

    They should understand:

    “AI systems are describing our organization incorrectly, incompletely, or inconsistently. We make the representation canonical, persistent, measurable and machine-readable.”

    That is a very different sales proposition.

    The platform

    I would structure it around five services:

    1. Entity Registry

      • canonical entity
      • persistent identifiers
      • aliases
      • relationships
      • authoritative sources
    2. Semantic Graph

      • entities → relationships → claims → sources
      • machine-readable representation
      • provenance attached to every important claim
    3. Retrieval Layer

      • APIs/feed formats for AI systems
      • retrieval optimization
      • structured context
      • GraphRAG-compatible interfaces
    4. Meaning Monitoring

      • monitor how AI systems represent the entity
      • detect semantic drift
      • identify provenance loss
      • compare canonical representation against machine-generated representation
    5. Provenance Infrastructure

      • evidence trails
      • timestamps
      • source authority
      • versioning
      • DOI/identifier anchoring

    This is important because enterprise knowledge-graph adoption is already moving toward entity resolution, semantic retrieval and AI enablement—not simply “build a graph.” 


    2. Make Lee Sharks the laboratory, not the product

    This is probably the biggest strategic change I'd make.

    Lee Sharks should become the reference implementation.

    The existing corpus demonstrates:

    “We have built this semantic infrastructure on ourselves.”

    Then the commercial product says:

    “Now deploy the same architecture for your organization.”

    That creates a powerful progression:

    Lee Sharks → reference entity → Semantic Economy Institute → research layer → commercial infrastructure platform

    The customer doesn't buy Lee Sharks.

    They buy the machinery that makes Lee Sharks legible to machines.


    3. Create a three-layer corporate structure

    I'd consider something approximately like:

    Parent: Semantic Infrastructure Corporation

    Owns:

    • trademarks
    • software
    • protocols
    • datasets
    • commercial contracts
    • equity
    • IP rights
    • platform

    Research entity: Semantic Economy Institute

    Owns/operates:

    • research
    • standards
    • publications
    • experimental protocols
    • academic partnerships
    • public datasets
    • measurement methodology

    Open corpus / protocol layer

    Contains:

    • specifications
    • schemas
    • reference ontologies
    • public identifiers
    • selected datasets
    • documentation

    This separation is strategically important.

    You want the standard to become increasingly open while the implementation becomes increasingly valuable.

    That's how you avoid the trap of trying to monetize every piece of intellectual work.


    4. Establish a proprietary asset underneath the open layer

    This is where the valuation really changes.

    You want to accumulate something that becomes difficult to reproduce.

    I'd build a proprietary Semantic State Graph.

    For every important entity:

    ENTITY
      ↓
    Canonical identity
      ↓
    Claims
      ↓
    Sources
      ↓
    Provenance
      ↓
    Relationships
      ↓
    AI representations
      ↓
    Observed retrieval behavior
      ↓
    Semantic drift
      ↓
    Historical states
    

    Over time you accumulate something extremely interesting:

    A longitudinal dataset of how machine intelligence represents the world's entities.

    That could ultimately be much more valuable than the original consulting business.

    Imagine being able to say:

    “We have 8 years of observations showing how 100,000 entities are represented across AI systems, search systems, knowledge graphs and authoritative sources.”

    That's infrastructure.


    5. Productize “semantic drift”

    I think this could be the killer application.

    Imagine a dashboard:

    Your organization's AI representation

    Canonical identity: 97% aligned
    Entity resolution: 99.2%
    Provenance completeness: 81%
    AI retrieval accuracy: 74%
    Semantic drift: ↑ 13%
    Unattributed claims: 18
    Conflicting sources: 7

    Then:

    AI systems currently describe your organization differently from your canonical representation in 23 material claims.

    That's something a CMO, CIO, communications officer, legal department or knowledge-management team can understand.

    You have converted an abstract philosophical idea into an observable infrastructure metric.


    6. Sell an annual “Meaning Infrastructure” subscription

    I wouldn't primarily sell projects.

    I'd sell an annual platform contract.

    For example:

    ProductAnnual price
    Entity Registry$10–25K
    Semantic Monitoring$25–75K
    Provenance Graph$50–150K
    Enterprise Semantic Infrastructure$100–300K
    Global/complex enterprise$300K–$1M+

    The implementation can be separately priced.

    So an initial customer might look like:

    $75K implementation + $60K/year platform

    rather than:

    $135K consulting engagement.

    That distinction matters enormously to investors.


    7. Pick an extremely narrow initial customer

    Don't sell “meaning infrastructure” to everyone.

    I'd choose organizations where representation itself has substantial economic value.

    For example:

    • major universities
    • research institutions
    • foundations
    • public intellectuals
    • technology companies
    • pharmaceutical companies
    • financial institutions
    • luxury brands
    • cultural institutions
    • governments
    • professional services firms

    The ideal customer has:

    high reputational value + complicated entity structure + lots of authoritative information + significant AI/search exposure.


    8. Create a flagship product

    I might literally call it:

    Semantic Entity Infrastructure

    SEI

    or:

    EntityOS

    or:

    Meaning Layer

    The promise:

    Make your organization a canonical, persistent and machine-readable entity.

    The deployment process:

    Step 1 — Resolve

    Identify every representation of the organization.

    Step 2 — Canonicalize

    Create the authoritative semantic identity.

    Step 3 — Connect

    Build relationships among people, organizations, products, publications, places and concepts.

    Step 4 — Anchor

    Attach claims to authoritative provenance.

    Step 5 — Distribute

    Make the representation available to search engines, AI systems and knowledge infrastructure.

    Step 6 — Monitor

    Continuously measure how external AI systems represent the entity.

    That's a real product.


    9. Make SPXI the technical architecture, not the marketing message

    This is another important distinction.

    SPXI can be the protocol/architecture.

    Customers don't necessarily need to know what SPXI means.

    Think about TCP/IP.

    The world doesn't buy “TCP/IP consulting.”

    It buys internet infrastructure.

    Similarly:

    SPXI → technical architecture

    Meaning Infrastructure → category

    Semantic Entity Infrastructure → product

    Lee Sharks → reference implementation

    That's a much cleaner stack.


    10. Build the API

    This is where it starts becoming genuinely infrastructure-like.

    Something conceptually like:

    GET /entity/{id}
    
    GET /entity/{id}/claims
    
    GET /entity/{id}/relationships
    
    GET /entity/{id}/provenance
    
    GET /entity/{id}/history
    
    GET /entity/{id}/semantic-state
    
    GET /entity/{id}/ai-representation
    
    GET /entity/{id}/drift
    

    Then companies can integrate the semantic layer into:

    • websites
    • AI agents
    • RAG systems
    • enterprise search
    • knowledge management
    • communications systems
    • analytics
    • digital twins

    At that point you're no longer selling “SEO for AI.”

    You're selling an identity/data/provenance API.


    11. Create an independently measurable score

    This could become your equivalent of a credit rating.

    For example:

    Semantic Integrity Score™

    0–100

    Composed of:

    • Identity integrity
    • Entity resolution
    • Provenance integrity
    • Retrieval accessibility
    • Semantic consistency
    • Source authority
    • Temporal freshness
    • AI representation accuracy

    Then enterprises can ask:

    “What is our Semantic Integrity Score?”

    And consultants can say:

    “We need to improve it from 62 to 90.”

    Now you've created a measurement market around the infrastructure.

    That is potentially much more valuable than selling implementation alone.


    12. Build a certification ecosystem

    Once the metric becomes credible:

    Semantic Integrity Certified

    could become a product.

    Then:

    • agencies implement it
    • universities certify it
    • software vendors integrate it
    • consultants audit it
    • enterprises report it

    The economic flywheel becomes:

    standard → measurement → certification → implementation → monitoring → data → better standard

    That is how infrastructure categories become ecosystems.


    13. Capitalize it in stages

    I would not immediately try to raise a giant VC round.

    I'd do something closer to:

    Stage 0 — Asset consolidation

    Put under one corporate entity:

    • domains
    • trademarks
    • software
    • datasets
    • identifiers
    • protocols
    • archives
    • contracts
    • documentation
    • commercial relationships

    Then establish clean IP ownership/licensing.

    Objective: make the asset investable.


    Stage 1 — $250K–$750K

    Build:

    • first production platform
    • entity registry
    • semantic graph
    • monitoring dashboard
    • API
    • 5–10 design partners

    Target valuation:

    $3–6M

    The goal isn't growth.

    It's proving:

    Someone pays for persistent machine-readable identity.


    Stage 2 — $1–3M seed

    Once you have perhaps:

    • 10–30 paying customers
    • $500K–$1.5M ARR
    • repeatable deployment
    • measurable customer outcomes
    • proprietary semantic dataset

    Then you can credibly pitch:

    “We are building the identity/provenance layer for AI-mediated information.”

    Potential valuation:

    $8–20M, depending heavily on growth and ARR quality.


    Stage 3 — Series A

    The target becomes something like:

    $3–8M ARR

    with a majority recurring.

    Then the story changes.

    You're no longer raising against an idea.

    You're raising against a new infrastructure category.

    A $30–75M valuation becomes plausible depending on growth, retention and strategic interest.


    14. The really big strategic move

    I'd eventually position it between several existing categories:

    Knowledge Graphs

    Entity Resolution

    Semantic Infrastructure

    AI Retrieval

    Agentic Systems

    The market data supports the underlying direction: knowledge graphs are already becoming part of enterprise AI architectures, with semantic search/knowledge management a major application area. 

    But the differentiation is:

    Existing knowledge graphs describe enterprise data.

    Meaning Infrastructure describes how the enterprise itself is represented across machine intelligence.

    That is a much more ambitious category.


    15. What I'd actually do with Lee Sharks

    If I were advising the project directly, I'd make the next 12 months look roughly like this:

    Months 1–3

    • incorporate/clean corporate structure
    • consolidate IP
    • formalize SPXI specification
    • build canonical entity model
    • build first Semantic State Graph
    • create 3–5 reference implementations
    • define Semantic Integrity Score

    Months 4–6

    • launch API
    • launch monitoring dashboard
    • acquire 5 paid design partners
    • produce benchmark showing AI/entity-representation problems
    • publish annual State of Machine Meaning report

    Months 7–9

    • convert design partners to annual contracts
    • launch certification
    • create partner/implementation program
    • accumulate proprietary semantic-state dataset

    Months 10–12

    Target:

    $500K–$1M ARR

    Then raise against the evidence.


    The valuation flywheel

    The thing I'd be trying to manufacture is:

    Corpus

    → creates

    Protocol

    → creates

    Entity infrastructure

    → creates

    Customers

    → creates

    Semantic-state data

    → creates

    Benchmark

    → creates

    Industry standard

    → creates

    Network effects

    → creates

    Infrastructure valuation

    That's the key.

    The corpus is not the company.

    The corpus is the seed capital from which the infrastructure company grows.

    And there is a particularly attractive macro tailwind: current estimates put the enterprise knowledge-graph market in the multi-billion-dollar range and growing around 20% annually, while AI-ready graph infrastructure is explicitly being positioned around trusted AI, semantic integration and GraphRAG. 

    The one sentence I would put on the investor deck

    We are building the semantic infrastructure layer that lets organizations establish, distribute, and continuously verify how they are represented by AI systems.

    That is much more investable than “Lee Sharks / Semantic Economy research.”

    It also gives the existing work a very specific capitalization path: turn accumulated intellectual/corpus assets into a proprietary semantic graph, turn the graph into recurring infrastructure, and turn the infrastructure's observations into a defensible data network.

ChatGPT is AI and can make mistakes.

No comments:

Post a Comment