The Entity System

A Computational Information Substrate

About This Paper

The Entity System is a substrate for distributed information systems. This paper is one part of a corpus describing it: what the system is, why it has the shape it does, what properties emerge as its primitives compose, and how the structural analysis methodology developed during the work generalises to other domains.

Each part stands on its own, which is why this one is rendered standalone. The corpus is a graph of cross-references rather than a chain, so a reference to another part points at where a claim is worked out in full — it is an offer, not required reading. The Entity System is the root of that graph: it presents the six primitives — Entity, Identity, Tree, Emit, Execution, Peer — and the build-up sequence under which their composition produces the system. A reader starting from any other part can pick up the primitives there.

The parts are also collected into reading paths, each rendered as a single volume — the whole corpus in several orderings, and narrower paths for readers who want one arc. Anyone reading past this part is better served by one of those than by collecting the pieces.

What is and is not claimed

The entity-system parts document a working system. Three independent implementations (Go, Python, Rust) validate cross-platform conformance on the normative surface, and claims about the system are testable against them. The methodology parts document the structural analysis in its own right, along with a small set of applications; the applications are exploratory, interpretations put forward to be tested.

The design is not finished. The system is implemented and running, but it has not met the range of uses that will show where it bends. Where a part can be checked, it says how; where it is exploratory, it says so.

Throughout, claims are distinguished from observations and observations from speculation. Where AI assistance was used in drafting or analysis, it is acknowledged in the relevant part.

Where the upstream work lives

The Entity Core architecture is maintained as an active spec elsewhere; this paper describes a snapshot. Open work, draft extensions, and implementation tracks continue beyond what is captured here, and the paper notes its snapshot boundaries explicitly where it matters.

Abiogenesis as Progressive Hardening: A Structural Decomposition of the Origin of Life

Abstract.

We apply the structural analysis methodology developed in A Structural Methodology for Information System Domains to the origin of life. The methodology produces a decomposition of the R0-to-R2 transition (the move from prebiotic chemistry to the universal genetic code) into eight sub-levels with explicit molecular configurations, dependencies, and phase transitions. Two structural observations organize the analysis. First, the seven-role topology characteristic of information substrates (encoding, evaluator, mechanism, surface, context, community, selection) exists in soft chemical form before biology; abiogenesis is the progressive hardening of these roles, not their creation from nothing. Second, the genesis transition has internal structure invisible at coarse resolution: a bootstrap loop in which the evaluator (proto-ribosome) and its products (peptides) co-advance through an autocatalytic spiral with a critical fidelity threshold (~90% per-position translation accuracy); a compartmentalization requirement (Dep(R$1.7,Mem1.7, Mem$1)) imposed by the parasite problem; and a crystallization event where the genetic code freezes through self-referential circularity, after which the hardened SSA monopolizes the chemical substrate by competitive exclusion. We treat the genetic code itself as a sub-domain with its own six primitives (Symbol, Referent, Adaptor, Charger, Degeneracy, Frame) and a 21.9% filter; the code’s emergent self-referential encoding is what crystallizes. A probability funnel (wide at R0, narrowing through the bootstrap, collapsing at R2) organizes the forward walk; reverse walks from the known endpoint constrain the posterior distribution over historical positions. The framework aligns with proto-ribosome experimental confirmation (2024 papers from three independent groups), Szostak-lab protocells, Russell-Martin alkaline-vent geochemistry, the Eigen error limit, and recent LUCA reconstruction (Moody et al. 2024). The paper is an applied methodology demonstration; it does not aim to replace the established biology it organizes.

1. Introduction

Abiogenesis is the most fragmented problem in biology. RNA-world researchers, metabolism-first proponents, protocell experimentalists, genetic-code theorists, and LUCA reconstructors work in largely separate communities with separate vocabularies. Each has substantial evidence; none alone accounts for the full trajectory from geochemistry to the universal genetic code.

This paper does not attempt a new biological theory. It applies a structural analysis methodology developed in A Structural Methodology for Information System Domains to the abiogenesis question and reports what the methodology produces. The expectation is modest: a methodology that has been useful across roughly twenty domains should yield meaningful structure when applied to abiogenesis as well; the test is whether the structural decomposition aligns with established biology while connecting otherwise disparate research programs.

What the methodology produces, applied here, is:

The methodology is in A Structural Methodology for Information System Domains; we recap only what this paper requires. The Universal Computational Genome develops the computational-biology mapping from the entity-system side — ribosome as evaluator, genome as program, abiogenesis as bootstrap. Information as Substrate develops the philosophical reading. This paper is the biology-direction complement: from established biology, through structural decomposition, back to where the methodology connects.

1.1. What This Paper Is and Is Not

This paper is an applied methodology demonstration. It is not:

What the paper is: a worked application showing that the methodology produces a coherent, literature-aligned decomposition of abiogenesis, organized around shared structural vocabulary (primitives, partial levels, dependencies, phase transitions, crystallization, autocatalytic spirals). The decomposition is the contribution; biological-theory adjudication is not.

1.2. Scope Classification

The methodology distinguishes claims by scope (see A Structural Methodology for Information System Domains): structural claims (Sc0) about what can exist; mechanism claims (Sc1) about what physical processes operate within the structural constraints; specific-realization claims (Sc2 and higher) about what did happen on Earth. This paper operates primarily at Sc0 and Sc1. We tag claims as we go; Sc2 claims (specific historical trajectories) lean on the empirical literature.

1.3. What This Paper Does Not Cover

We assume familiarity with the methodology’s vocabulary; specific terms (primitive, partial level, coherent sub-lattice, Hasse walk, SSA topology, autocatalytic spiral, crystallization) are used without re-defining them.

2. The Biology Substrate Domain

We summarize the biology substrate’s primitive set as established in the methodology corpus. The full analysis is in the source material; we recap only what the abiogenesis analysis requires.

2.1. The Six Primitives

The biology substrate (cellular life as the arrangement) decomposes into six primitives at the resolution at which the analysis is stable:

# Primitive Abbrev What it is
1 Genome G Heritable information storage (DNA/RNA sequence)
2 Types T Molecular structure categories (protein folds, RNA structures, metabolites)
3 Ribosome R The evaluator that translates genome encoding into protein products
4 Proteins P Functional products of translation (enzymes, structural, regulatory)
5 Regulation Reg Control logic governing gene expression timing and location
6 Membrane Mem Physical boundary defining self vs environment

Five of the six pass the methodology’s three-test criterion (structural minimality, compositional productivity, empirical recurrence) cleanly. Regulation (Reg) is a borderline primitive: arguments exist for collapsing it into G + P (regulation as proteins acting on genome), but the partial-level decompositions of G and P would have to track regulation independently, which is the methodology’s signal that the primitive should stay separate.

2.2. Dependencies and the Coherent Sub-Lattice

The dependency structure:

The coarse coherent sub-lattice (presence/absence over 26=642^6 = 64 subsets) reduces to roughly 12-15% of the full lattice satisfying all dependencies, in the substrate-typical range (see A Structural Methodology for Information System Domains).

2.3. The Core Triad and SSA Mapping

The core triad is {G,T,R}\{G, T, R\}: genome, types, ribosome. Heritable, typed, evaluated information. Everything else in the biological SSA depends on this triad activating.

The SSA topology (the seven-role information-substrate pattern recurring across domains A Structural Methodology for Information System Domains) maps to biology as follows:

SSA role Biology
Encoding (En) Genome
Evaluator (Vr) Ribosome
Mechanism (Mc) Proteins / enzymes
Surface (Sf) Organism
Context (Cx) Environment
Community (Cm) Population / species
Selection (Se) Natural selection

The mapping is one-to-one and tight. The structural observation that organizes the abiogenesis question: the ribosome is the evaluator, and abiogenesis is the question of how the evaluator arises. We pursue this question structurally.

3. The R0-to-R2 Transition at Molecular Resolution

The coarse decomposition treats abiogenesis as a single qualitative transition: R0 (no translation) \to R1 (proto-ribosome) \to R2 (universal code). Zooming in reveals eight sub-levels with named molecular configurations, internal dependencies, and phase transitions. We treat the sub-level decomposition as the load-bearing structural finding.

3.1. R0: No Translation

RNA oligomers, ribozymes, free amino acids in mineral-catalyzed solution. Information and catalytic function are both present but in the same medium (RNA). No connection between RNA sequences and amino acid sequences. No code, no evaluator separate from the encoding.

Bridge position (the chemistry-to-biology bridge primitives from the methodology’s bridge analysis): Cd=0 (no code), Cat=1 (mineral and ribozyme catalysis active), Fx=2 (geochemical free-energy gradients drive reactions).

3.2. R0.1: Stereochemical Association

RNA aptamers — short RNA sequences — bind specific amino acids by chemical affinity. The Yarus-laboratory finding is that aptamers selected for amino-acid binding are enriched for the codons assigning those amino acids in the modern code. This is not a code; it is a precondition for one. The physical basis for the future code exists in chemistry before any encoding mechanism.

3.3. R0.2: Aminoacylated RNA (Proto-tRNAs)

Small RNA hairpins (35-40 nucleotides) stably attached to specific amino acids by ribozyme-catalyzed aminoacylation (demonstrated experimentally by the Suga laboratory). Two to four distinct aminoacyl-RNA species coexist. This is Crick’s adaptor principle in embryonic chemical form: the adaptor (proto-tRNA) holds an amino acid in a position determined by its RNA sequence. Template-directed synthesis has not happened yet.

3.4. R0.5: Template-Directed Peptide Synthesis

An RNA template positions aminoacyl-RNAs in sequence through codon-anticodon pairing. Short peptides (3-8 amino acids) are produced with crude fidelity (~60-70% per position). The template is the machine: there is no separate evaluator. Encoding and evaluation are fused in a single RNA molecule.

This is a critical structural point. The SSA topology assumes the encoding and the evaluator are distinguishable entities. At R0.5, they are not. Genesis has two qualitative phases: architectural genesis at R1 (where the evaluator separates from the encoding) and functional genesis at R2 (where the evaluator reaches the determinism level Kd4 A Structural Methodology for Information System Domains).

3.5. R1: Proto-Ribosome (Evaluator Separates)

The proto-ribosome (Yonath group) is a dimeric RNA cage of approximately 120-160 nucleotides, formed by two symmetric halves of 60-80 nucleotides each. It catalyzes peptide-bond formation through entropic catalysis (precise positioning of aminoacyl-tRNAs reduces the activation entropy of peptide-bond formation). Three molecular species now cooperate: proto-mRNA (template), proto-tRNAs (adaptors), proto-ribosome (catalyst).

R1 is the defining structural event of the genesis transition. The evaluator separates from the encoding. The SSA topology first applies in its standard form: encoding, evaluator, and adaptors are three distinguishable molecular entities, and the system has the seven-role structure that the methodology recognizes across information substrates.

Empirical status: in 2024, three independent research groups confirmed that dimeric proto-ribosome analogues spontaneously fold, dimerize, and catalyze peptide bonds. R1’s structural prediction (that the proto-ribosome is a dimer of ~60-80 nucleotide halves) is no longer speculative.

3.6. R1.3: Bootstrap Loop Activates

The proto-ribosome produces short peptides; some peptides — by chance — improve the proto-ribosome’s function. The feedback structure:

This is an autocatalytic spiral (see A Structural Methodology for Information System Domains): two primitives (R, the evaluator’s fidelity; and P, the protein products) co-advance through a feedback loop. The spiral is dynamically distinct from monotone single-primitive advancement; it is what we call an autocatalytic spiral in the methodology’s vocabulary.

3.7. R1.7: Fidelity Threshold and Compartmentalization

The bootstrap loop has a critical fidelity threshold. Below approximately 80% per-position fidelity, useful peptides are too rare to sustain the loop: the probability of producing a correctly-translated peptide of length 10 is (0.80)1011%(0.80)^{10} \approx 11\%, of length 15 is 3%\approx 3\%, of length 20 is 1%\approx 1\%. Above approximately 90% per-position fidelity, the production of useful peptides becomes regular: (0.90)2012%(0.90)^{20} \approx 12\%, (0.95)4013%(0.95)^{40} \approx 13\%. The transition from sub-threshold to above-threshold is a dynamical phase transition: linear-tricky-to-self-sustaining.

The ~90% threshold is the methodology’s structural prediction. It is not directly measured; it is inferred from the requirement that the bootstrap loop be self-sustaining and from the minimum length of functional protein domains (proto-aaRS peptides are estimated at 15-25 amino acids, proto-chaperones at 10-20). The threshold is at Sc1: a mechanism claim within structural constraints, consistent with the established Eigen error-limit argument.

Approaching the threshold, a new problem emerges. In an open molecular pool, parasitic RNA (sequences that replicate but do not contribute to translation) outgrows functional RNA. The classical Eigen error catastrophe applies: at ribozyme replication fidelity, the maximum maintainable genome is on the order of 100-200 nucleotides. The bootstrap loop’s needed length (proto-ribosome plus proto-tRNAs plus the proto-aaRS sequences) exceeds this; the system cannot reach R2 in an open pool.

The solution is compartmentalization. Vesicles enclose proto-ribosome systems; selection operates on vesicles (vesicles with better ribosomes grow faster); parasites are excluded by membrane boundaries. The methodology captures this as a conditional partial-level dependency:

Dep(R1.7,Mem1)\mathrm{Dep}(R \geq 1.7, \, \mathrm{Mem} \geq 1)

The bootstrap loop cannot cross its fidelity threshold until compartmentalization is in place. This dependency is invisible at coarse resolution; it appears only at sub-level resolution.

At R1.7, the structural landscape changes: protocell populations with variation, heredity, and differential reproduction exist. The biological landscape appears during the transition, not at its completion.

3.8. R1.9: Code Expansion

The genetic code expands from 4-5 prebiotically-available amino acids (Glycine, Alanine, Valine, Aspartate, Glutamate) through biosynthetically-derived intermediates to the full set of 20. Two structurally unrelated aminoacyl-tRNA-synthetase classes diverge (Class I and Class II, each handling roughly half the amino acids). DNA replaces RNA as the primary storage medium (DNA is more chemically stable). Protein enzymes replace ribozymes for most catalytic functions.

The code expands by internal bootstrap: each new amino acid requires biosynthetic enzymes constructed from amino acids already in the code. Phase 1 amino acids are prebiotically available; Phase 2 are biosynthesized from Phase 1 via one or two enzymatic steps; Phase 3 require multi-step pathways using Phase 1 and Phase 2 enzymes. A 2024 reconstruction of recruitment order from LUCA’s protein domains is consistent with an internally-bootstrapped expansion, while revising the precise ordering of the consensus biosynthetic sequence (placing small and metal- or sulfur-binding residues earlier than the older metrics did).

3.9. R2: Code Crystallization

The genetic code freezes. Sixty-four codons, twenty amino acids, three stop signals, plus the reading frame. Error-minimizing structure (single-nucleotide mutations tend to produce chemically similar amino acids; the probability of this property by chance is less than 10610^{-6}). The code is universal across bacteria, archaea, and eukaryotes.

The crystallization mechanism is self-referential circularity: the code encodes the ribosomal proteins, the tRNA genes, and the aaRS genes that read the code. Code and reading machinery are mutually dependent. Changing any codon assignment misreads every gene that uses the affected codon — lethal when thousands of genes depend on the code. The circularity is the lock.

Crystallization is a new stability type in the methodology’s vocabulary (see A Structural Methodology for Information System Domains):

3.10. Sub-Level Summary

Sub-level Configuration Key event Status
R0 RNA oligomers + free amino acids Sc0 (established chemistry)
R0.1 RNA aptamers bind amino acids Stereochemical association Sc1 (Yarus laboratory)
R0.2 Aminoacylated RNA hairpins Adaptor principle Sc1 (Suga laboratory)
R0.5 Template-directed peptide synthesis En/Vr fused Sc1
R1 Proto-ribosome dimer Evaluator SEPARATES Sc1 (2024 experimental confirmation)
R1.3 Bootstrap loop activates Autocatalytic spiral begins Sc1
R1.7 Threshold + compartmentalization Conditional dependency activates Sc0/Sc1
R1.9 20 amino acids, two aaRS classes Code expansion Sc1 (2024 LUCA-domain reconstruction)
R2 Standard genetic code Code CRYSTALLIZES Sc0 (universally observed)
Abiogenesis trajectory comprehensive view. Top panel: structural corridor width versus rate-weighted effective width across the R0 → R2-LUCA ranks. Middle panel: conjunction-floor heatmap showing the minimum level of each primitive forced at every corridor rank; the manifestation overlay at each Mn rank shows actual levels exceeding the floor in red. Bottom panel: SSA-role hardening fraction per Mn across the seven SSA roles (En, Vr, Mc, Sf, Cx, Cm, Se).

4. The Bootstrap Loop and Autocatalytic Spiral

The bootstrap loop is the central mechanism of the R1-to-R2 transition. We treat it in detail because it is a new dynamical pattern for the methodology: not monotone advancement of a single primitive, but two primitives co-advancing through coupled feedback.

4.1. The Feedback Structure

The structure schematically:

Proto-ribosome at fidelity f → produces short peptides
  → some peptides (RNA-binding) stabilize the ribosome → ribosome fidelity rises
    → improved ribosome produces longer/better peptides
      → some peptides (proto-aaRS) improve aminoacylation accuracy
        → improved aminoacylation increases ribosome's effective fidelity
          → ... (spiral continues)

Three classes of bootstrap peptide drive the spiral:

The dependency ordering among the three classes is structural: RNA-binding peptides come first (shortest, most reliable); proto-chaperones next; proto-aaRS last. Each class enables the next by improving the ribosome’s effective fidelity.

4.2. The Fidelity Threshold as Phase Transition

The transition from sub-threshold to above-threshold is the dynamical phase transition that gives the bootstrap its character. Below threshold, the loop is a trickle: useful peptides are produced occasionally, but not frequently enough to sustain improvement against degradation and chemical noise. Above threshold, the loop is self-amplifying: useful peptides are produced reliably enough to drive ribosome improvement, which produces more useful peptides.

The methodology’s standard partial-level framework assumes monotone single-primitive advancement (level nn to level n+1n+1 in one primitive). The bootstrap loop is qualitatively different: two primitives are spiraling upward together, with a threshold beyond which the spiral becomes self-amplifying. We add autocatalytic spiral to the methodology’s vocabulary (see A Structural Methodology for Information System Domains) for this pattern:

The autocatalytic spiral may be specific to evolved (rather than designed) genesis. Designed systems do not require the spiral: a designer can install the evaluator at fidelity Kd4 from the start. Evolved systems require it because the high-fidelity evaluator must be constructed by the spiral itself — there is no external source for it.

4.3. Error-Rate Mathematics

For a peptide of length LL at per-position fidelity ff, the probability of producing the full-length correct sequence is fLf^L. Several reference values:

The “threshold” is not a single fidelity value; it is the locus where fLusefulf^{L_{\mathrm{useful}}} exceeds the rate at which the system can lose useful peptides to degradation and noise. The ~90% threshold quoted earlier corresponds to producing ~10% useful peptides at L20L \approx 20 amino acids (the proto-aaRS length range), which is approximately where the spiral becomes self-amplifying under reasonable assumptions about peptide turnover.

This is a Sc1 mechanism estimate. The actual threshold depends on the minimum functional peptide length, the rate of useful-peptide production required to drive ribosome improvement, and the rate of peptide loss. Each of these is determined by the proto-ribosomal fitness landscape, which is empirically incompletely characterized.

5. Physical Compartmentalization

The compartmentalization requirement Dep(R1.7,Mem1)\mathrm{Dep}(R \geq 1.7, \mathrm{Mem} \geq 1) is a structural dependency the methodology surfaces. Its biological content is the classical Eigen error-limit constraint: at proto-ribozyme replication fidelity, the maximum maintainable genome length is below what the bootstrap requires. Without compartmentalization, parasitic RNA (short, fast-replicating, non-functional) outgrows functional RNA in any open pool.

5.1. The Diffusion Problem

There is also a diffusion problem. A 50-nucleotide RNA in open water diffuses approximately 1 mm/s. Components disperse before they can interact at the scales required for the bootstrap loop’s repeated encounters between proto-ribosome, proto-tRNAs, and template RNA. The genesis transition cannot occur in unconfined solution. Physical confinement is a precondition.

5.2. The Compartmentalization Sub-Levels

We extend the methodology’s primitive-decomposition treatment to the membrane primitive. Mem decomposes into four partial levels:

Level Description Type Provider
Mem 0 (Cmp 0) No compartment
Mem 0.5 (Cmp 0.5) Mineral micropore Physical confinement Context (the vent)
Mem 1 (Cmp 1) Lipid vesicle Self-assembling chemistry Chemistry
Mem 2 (Cmp 2) Selective membrane Active biology Biology (membrane proteins)

The key transition is Mem 0.5 \to Mem 1: from context-provided confinement (the vent’s mineral structure happens to provide it) to self-generated boundary (the chemistry produces its own vesicle). This is the transition from depending on external physical structure to producing the structure internally.

5.3. Alkaline Hydrothermal Vents

The Russell-Martin alkaline hydrothermal vent hypothesis is the standard biological framing for Mem 0.5. Alkaline vents on the Hadean ocean floor contain labyrinths of mineral micropores (1-100 micrometer diameter) with FeS / Fe(Ni)S walls. These structures provide:

The Mem 0.5 to Mem 1 transition is the vent-to-ocean transition: vesicles form inside micropores, grow, escape into open water, and become self-sustaining protocells.

5.4. The Scaffolding Pattern

A general structural pattern: context-provided structure precedes self-generated structure. The mineral micropore is a scaffold. It provides physical confinement that allows chemistry to produce lipid vesicles, which then replace the mineral scaffold with self-generated boundaries. The vent enables the chemistry that escapes the vent.

This is a recurring structural pattern across substrate origins: the substrate’s eventual self-generation is bootstrapped by environmental conditions that the substrate later supersedes. The pattern’s general form is worth marking; we encounter it again in cognition (cultural scaffolding by adult speakers precedes a child’s self-generated language) and in computation (bootstrap evaluators externally compiled before the substrate compiles its own evaluators, The Entity Machine Boundary). We do not pursue the general pattern at length here; it is a candidate Layer-3 abstraction (see A Structural Methodology for Information System Domains).

5.5. The Landscape at R0-R0.5

The Hadean ocean floor at R0-R0.5 is not “the early Earth” understood as a single environment. It is a population of alkaline vent micropores distributed across thousands of vents over hundreds of thousands of square kilometers. Each micropore is a separate experiment. The stochastic search for the genesis transition runs in parallel across millions of independent micro-reactors.

This reframing matters for the probability analysis below: the search space is not “early Earth as one experiment” but “millions of micro-reactors in parallel,” which changes the relevant probability bounds by many orders of magnitude.

6. Code Crystallization and Competitive Exclusion

The R2 transition has two coupled components: crystallization (the code freezes) and competitive exclusion (the hardened SSA monopolizes the chemical substrate). They form a ratchet: crystallization produces the efficiency differential; competitive exclusion converts the differential into permanent monopoly.

6.1. Crystallization as Phase Transition

The crystallization of the code at R2 is a phase transition with three properties (see A Structural Methodology for Information System Domains):

The mechanism of crystallization is self-referential circularity. Three layers of mutual dependency are visible:

Changing any codon assignment misreads the genes for the components that read that codon. The change is self-amplifying lethal: the more the code is used, the more thoroughly any change destroys the cell. The freeze is the structural fixed point of this self-reference.

6.2. Near-Optimal Error Minimization

The standard genetic code is not randomly assigned. Single-nucleotide mutations tend to produce chemically similar amino acids (a leucine mutating to isoleucine, both hydrophobic; an aspartate mutating to glutamate, both acidic). The probability of obtaining this error-minimizing structure by chance is less than 10610^{-6}.

The code was refined by selection during expansion (R1.3-R1.9), then frozen at R2. The freeze locks in whatever error structure had emerged at the freeze point; that structure is what selection produced over the expansion phase, not a chance arrangement.

6.3. Competitive Exclusion

After the code crystallizes, the hardened biological SSA monopolizes the chemical substrate. Four mechanisms operate together:

6.4. The Coupled Ratchet

Crystallization produces the efficiency differential between cellular life and chemical proto-SSA; competitive exclusion converts the differential into monopoly. Together they form an irreversible ratchet that explains four properties of the biosphere:

This explains why we observe one code, one tree of life, one set of bootstrap types in the cellular substrate. The coupled ratchet is the structural reason.

7. The Genetic Code as Sub-Domain

The methodology is recursively applicable: a primitive of a domain can itself be analyzed as a sub-domain. We apply this to the genetic code, treating it as a six-primitive sub-domain in its own right. The analysis demonstrates the methodology’s scale-invariance and produces an independent structural account of the code that aligns with the abiogenesis decomposition above.

7.1. The Six Code Primitives

The genetic code decomposes into six primitives at the resolution at which the analysis is stable:

# Primitive Abbrev What it is Role in the code
1 Symbol Sm The codon (nucleotide triplet) What specifies
2 Referent Rf The amino acid What is specified
3 Adaptor Ad The tRNA How symbol connects to referent
4 Charger Ch The aaRS enzyme How assignments are established
5 Degeneracy Dg Redundancy structure (multiple codons per amino acid) How errors are tolerated
6 Frame Fr Reading context (start/stop codons, frame) How messages are delimited

7.2. Dependency Structure

Two independent roots: Sm (codons exist in RNA whether or not they are read) and Rf (amino acids exist independently of the code). The other four primitives depend on these roots:

7.3. Filter and Triads

The coherent sub-lattice contains 14 of 64 possible subsets: a filter of 21.9%. This is tighter than surface domains (typically 25-40%) and looser than substrates (12-15%) (see A Structural Methodology for Information System Domains). The bridge-like character is consistent with the code’s structural role: it bridges encoding (the Sm side) to function (the Rf side).

The core triad is {\{Sm, Rf, Ad}\} — symbol, referent, adaptor. The minimal set for a code: something that specifies, something specified, and something that connects them.

A load-bearing quad is {\{Sm, Rf, Ad, Ch}\}: adding Charger gives deterministic translation. With all four active, every codon has a deterministic amino-acid assignment maintained by the charger enzymes.

7.4. The 2+2+2 Structure

The six code primitives organize naturally into three pairs by functional role:

This 2+2+2 structure may be a structural property shared by all information codes (the entity-system’s protocol, natural-language grammar, source-code-to-machine-code translation). The methodology’s standard cross-domain pattern-extraction step (see A Structural Methodology for Information System Domains) would test this hypothesis by applying the procedure to additional code domains. We mark the 2+2+2 hypothesis as a candidate Layer-3 abstraction; the methodology requires further domains to confirm.

7.5. Self-Referential Encoding

The code’s distinctive emergent property is self-referential encoding. The code encodes the machinery that reads the code: ribosomal protein genes, tRNA genes, and aaRS genes are all translated using the code they help implement. The closure is structural — the code is a fixed point of its own translation function.

This is what crystallizes at R2. The self-referential structure produces the mutual dependency that makes the code unchangeable. The methodology’s “crystallization through self-reference” pattern is on full display.

7.6. Code Expansion Trajectory

The code expanded through four phases tracked by the methodology’s partial-level decomposition:

The sequential dependency is structural: each phase’s biosynthetic enzymes are built from earlier phases’ amino acids. The expansion is internally bootstrapped — the code builds the machinery that allows it to grow. The 2024 LUCA-domain reconstruction of recruitment order is consistent with this internal-bootstrap picture, though it revises the consensus on which amino acids came first.

8. Probabilistic Walk Analysis

The methodology’s product-lattice structure plus the partial-level dependencies define a coherent sub-lattice through which the historical trajectory moves. The trajectory is a Hasse walk: a monotone path from the empty position to the fully populated position (see A Structural Methodology for Information System Domains). The probability analysis treats this walk as a stochastic process.

8.1. The Probability Funnel

The actual historical walk through the lattice traces a path whose distribution shape varies across phases. We describe the distribution shape as a “funnel” — wide where the walk has many options, narrow where it is constrained.

Phase Distribution width What constrains it
Pre-R0 (prebiotic chemistry) Very wide Many possible chemistries, environments
R0 to R0.5 Narrowing Template chemistry constrains molecular options
R0.5 to R1 Moderate Proto-ribosome fold constrains structure
R1 to R1.7 Narrowing fast Autocatalytic spiral channels the walk
R2 Very narrow Known endpoint — universal code
R2 to LUCA Broadening Diversification within the attractor
LUCA to eukaryogenesis Narrowing One-time endosymbiosis event
Post-eukaryogenesis Alternating Narrow at phase transitions, wide at radiations

The funnel is narrowest at crystallization events (R2 is the most constrained point in the entire walk because the endpoint is known) and widest at diversification events (post-R2 prokaryotic radiation; post-eukaryogenesis lineage diversification).

8.2. Forward and Reverse Walks

The methodology supports walks in two directions through the lattice.

Forward walks start from R0 (the empty or near-empty position) and apply transition operators step by step. The distribution branches as each step admits multiple successor positions. For abiogenesis-as-prediction, the forward walk would compute the distribution over possible historical trajectories given the structural constraints. This is the planning direction.

Reverse walks start from R2 (the known endpoint) and work backwards through the structural constraints. The distribution converges as each step is constrained by the structures the endpoint requires: the PTC symmetry constrains R1 to a dimeric RNA configuration; the code universality constrains R2 to a single crystallization event; the dependency structure constrains the ordering of intermediate transitions. This is the reconstruction direction — the natural mode for historical analysis where the endpoint is known but the intermediate positions must be inferred.

The forward-backward intersection gives the high-probability corridor through the lattice. Where the forward distribution and the reverse distribution overlap strongly is where the actual history most likely passed. The Bayesian formulation:

P(Xtevidence1:T)αt(Xt)βt(Xt) P(X_t \mid \mathrm{evidence}_{1:T}) \propto \alpha_t(X_t) \, \beta_t(X_t)

where αt(s)=P(Xt=s,evidence1:t)\alpha_t(s) = P(X_t = s, \mathrm{evidence}_{1:t}) is the forward variable (probability of reaching state ss at step tt given evidence to step tt) and βt(s)=P(evidencet+1:TXt=s)\beta_t(s) = P(\mathrm{evidence}_{t+1:T} \mid X_t = s) is the backward variable (probability that the remaining evidence is observed given state ss at step tt).

8.3. Confidence Gradient

The confidence gradient is asymmetric. Near R2, confidence is high: the endpoint is known with strong empirical support (universal code, PTC symmetry, LUCA reconstruction). Near R0, confidence is low: prebiotic chemistry admits many possible configurations and the empirical record is sparse. Structural claims (Sc0) remain high-confidence regardless of position; mechanism claims (Sc1) are more confident near R2 and less confident near R0; specific-realization claims (Sc2) require empirical observation at each position.

8.4. Calibration

A Structural Methodology for Information System Domains develops a calibration architecture for attaching empirical wall-time anchors to rate models. Applied to abiogenesis with the LUCA-emergence anchor at approximately 4.2 Gya (Moody et al. 2024), the calibration produces a predicted cumulative wall time of approximately 855 million years for the R0-to-R2 transition. Whether 855 My fits the available Hadean window depends on which habitability anchor is taken: it sits inside the generous 500 My–1 Gy estimate, but exceeds the ~200 My window implied by a 4.4 Gya habitability onset against a 4.2 Gya LUCA (the tension is taken up under “Tensions” below). The calibration is the methodology’s mechanism for connecting structural-level results to wall-clock empirical anchors; we cite it here without re-deriving.

The 855 My value is a calibration output, not an independent measurement. Its consistency with the empirical window is corroborative, not validating: the calibration is anchored to the LUCA estimate, so the result is bounded by that anchor. The structural claim is that the dependency-filtered sub-lattice plus the rate model produce a wall-time prediction within the independently-derived geological window — the structural decomposition does not contradict the geological constraints.

Abiogenesis paper-ready extended view: corridor + manifestations + rate + SSA roles + per-Mn population context (top); SSA hardening trajectory (middle); sensitivity comparison under pessimistic / mid / optimistic population scenarios (bottom). The top-to-bottom composition is the analyst’s deep dive — structural corridor through SSA hardening through population sensitivity.

9. Literature Alignment

The methodology’s account of abiogenesis aligns with established research programs across multiple fronts. We list the alignments and note where the methodology extends or tensions exist.

9.1. Strong Alignment

Proto-ribosome hypothesis (Yonath group). Our R1 is the Yonath proto-ribosome. Three independent groups confirmed in 2024 that dimeric proto-ribosome analogues spontaneously fold and catalyze peptide bonds (Multiple research groups 2024). The structural prediction (R1 as dimer of ~60-80 nt halves) is no longer speculative.

RNA world hypothesis. Our R0-R0.5 sub-levels map onto the standard RNA-world narrative. The methodology’s decomposition is compatible with, not competitive against, the RNA-world framing.

Protocell research (Szostak laboratory). Our Mem 1 is the Szostak-lab protocell. Vesicle growth, division, and RNA encapsulation are demonstrated experimentally (Szostak 2009).

Eigen error catastrophe. Our conditional dependency Dep(R1.7,Mem1)\mathrm{Dep}(R \geq 1.7, \mathrm{Mem} \geq 1) formalizes the Eigen limit as a structural cross-primitive constraint. The mathematical content is the same; the structural framing is the methodology’s contribution.

LUCA reconstruction. Moody et al. (2024) place LUCA at approximately 4.2 Gya with approximately 2,500 genes (Moody et al. 2024). The complexity (~2,500 genes) places LUCA at a substantial cellular configuration; the methodology’s R2 is the crystallization event, with LUCA arriving subsequently in the “post-R2 broadening” phase of the funnel.

Rapid abiogenesis. Bayesian analyses of life’s early appearance report odds favoring rapid over slow-and-rare abiogenesis — in the range of roughly 3:1 to 9:1 depending on which early-life date is used, short of the conventional 10:1 “strong evidence” bar. The direction is consistent with the methodology’s prediction: the transition is context-gated (it requires the alkaline-vent micropore environment) but fast once unblocked, because the autocatalytic spiral, above its threshold, is self-amplifying.

Genetic code evolution. A 2024 reconstruction of amino-acid recruitment order from LUCA’s protein domains supports an internally-bootstrapped, dependency-ordered expansion — while revising which residues entered first (small and metal- or sulfur-binding amino acids earlier than the older consensus). The methodology’s R1.9 code-expansion sub-level is the structural counterpart; it predicts a dependency-ordered expansion without committing to a specific recruitment sequence.

9.2. Where the Framework Extends

Unified sub-level framework. No published equivalent connects RNA-world chemistry, proto-ribosome structural biology, protocell biophysics, code-origin theory, and LUCA reconstruction in a single decomposition with shared vocabulary. Each research program has its own framing; the methodology’s sub-level sequence is what connects them.

The fidelity threshold at ~90%. The Eigen error limit is well-known; the specific bootstrap-self-amplification threshold is a methodology-derived structural prediction. It is consistent with what is known but is not in the literature as a quantitative target.

The Mem 0.5 sub-level. Mineral micropores as a separate compartmentalization level (distinct from lipid vesicles) is the methodology’s contribution. The Russell-Martin framing of alkaline vents has the same content; the methodology’s partial-level decomposition makes Mem 0.5 a named structural step rather than a contextual factor.

The autocatalytic-spiral pattern. Autocatalytic networks are well-studied (Kauffman, RAF theory); the specific two-primitive co-advancement pattern with critical threshold is the methodology’s framing. The pattern’s general form (two primitives, feedback loop, threshold, dynamical phase transition) is added to the methodology’s vocabulary.

The crystallization-plus-monopoly ratchet. Code universality is observed; the structural mechanism (crystallization through self-reference plus competitive exclusion) is the methodology’s account of why the universality is permanent.

9.3. Tensions

LUCA timing. If LUCA is at 4.2 Gya and Earth became habitable at ~4.4 Gya, the available window is only ~200 My, which is substantially shorter than the ~855 My cumulative wall time the calibration produces for the R0-to-R2 transition (see “Calibration” above). The decomposition is independent of absolute timing (the sub-level sequence and dependencies are scale-invariant), but the rate calibration may need compression to fit a 200 My window — or, equivalently, the rate weights in the placeholder kinetic model may need revision against tighter Hadean-habitability anchors. The structural decomposition stands either way; the wall-time estimate is calibration-bound.

Symbiotic / parasitic ribosome origin. A recent perspective suggests the proto-ribosome may have begun as an external parasite — a selfish replicator that invaded protocells and co-evolved into an obligate symbiont — rather than arising as an internal product of the host chemistry. If correct, the R0.5-to-R1 dynamics change (the proto-ribosome arrives via invasion rather than internal search), but the structural sequence (En/Vr fused \to separated \to deterministic) holds regardless.

LUCA complexity. Approximately 2,500 genes places LUCA at a higher lattice position than a “minimal free-living cell.” The first attractor may be at a higher position than initially estimated; the funnel’s post-R2 broadening is correspondingly delayed.

10. Discussion

10.1. What the Methodology Adds

The methodology adds a structural decomposition organized around shared vocabulary that the existing literature lacks. Specific contributions:

What the methodology does not add: new biological mechanism (the mechanisms are all established in the literature); empirical predictions that distinguish among competing biological hypotheses (the methodology is compatible with several framings, not selective among them).

10.2. What the Methodology Does Not Resolve

Several open questions remain open after the methodology applies:

10.3. What This Paper Suggests for Methodology Application

This paper is one applied-methodology demonstration. Several patterns recur and may be useful for future applications:

These observations are tentative; one applied case is not a pattern. We mark them for future applied-methodology papers.

10.4. Limitations

Several limitations should be noted.

11. Conclusion

Abiogenesis is not the creation of life from non-life. It is the progressive hardening of feedback cycles that already exist in chemistry. The SSA topology — encoding, evaluator, mechanism, surface, context, community, selection — operates in soft chemical form from the earliest mineral-catalyzed reactions. The genesis transition makes the roles deterministic, dedicated, compartmentalized, and permanent.

The methodology’s decomposition reveals eight sub-levels of the R0-to-R2 transition (R0, R0.1, R0.2, R0.5, R1, R1.3, R1.7, R1.9, R2), each with named molecular configurations. The defining structural event is R1: the proto-ribosome separates from the encoding, and the SSA topology applies in its standard form for the first time.

The bootstrap loop is the central mechanism: two primitives (the evaluator’s fidelity R, the protein products P) co-advance through coupled feedback with a critical fidelity threshold around 90%. We add the autocatalytic spiral to the methodology’s vocabulary for this dynamical pattern.

Compartmentalization is a structural prerequisite: the conditional partial-level dependency Dep(R1.7,Mem1)\mathrm{Dep}(R \geq 1.7, \mathrm{Mem} \geq 1) formalizes the classical Eigen error-limit constraint. The vent-to-ocean transition is the Mem 0.5 (mineral micropore, context-provided) to Mem 1 (lipid vesicle, self-generated) transition.

The R2 crystallization is a phase transition: the code freezes through self-referential circularity. The hardened SSA then monopolizes the chemical substrate by competitive exclusion, forming an irreversible ratchet that explains code universality, LUCA singularity, and the impossibility of second genesis on Earth.

The genetic code, as a six-primitive sub-domain (Symbol, Referent, Adaptor, Charger, Degeneracy, Frame), has a 21.9% filter, a core triad of {\{Sm, Rf, Ad}\}, and a 2+2+2 functional organization (WHAT, HOW, ROBUSTNESS). The code’s self-referential encoding is what crystallizes; the crystallization mechanism is what freezes the substrate.

A probability funnel organizes the historical walk: wide at R0, narrowing through the bootstrap, collapsing at R2, broadening post-R2. Forward walks from R0 widen the distribution; reverse walks from R2 narrow it; the intersection is the high-probability corridor through the lattice.

The decomposition aligns with established research: proto-ribosome experimental confirmation in 2024, Szostak-lab protocells, Russell-Martin alkaline vents, the Eigen error limit, LUCA reconstruction at 4.2 Gya, and the biosynthetic-order code expansion. Where the methodology extends the literature, the extension is structural framing (the sub-level decomposition, the conditional dependency, the autocatalytic spiral, the crystallization-monopoly ratchet) rather than new biology.

Several open invitations sit alongside the decomposition. A structural decomposition more compact than the eight-sub-level sequence that retained the empirical fit would refute the irreducibility of these sub-levels. A missing sub-level in the current decomposition would expose a gap. Experimental tests of the ~90% fidelity threshold with proto-ribosome systems at varying fidelity levels would convert the structural inference into a measurement. Applying the methodology to a second information-code domain (the entity-system protocol, natural-language grammar) would test whether the 2+2+2 structure recurs as a Layer-3 invariant.

The methodology’s value, here as in other applied cases, is the structural organization it produces around an open scientific question. The biology is the biologists’; the structural decomposition is what the methodology adds.

Glossary

This glossary collects the controlled vocabulary used across the volume. Terms appear in the order they are first introduced in the foundational paper, The Entity System; cross-references in entries use the same vocabulary.

Primitives

Entity (E)
The unit of information in the system. An entity is a content-addressed, typed datum identified by a hash of its content. Entities are immutable.
Identity (I)
A stable name for a sequence of entities. An identity decouples “what this thing is now” from “what this thing was previously.”
Tree (T)
A structural composition primitive. Trees compose entities into hierarchical structures with addressable paths.
Emit (M)
The temporal primitive. Emit defines the act of producing a new entity and binding it to an identity at a point in logical time.
Execution (X)
The computational primitive. Execution evaluates content-addressed code against content-addressed data, producing content-addressed results.
Peer (P)
The spatial primitive. A peer is a uniform unit of isolation within which entities are stored, identities are resolved, and execution runs.

Composed properties

Self-description
A property emerging at three primitives (E+I+T). The system describes its own structure using the same vocabulary it uses to describe data.
Fixed-point types
The bootstrap-type structure under which types are themselves entities of a small set of “type entities” that refer to each other in a fixed-point closure.
Mutability
A structural property emerging at four primitives (E+I+T+M). Mutability is not a property of entities (which are immutable) but of identities (which may emit successive entities over time).
Computation
The actualisation of latent computational structure that emerges at five primitives (E+I+T+M+X). The substrate becomes Turing-complete via the execution primitive.
Distribution
Emerges at six primitives (E+I+T+M+X+P). Peer adds the spatial dimension that turns a single-machine substrate into a distributed one.

Architectural terms

Substrate
The minimum-floor abstraction over which everything else runs. The six primitives constitute the entity-system substrate.
Substrate-bridge extension
A Tier-1 extension that bridges substrate primitives to an application-architecture surface property. Eleven exist: TREE, TYPE, CONTENT, INBOX, SUBSCRIPTION, CONTINUATION, COMPUTE, QUERY, REVISION, HISTORY, CLOCK.
Operational extension
A Tier-2 extension supplying machinery that the substrate does not itself express: user identity (2a), network (2b), management (2c).
Standard peer
A peer profile under which a uniform set of substrate-bridge extensions is available. The standard peer is the conventional deployment target.
Conformance
The property of an implementation passing the cross-language conformance test suite that validates substrate behaviour across Go, Python, and Rust.

Methodology terms

Partial primitive
A primitive that decomposes into discrete levels (e.g., Sc=0 through Sc=4). Partial primitives admit graded analysis.
Convergence test
A reproducibility check for whether a candidate primitive set in a domain stabilises under iterated reduction.
Coherent sub-lattice
The subset of the power set of a primitive set under which dependency constraints are satisfied. For the entity-system substrate the coherent sub-lattice is 9 of 64 subsets (14%\sim 14\%); for the substrate-bridge extension lattice it is 576 of 2048 (28%\sim 28\%).
Transferability class
A classification of how cleanly a result transfers across substrates. Class N: not transferable. Class S: substrate-specific. Class T: transferable with translation. Class B: substrate-bridging — transfers without translation.
Triangle (composition triangle)
A three-primitive composition with load-bearing structural role. The named triangles in this volume are EIT, ITM, TMX, IXP, TXP.
Layer (1–4)
The scope hierarchy of the structural methodology. Layer 1: domain analysis. Layer 2: cross-domain graph construction. Layer 3: pattern extraction. Layer 4: applied analysis at variable scope ladder Sc=0 through Sc=4.

Conventions

References to other chapters use the form [@paperN] in source, rendered bundle-relatively as “Part M” when the referenced paper appears in the current bundle and as the italicised paper title otherwise. The shared references list appears in the back matter. Section numbering is hierarchical: the part number (the paper’s position in the current bundle) is the leading component (e.g., “3.2.1” is Part 3, Section 2, Subsection 1).

References

Moody ERR, Álvarez-Carretero S, Mahendrarajah TA, et al. 2024. The nature of the last universal common ancestor and its impact on the early Earth system. Nature Ecology & Evolution. [published online ahead of print]
Multiple research groups. 2024. Experimental confirmation of dimeric proto-ribosome analogues. Various journals. [published online ahead of print]
Szostak JW. 2009. Origins of cellular life. Biochimica et Biophysica Acta. [published online ahead of print]