Biology Substrate: Canonical Domain Analysis
Status: Canonical reference. Full 12-step analysis of the biological information substrate — the En primitive of the SSA in the biology arrangement.
Position in the topology: Substrate domain. Realization chain: Physics → Chemistry → Biology substrate → Organism architecture → Ecosystem → Environment (context). Connected downward to chemistry through the chemistry-to-biology bridge (analysis-chemistry-to-biology-bridge.md); upward to organism architecture through the biology-to-organism bridge (deferred). The genetic code, biology's load-bearing crystallization, is analyzed at sub-domain resolution in analysis-genetic-code-sub-domain.md.
Reference (not basis for copying): v1_revision/v1_biology_domain_analysis/bio_v1/biology-domain.md contains the v1-era analysis. This document re-does the analysis under current methodology (12 steps, SSA framework, Layer 4 vocabulary, conditional partial-level dependency notation).
Step 1 — Information Gathering
1.1 What we're analyzing
The biological information substrate — the molecular machinery that life uses to store, read, execute, regulate, and bound its information. Specifically: how DNA stores information; how transcription makes that information available; how the ribosome converts it into proteins; how proteins do biology's work; how regulation controls expression; how membranes define cells.
Not in scope here: organism-level traits (development, physiology, behavior — those belong to organism architecture); ecosystem-level dynamics (competition, predation, coevolution); environmental context (climate, geology). Each is a separate domain in the biology arrangement.
1.2 Sources
- Central dogma — Crick (1958, 1970): directional flow of sequence information
- Operon model — Jacob & Monod (1961): gene regulation as a control system
- Genetic code — Nirenberg, Khorana, Holley (1960s): 64 codons → 20 amino acids
- Ribosome structure — Ramakrishnan, Steitz, Yonath (Nobel 2009): atomic-resolution translation machine
- Epigenetics — Waddington (1942) and modern molecular epigenetics
- Systems biology — Alon (2007): gene regulatory networks as computational circuits, network motifs
- RNA world — Gilbert (1986): RNA as proto-substrate combining storage and catalysis
- Major evolutionary transitions — Maynard Smith & Szathmáry (1995): substrate-level discontinuities
- Reference textbook — Alberts et al., Molecular Biology of the Cell
1.3 The landscape (instances we will position)
| Instance | What it is |
|---|---|
| Viroid | Naked circular RNA; minimal information substrate |
| RNA virus (TMV, polio) | Parasitic; genome + proteins, host-dependent |
| Retrovirus (HIV) | Reverse transcriptase, complex regulation, host-dependent |
| Mycoplasma | Smallest free-living genome (~500 genes) |
| E. coli | Model prokaryote (~4,300 genes, operons) |
| S. cerevisiae | Simple eukaryote, chromatin, organelles |
| C. elegans / Drosophila / Arabidopsis | Multicellular models |
| H. sapiens | Near-full elaboration across all primitives |
Step 2 — Landscape Analysis
2.1 What recurs across all biology
Every cellular organism, from Mycoplasma to mammal, exhibits:
- Persistent information — nucleic acid sequences encoding genes
- Information transfer — RNA polymerase reading DNA into mRNA without consuming it
- Information execution — the ribosome reading mRNA into protein
- Functional output — proteins doing biology's work
- Expression control — regulation of when, where, and how much each gene is expressed
- Boundary — membrane defining the cell and selectively controlling its exchange with the outside
Viruses exhibit subsets of these primitives but require a host that exhibits all of them — viruses are dependency-violating manifestations stabilized by host coupling, not autonomous instances of the substrate.
2.2 Why biology validates the methodology
Biology's primitives were not chosen by an analyst. They were found by 4 billion years of selection on chemical substrates with no human bias. If the methodology's structural patterns (pair-relationships, core triads, lattice positions, attractors) replicate in biology, that replication is not artefactual. Biology is the strongest validation case the methodology has.
Step 1b — Domain Type Declaration
Substrate domain — primitives describe what life IS, not what anyone designs. Predicted properties from R11:
- Tight filter (~12-20%) — confirmed at 12.5%, the tightest of any domain analyzed
- Information-flow core triad — confirmed: {G, T, R} = central dogma = information flow
- Dependency chain depth: deep (the central dogma forms a sequential chain)
Biology is the canonical instance of an information substrate — the En primitive of the SSA. Its evaluator (R, the ribosome) is biology's load-bearing crystallization (Kd4 evaluator determinism). Its Mc primitive (developmental mechanisms, analyzed in the biology-to-organism bridge) connects substrate to organism architecture (Sf in the SSA). Its context root (environment, analyzed separately) provides the SSA's Cx. Its community structure (ecosystem, analyzed separately) provides the SSA's Cm. Its selection mechanism (natural selection, in the ecosystem analysis) provides Se.
This document covers En only. The remainder of the SSA components for the biology arrangement live in their own documents.
Step 3/3b — Primitives and Partial Levels
3.1 Six primitives
| # | Primitive | What it is |
|---|---|---|
| 1 | Genome (G) | Persistent nucleic acid encoding genetic information |
| 2 | Transcription (T) | RNA polymerase reading DNA into mRNA without consuming the source |
| 3 | Translation (R) | Ribosome reading mRNA into protein (the evaluator) |
| 4 | Protein (P) | Folded amino acid chains with specific function |
| 5 | Regulation (Reg) | Machinery controlling when, where, and how much each gene is expressed |
| 6 | Membrane (Mem) | Lipid bilayer boundary with selective permeability |
3.2 Stability under 3/3b iteration
Splitting candidates rejected:
- G into "structure" and "sequence." Rejected — the physical molecule and its information content are inseparable. The base-pairing rules mean structure encodes information and information determines structure. G's internal structure is captured by partial levels (G0 through Full G).
- T into transcription and RNA processing. Rejected — RNA processing (capping, splicing, polyadenylation) is eukaryote-specific elaboration of T, captured at T3. Prokaryotes transcribe without processing; the difference is a partial level, not a new primitive.
- R into ribosome and genetic code. Rejected — the genetic code (codon → amino acid mapping) is implemented BY the ribosome and tRNAs. They are operationally inseparable; the code has no existence independent of translation machinery. The code's internal structure is treated separately as a sub-domain (
analysis-genetic-code-sub-domain.md). - Reg into transcriptional and post-transcriptional. Rejected — both share the same structural role (expression control) and often coordinate. They are levels of the same regulatory gradient.
Collapsing candidates rejected:
- T and R into "Expression." Rejected — they are independently regulable (post-transcriptional regulation acts BETWEEN them), use different molecular machines (RNA polymerase vs. ribosome), and are temporally separate in eukaryotes (decoupled by nuclear envelope). Independent regulability is the decisive test.
- P and Reg into "Function." Rejected — many proteins have no regulatory function (structural, metabolic). Regulation acts ON expression; proteins are the OUTPUT of expression. Different structural roles.
Additions considered and rejected:
- Replication as a 7th primitive. Replication is G-internal — the genome copying itself uses the same primitive (G's internal mechanism), not a separate one.
- Metabolism as a 7th primitive. Metabolism is below the information level — it provides energy, not information structure. Belongs to the chemistry-to-biology bridge (Fx primitive there) rather than to the biology substrate.
Verdict: 6 primitives stable under iteration.
R6 (count-sensitivity honesty): v1 also identified 6 primitives with the same set. Confirmed independently here. Alternatives considered (5: merging T+R; 7: splitting Reg or adding Replication or Metabolism) all reduce to existing primitive structure or fall outside the substrate level.
3.3 Partial levels
Genome (G) — persistent information:
| Level | Description | Distinguishing features | Example |
|---|---|---|---|
| G0 | No genome | No persistent information store | Prions, prebiotic chemistry |
| G1 | Information-carrying polymer | Nucleic acid sequence; no encoded genes | Viroids (~250-400 nt) |
| G2 | Gene-encoding genome | Discrete genes encoding functional products | Simple RNA viruses |
| G3 | Organized multi-gene genome | Operons, gene order, regulatory regions | E. coli (~4.6 Mb) |
| G4 | Chromatinized genome | Histone packaging, epigenetic marks, accessibility | Yeast (~12 Mb), human (~3.2 Gb) |
| Full G | Self-describing genome | Transposable elements, self-regulatory circuits, epigenetic inheritance | Complex eukaryotes |
Phase transition: G1 → G2 (gene encoding). Below: chemistry. Above: information specifies function. The threshold where "molecules become genetic."
Transcription (T) — information transfer:
| Level | Description | Distinguishing features | Example |
|---|---|---|---|
| T0 | No transcription | Genome not read | Prions; some replicators |
| T1 | Simple template copying | Non-enzymatic / simple enzymatic; no regulation of initiation | RNA-world ribozymes |
| T2 | Regulated initiation | Promoters, sigma factors, terminators | Prokaryotes |
| T3 | Processed transcription | Capping, splicing, polyadenylation | Eukaryotes |
| Full T | Complex regulated transcription | Alternative splicing, RNA editing, regulatory ncRNAs | Mammals (~95% multi-exon genes alt-spliced) |
Phase transition: T1 → T2 (regulated initiation). Below: constitutive or none. Above: gene-level information selection.
Translation / Ribosome (R) — the evaluator:
| Level | Description | Distinguishing features | Example |
|---|---|---|---|
| R0 | No translation | No protein synthesis | Viroids; pre-code RNA world |
| R1 | Proto-translation | Simple peptide synthesis, ambiguous code, proto-ribosome | Hypothetical (no extant; supported by ribosome structural studies) |
| R2 | Standard genetic code | Full 64-codon → 20-aa code; tRNAs, aaRSs | All cellular life |
| R3 | Quality-controlled translation | Proofreading, chaperones, mRNA QC | All modern free-living cells |
| Full R | Regulated localized translation | Heterogeneity, mRNA-specific rates, co-translational insertion | Eukaryotes |
Phase transition: R0 → R2 at coarse — the abiogenesis bottleneck.
At sub-resolution (per abiogenesis_analysis_v1/), R0 → R2 decomposes into R0 → R0.5 → R1 → R1.7 → R1.9 → R2 — an autocatalytic spiral coupling R and P fidelity, terminating at the R2 crystallization (universal genetic code, frozen since LUCA). The genetic code is biology's load-bearing crystallization: irreversible, enabling, universal. See analysis-genetic-code-sub-domain.md for the sub-domain analysis.
This parallels X0 → X2 in the entity system (open dispatch — the most consequential transition). Both mark where "the system can convert information to function via a general mechanism."
Protein (P) — functional output:
| Level | Description | Distinguishing features | Example |
|---|---|---|---|
| P0 | No proteins | RNA-only or non-biological | Pre-code; viroids |
| P1 | Simple peptides | Short, low-specificity catalysis | Hypothetical early; antimicrobial peptides |
| P2 | Folded domain proteins | Stable tertiary structure, single domain | Simple enzymes, structural proteins |
| P3 | Multi-domain proteins | Modular architecture, complexes | Bacteria, viruses (multi-domain coat) |
| P4 | Post-translationally modified proteome | Phosphorylation, ubiquitination, glycosylation | Eukaryotes |
| Full P | Proteome-scale networks | Interactomes, signaling cascades, systems-level regulation | Complex eukaryotes |
Phase transition: P1 → P2 (folding). Random peptides → molecular machines.
Regulation (Reg) — expression control:
| Level | Description | Distinguishing features | Example |
|---|---|---|---|
| Reg0 | No regulation | Constitutive expression | Hypothetical |
| Reg1 | Simple feedback | Product inhibition, basic repression/activation | Minimal bacteria |
| Reg2 | Operon-level | Coordinated gene clusters, sigma factors, two-component | E. coli |
| Reg3 | Combinatorial | Multiple TFs per gene, enhancer logic, GRNs as circuits | Eukaryotes |
| Reg4 | Epigenetic | Heritable silencing/activation without sequence change | Multicellular (X-inactivation, imprinting) |
| Full Reg | Developmental programs | Morphogen gradients, cell fate determination, regulation as spatial computation | Complex development |
Phase transitions: Reg2 → Reg3 (combinatorial; cell-type-specificity becomes possible) and Reg3 → Reg4 (epigenetic inheritance; stable differentiation possible).
Membrane (Mem) — boundary:
| Level | Description | Distinguishing features | Example |
|---|---|---|---|
| Mem0 | No membrane | Open system | Pre-cellular; viroids |
| Mem1 | Simple lipid boundary | Bilayer, passive permeability | Protocells; Mycoplasma |
| Mem2 | Selective boundary | Channels, pumps, transporters | Bacteria |
| Mem3 | Compartmentalized interior | Internal organelles (nucleus, ER, mito) | Eukaryotes |
| Mem4 | Multi-cell coordinated boundaries | Cell-cell junctions, ECM, tissue barriers | Multicellular |
| Full Mem | Integrated multi-scale | Organ-level barriers, immune self/non-self | Complex organisms |
Phase transition: Mem2 → Mem3 (eukaryogenesis). Internal compartmentalization. Singular event in Earth history (mitochondrial endosymbiosis).
3.4 Phase transition summary
| Primitive | Levels | Phase transition | Significance |
|---|---|---|---|
| G | 6 | G1 → G2 | Chemistry → biology |
| T | 5 | T1 → T2 | Constitutive → selective |
| R | 5 | R0 → R2 (sub-res spiral) | The abiogenesis bottleneck — biology's load-bearing crystallization |
| P | 6 | P1 → P2 | Peptides → molecular machines |
| Reg | 6 | Reg2 → Reg3, Reg3 → Reg4 | Combinatorial; epigenetic inheritance |
| Mem | 6 | Mem2 → Mem3 | Eukaryogenesis |
Total raw partial-level positions: 6 × 5 × 5 × 6 × 6 × 6 = 32,400. The dependency structure filters this dramatically (see Step 7).
3c Evaluator identification
Biology has an evaluator: R (Translation / Ribosome), at Kd4 (deterministic).
The ribosome reads mRNA codons, charges tRNAs with their cognate amino acids via aminoacyl-tRNA synthetases, and catalyzes peptide bond formation. The mapping from codon to amino acid (the genetic code) is universal, near-optimal for error minimization, and frozen since LUCA (~4.2 Gya). This is the load-bearing crystallization that makes biology biology — an information-substrate evaluator at maximum determinism.
Above R0 → R2 there is no further evaluator-determinism advance. R3 adds quality control (proofreading, chaperones, mRNA quality control) but doesn't change the evaluator's deterministic character. Full R adds regulation around translation but the evaluator itself stays Kd4.
The genetic code's structural analysis is analysis-genetic-code-sub-domain.md. Sub-domain primitives (codon table, tRNA charging, recognition specificity, aaRS class divide) and their crystallization properties are covered there.
Steps 4–6 — Dependencies, Pairs, Load Classification
4.1 Primitive-presence dependencies
T → G transcription reads from a genome
R → T translation reads mRNA produced by transcription
P → R proteins are produced by translation
Reg → G regulation targets genes
Reg → P transcription factors ARE proteins (the bootstrap cycle)
Mem → P membranes require membrane proteins for selective permeability
The Reg → P → R → T → G → Reg cycle is the self-referential bootstrap — the biological equivalent of the entity system's self-description fixed point (types describe themselves). For coherent subset analysis, the cycle resolves by noting Reg requires the full expression chain (G, T, R, P) to produce its TF proteins.
DAG (resolving the cycle):
G ← T ← R ← P ← {Reg, Mem}
Nearly linear — the tightest chain of any analyzed domain. Late branching at {Reg, Mem} (both depend on P; neither depends on the other).
4.2 Conditional partial-level dependencies
In current notation Dep(A ≥ x, B ≥ y):
| Constraint | Reasoning |
|---|---|
| Dep(T ≥ 2, G ≥ 2) | Regulated transcription requires gene-encoding genome |
| Dep(R ≥ 2, T ≥ 2) | Standard translation requires regulated transcription (at minimum, ribosomal gene transcription) |
| Dep(P ≥ 2, R ≥ 2) | Folded proteins require the full genetic code |
| Dep(Reg ≥ 2, P ≥ 2) | Operon regulation requires folded TF proteins |
| Dep(Reg ≥ 4, Mem ≥ 3) | Epigenetic inheritance across cell division requires nuclear compartment |
| Dep(Mem ≥ 3, P ≥ 3) | Organelle biogenesis requires multi-domain proteins |
| Dep(Mem ≥ 4, Reg ≥ 3) | Multi-cell coordination requires combinatorial regulation |
| Dep(Reg ≥ 5, Mem ≥ 4) | Developmental programs (morphogen gradients) require multi-cell coordinated boundaries |
These tighten the coherent sub-lattice at fine resolution.
5.1 Pair enumeration
C(6,2) = 15 pairs.
6.1 Load classification
Heavy (7):
| Pair | Name | Content |
|---|---|---|
| G-T | Transcription machinery | RNA polymerase, promoter recognition, elongation, termination |
| G-R | The genetic code | 64 codons → 20 aa, wobble, degeneracy, codon usage, tRNA charging |
| G-Reg | Gene regulation | Promoters, enhancers, silencers, epigenetics, chromatin remodeling |
| T-Reg | Transcriptional regulation | TF binding, combinatorial control, enhancer logic |
| R-P | Translation fidelity + folding | Codon-anticodon accuracy, chaperone-assisted folding, QC |
| P-Reg | Feedback regulation | Protein products regulating their own genes (the bootstrap loop) |
| P-Mem | Membrane proteins | Channels, transporters, receptors (~30% of all genes) |
Medium (4):
| Pair | Name | Content |
|---|---|---|
| G-P | Genotype-phenotype map | The fundamental question; mostly mediated through T, R, Reg |
| T-R | Expression pipeline | mRNA as intermediate; coupled in prokaryotes, decoupled in eukaryotes |
| T-P | Protein expression levels | Transcription rate → protein abundance; stochastic gene expression |
| Reg-Mem | Spatial regulation | Morphogen gradients, cell-cell signaling, compartment-specific regulation |
Light (4):
| Pair | Name | Content |
|---|---|---|
| G-Mem | Genome compartmentalization | Nuclear envelope, nucleoid organization, viral capsid |
| T-Mem | mRNA export | Nuclear pore transport, mRNA localization (e.g., dendritic) |
| R-Reg | Translational regulation | miRNA, ribosome stalling, mRNA stability, initiation control |
| R-Mem | Membrane-bound translation | SRP pathway, co-translational ER insertion |
Negligible: none. All 15 pairs have biological content. Consistent with biology's tight filter — dense integration means every pair has structural content.
Distribution: 7/4/4/0 (47% heavy). Above the typical ~40% for 6-primitive domains.
6.2 Anchor analysis
- G (Genome): in 4 of 7 heavy pairs (G-T, G-R, G-Reg) plus medium G-P. Primary anchor.
- Reg (Regulation): in 3 of 7 heavy pairs (G-Reg, T-Reg, P-Reg). Secondary co-anchor.
- P (Protein): in 3 of 7 heavy pairs (R-P, P-Reg, P-Mem). Secondary co-anchor.
G is the primary anchor — the genome hub everything connects to. This matches the dependency DAG (G has no dependencies and everything depends on it).
Cross-domain comparison:
| Domain | Primary anchor | What it concentrates |
|---|---|---|
| Biology | G (Genome) | Information storage |
| Chemistry | Rx (Reaction) | Material transformation |
| Entity system | E + I (Entity + Identity) | Self-description |
| Cognition | Rp + Ct (Representation + Context) | Semantic relation |
Substrate domains have single-hub anchors with sequential dependency chains. The hub differs by what the substrate IS.
9.1 Load-bearing compositions
Five firm triangles:
| Triangle | Name | Emergent property |
|---|---|---|
| G-T-R | Central dogma | Gene expression: information → function. Biology's defining property. |
| G-Reg-T | Transcriptional regulatory circuit | Combinatorial gene control, enhancer logic, developmental patterning |
| P-Reg-G | Autoregulatory feedback | Homeostasis. Self-referential bootstrap. Negative autoregulation, bistability. |
| P-Mem-Reg | Signal transduction | Environment sensing. Cascade from receptor to gene expression change. |
| R-P-Mem | Co-translational insertion | Secretory pathway. SRP. Eukaryote-required. |
One borderline quad:
| Quad | Name | Note |
|---|---|---|
| G-T-R-P | Full expression pipeline | Decomposes into GTR + R-P overlapping at R; borderline per R3. |
No firm quads. Biology's emergent properties decompose into triangles. The entity system has one firm quad (E-I-T-X, self-sustainability); biology bundles identity into G (sequence IS identity) and execution into R (the ribosome combines code-reading and synthesis), so GTR carries what takes EITX in the entity system. Biology is more compressed at the composition level.
Anchor-pair clustering: All five triangles contain at least one anchor primitive (G appears in 3, Reg in 3, P in 4). 100% anchor coverage.
9.2 Core triad
{G, T, R} — the central dogma. Information storage + transfer + execution.
Cross-domain comparison:
| Domain | Core triad | What it does |
|---|---|---|
| Biology | {G, T, R} | Information flow (genome → transcript → protein) |
| Chemistry | {El, Bd, Rx} | Material transformation (atoms → molecules → reactions) |
| Entity system | {E, I, T} | Self-description (entity → identity → tree) |
| Cognition | {Rp, Ct, As} | Semantic relation (representation → context → association) |
All four core triads are 3-primitive irreducible cores naming the substrate's defining flow. Biology's flow is information; chemistry's is matter; entity system's is identity; cognition's is meaning.
Steps 7–8 — Lattice and Hasse Walks
7.1 Coherent sub-lattice at primitive-presence resolution
Valid subsets given the dependency DAG:
| # | Subset | Biological identity |
|---|---|---|
| 1 | {} | No biology (chemistry only) |
| 2 | {G} | Dormant DNA, viroid fragment |
| 3 | {G, T} | RNA world remnant (transcription/ribozyme without translation) |
| 4 | {G, T, R} | Minimal gene expression (central dogma without regulation or boundary) |
| 5 | {G, T, R, P} | Naked replicator |
| 6 | {G, T, R, P, Reg} | Regulated replicator (chromosome with logic, no boundary) |
| 7 | {G, T, R, P, Mem} | Minimal cell (expression + compartment, no regulation) |
| 8 | {G, T, R, P, Reg, Mem} | Full cell |
8 / 64 = 12.5% coherent. The tightest filter of any domain analyzed.
Comparison: biology 12.5%, entity system 14%, chemistry 20%, cognition substrate 14%. Substrate domains cluster at 12-20%. Biology's especially tight filter reflects the sequential dependency chain — each primitive requires nearly all the previous ones.
7.2 Host-dependent / parasitic manifestations
Viruses occupy positions in the lattice that violate the dependency structure — they have P (protein products) without their own R (ribosome). This works because viruses depend on a host whose complete primitive set provides the missing capabilities.
This is structurally significant: the coherent sub-lattice describes self-sufficient systems. Host-dependent systems can occupy otherwise-incoherent positions by relying on a host that provides the missing primitives. The dependency filter tells you what's needed for self-sufficiency; violations indicate host-dependency.
Terminology note: "Host-dependent" is a structural descriptor — it means the manifestation occupies a dependency-violating position. It says nothing about the relationship quality. The relationship can be parasitic (exploitative — biological viruses, malware), symbiotic (mutually beneficial — mitochondria, software you install), or commensal (neutral). Biological viruses are host-dependent AND parasitic. Most software is host-dependent AND symbiotic. The structural dependency and the relationship quality are orthogonal.
This is methodologically important: the methodology's coupling vocabulary (Cp, per Layer 4) covers the relationship dimension. Position-in-lattice is structural; coupling-character is relational. Both are needed to describe a host-dependent manifestation.
8.1 Hasse walks (primitive-presence)
Two complete monotone build-up paths from {} to {G, T, R, P, Reg, Mem}:
- Path α (central dogma → regulation → boundary): G → GT → GTR → GTRP → GTRP-Reg → Full. Information → expression → regulation → compartmentalization.
- Path β (central dogma → boundary → regulation): G → GT → GTR → GTRP → GTRP-Mem → Full. Information → expression → compartmentalization → regulation.
Only 2 paths. The fewest of any analyzed domain (entity system: 3; chemistry: 4; UI: 4). Biology has the most constrained build-up — direct consequence of the nearly-linear dependency chain.
8.2 Build-up narrative
| Position | Biological identity | Corresponds to |
|---|---|---|
| {} | No biology | Pre-biotic chemistry |
| {G} | Dormant information | Viroid, spore, desiccated DNA |
| {G, T} | Information readout | RNA world (ribozymes as both template and enzyme) |
| {G, T, R} | Information → function | Proto-cell: central dogma active, unregulated, unbounded |
| {G, T, R, P} | Functional system | Naked replicator |
| {G, T, R, P, Reg} | Regulated system | Chromosome with logic |
| {G, T, R, P, Mem} | Bounded system | Minimal cell, no regulation |
| {G, T, R, P, Reg, Mem} | Full cell | LUCA and all descendants |
The build-up IS the abiogenesis narrative. Not because the methodology forced it but because both trace the same bootstrapping problem — how to build a self-sustaining information substrate from chemistry. The order: information first (G), then readout (T), then execution (R), then function (P), then control (Reg) and boundary (Mem). The full reverse-walk reconstruction of abiogenesis is in abiogenesis_analysis_v1/.
8.3 Sub-resolution: the abiogenesis trajectory
The R0 → R2 transition (the central abiogenesis bottleneck) decomposes at sub-resolution into a multi-step walk through a product corridor:
R0 → R0.5 → R1 → R1.7 → R1.9 → R2
This is an autocatalytic spiral: R (ribosome accuracy) and P (protein quality) co-advance through positive feedback. Each cycle improves both. Below the critical threshold, advancement is linear (trickle); above it, exponential (flood). Terminates at R2 with the crystallization of the genetic code — frozen since LUCA, irreversible, enabling, universal.
This is the canonical instance of the autocatalytic spiral in the methodology. See abiogenesis_analysis_v1/ for the full sub-resolution analysis with conditional partial-level dependencies, internal phase transitions, and rate variation.
Step 10 — Emergent Property Map
| Property | Required composition | Required regime | Prediction |
|---|---|---|---|
| Heredity | G internal (replication) | G ≥ 2, high-fidelity copying | Information persists across generations |
| Gene expression | G-T-R triangle | G ≥ 2, T ≥ 1, R ≥ 2 | Information → function conversion |
| Metabolic specificity | R-P, G-R | R ≥ 2, P ≥ 2 | Enzymes with specific catalytic activity |
| Homeostasis | P-Reg-G | P ≥ 2, Reg ≥ 1, G ≥ 2 | Self-regulating steady states; negative autoregulation |
| Conditional behavior | G-Reg, T-Reg | G ≥ 3, Reg ≥ 2 | Different environments → different gene expression |
| Cell identity | G-Reg + Mem | Reg ≥ 4, Mem ≥ 3 | Same genome, different stable expression states |
| Signal transduction | P-Mem-Reg | P ≥ 3, Mem ≥ 2, Reg ≥ 2 | External signal → gene expression change |
| Multicellularity | Mem4 + Reg4 + P4 | Mem ≥ 4, Reg ≥ 4, P ≥ 4 | Coordinated multi-cell behavior |
| Adaptive immunity | G + Reg + P + Mem | G-Full, Reg ≥ 3, P ≥ 4, Mem ≥ 3 | Somatic recombination, clonal selection |
| Neural computation | P-Full + Mem-Full + Reg ≥ 4 | Full P, Full Mem, Reg ≥ 4 | Ion channels, synapses, activity-dependent expression |
| Evolution | Full substrate, imperfect-fidelity regime | G ≥ 2 (imperfect), R ≥ 2, Reg ≥ 1, P ≥ 2 | Self-modification of the substrate; requires the full system |
Each is a testable prediction: a system at or above the regime should exhibit the property; below should not. Viruses (below thresholds) exhibit none of the self-sustaining properties. E. coli (mid-range) exhibits heredity through homeostasis and conditional behavior, but not cell identity or multicellularity. Mammals (near-full) exhibit all.
Steps 11–12 — Structural Patterns and Literature Alignment
11.1 Cross-domain patterns
Anchor-pair clustering. G is the single hub with highest pair-connectivity (4 of 7 heavy pairs). Replicates across domains (entity system E+I hub, chemistry Rx hub, cognition Rp hub).
Core triad emergence. {G, T, R} is the irreducible core. Replicates: chemistry {El, Bd, Rx}, entity system {E, I, T}, cognition {Rp, Ct, As}. Every analyzed substrate has a core triad.
Filter stringency as substrate signature. Biology at 12.5% is the tightest of any domain. Substrate domains cluster at 12–20% (biology, entity system, chemistry, cognition). Surface domains (UI 52%, capability 56%) are looser. R11's prediction holds.
Triangle count stability. 5 firm triangles. Same as entity system (5), chemistry (5 + 1 borderline quad), cognition (5). The count ≈ 5 for 6-primitive substrate domains is consistent across analyses.
Phase transition concentration. Biology's R0 → R2 (abiogenesis) and Mem2 → Mem3 (eukaryogenesis) parallel the entity system's X0 → X2 (open dispatch) and E2 → Full E. Phase transitions cluster at "where the system becomes qualitatively new." Replicates across substrate analyses.
Self-referential bootstrap. Biology (proteins regulate their own genes) parallels entity system (types describe themselves), PL (metacircular evaluator), DB (schema-as-data). Self-reference appears in every substrate domain analyzed — likely a structural necessity for self-sustaining information substrates.
11.2 SSA mapping
Biology is the canonical instance of the Situated Substrate Architecture:
| SSA primitive | Biology mapping |
|---|---|
| Encoding (En) | Genome (G) |
| Evaluator (Vr) | Ribosome (R), at Kd4 |
| Mechanism (Mc) | ~12 developmental mechanisms (analyzed in biology-to-organism bridge) |
| Surface (Sf) | Organism architecture (~9 primitives, separate analysis) |
| Context (Cx) | Environment (~6 primitives, separate analysis) |
| Community (Cm) | Ecosystem (~9 primitives, separate analysis) |
| Selection (Se) | Natural selection (within ecosystem analysis) |
Biology is the SSA's empirical foundation — every SSA component has a biology instance, all confirmed across 4 billion years of selection. The two other confirmed SSA arrangements (entity system and cognition) inherit their structure from this template, though each implements it with different evaluator determinism, different Mc count, and different selection mechanism.
12.1 Literature alignment
- Central dogma (Crick 1958, 1970). The build-up sequence G → GT → GTR independently reconstructs the central dogma as a structural necessity. The directional flow corresponds to the dependency chain. The methodology explains why the central dogma has the direction it does.
- RNA world (Gilbert 1986). The lattice position {G, T} corresponds exactly — RNA as both information store and catalytic agent (ribozyme). Methodology predicts this is a coherent intermediate; RNA world hypothesis asserts it was a historical one.
- Operon model (Jacob-Monod 1961). Maps to G-Reg-T triangle. Operons are the minimal structure with this triangle active. Methodology's pair-bundle analysis identifies G-Reg and T-Reg as heavy pairs — exactly where operon biology concentrates.
- Network motifs (Alon 2007). Map to P-Reg-G triangle. Negative autoregulation, feed-forward loops, toggle switches. Methodology independently identifies this triangle as load-bearing; systems biology confirms it as cellular computation's substrate.
- Major evolutionary transitions (Maynard Smith & Szathmáry 1995). Hasse walk maps to recognized transitions: RNA world → translation (GTR activation), unregulated → regulated (Reg), unbounded → bounded (Mem), prokaryote → eukaryote (Mem2 → Mem3), unicellular → multicellular (Reg3 → Reg4, Mem3 → Mem4). Sequence and discontinuities both agree.
12.2 What the methodology adds beyond existing biology
The methodology does not discover new biology. It provides:
- Structural coordinate system. "E. coli is (G3, T2, R3, P3, Reg2, Mem2)" is more informative than a phylogenetic label.
- Phase transition identification. The hard transitions (R0 → R2, Mem2 → Mem3) are identified structurally rather than by historical contingency.
- Pair-load explanation. Why G-Reg is biology's biggest research area: it's the heaviest pair connecting primary anchor (G) to secondary co-anchor (Reg).
- Cross-domain comparison framework. Biology can be formally compared to entity system, chemistry, cognition — structural, not metaphorical.
Manifestation Landscape
| System | G | T | R | P | Reg | Mem | Notes |
|---|---|---|---|---|---|---|---|
| Viroid | 1 | 0 | 0 | 0 | 0 | 0 | Naked RNA |
| TMV (RNA virus) | 2 | 0* | 0* | 2 | 1 | 0 | *Host-dependent |
| Mycoplasma | 3 | 2 | 2 | 2 | 1 | 1 | Smallest free-living |
| E. coli | 3 | 2 | 3 | 3 | 2 | 2 | Model prokaryote |
| Yeast | 4 | 3 | 3+ | 4 | 3 | 3 | Simple eukaryote |
| C. elegans | 4 | F | F | 4 | 4 | 4 | Simple multicellular |
| Drosophila | 4 | F | F | 4 | 4 | 4 | Complex multicellular |
| Arabidopsis | 4 | F | F | 4 | 4 | 4 | Model plant |
| Human | F | F | F | F | F | F | Near-full elaboration |
(F = Full; 0* = host-dependent, dependency-violating)
Attractor positions
- Parasitic / host-dependent agent (G2, T0*, R0*, P1-2, Reg0-1, Mem0). Viruses, viroids. Dependency-violating; stabilized by host coupling.
- Minimal free-living cell (G3, T2, R2-3, P2-3, Reg1-2, Mem1-2). Mycoplasma, simple bacteria.
- Complex unicellular eukaryote (G4, T3, R3-Full, P4, Reg3, Mem3). Yeast, protists.
- Complex multicellular (G4-Full, Full T, Full R, P4-Full, Reg4-Full, Mem4-Full). Animals, plants, fungi.
(Full everything is not an attractor — it is the current frontier.)
Walls and fences (analyst judgment, per methodology §6.3 caveat)
| Transition | Tentative character | Evolutionary frequency | Reasoning |
|---|---|---|---|
| RNA world → translation (R0 → R2) | Wall | Once in 4 Gya | Requires the genetic code; an enormously complex molecular machine; happened once on Earth |
| Prokaryote → eukaryote (Mem2 → Mem3) | Wall | Once | Requires endosymbiosis; a singular event |
| Unicellular → multicellular (Reg3 → Reg4, Mem3 → Mem4) | Fence | 25+ times independently | Additive; existing machinery extended |
| Asexual → sexual (G internal) | Fence | Multiple times | G-internal recombination enhancement |
| Aquatic → terrestrial | Fence | Multiple times in plants, animals, fungi | Environmental adaptation through Mem and Reg elaboration |
Per methodology §6.3, these wall/fence assignments are analyst judgments, not derived from the structural model alone. They use external evidence (evolutionary frequency) to assign character. The structural finding is: walls correspond to once-in-history events; fences correspond to independent re-emergences. The model predicts the existence of an obstruction; the wall-vs-fence character requires the deeper substructure analysis flagged in methodology §10 question 10.
Summary
Domain: Biology substrate (En primitive of the SSA).
Primitive set: {G, T, R, P, Reg, Mem}.
Filter stringency: 8/64 = 12.5%. Tightest of any domain analyzed.
Pair distribution: 7 heavy / 4 medium / 4 light / 0 negligible. 47% heavy.
Core triad: {G, T, R} — central dogma. Information flow.
Primary anchor: G (Genome) — biology is information-centric.
Phase transitions: G1 → G2 (gene encoding); T1 → T2 (regulated initiation); R0 → R2 (abiogenesis bottleneck — sub-resolution autocatalytic spiral terminating at code crystallization); P1 → P2 (folding); Reg2 → Reg3 (combinatorial); Reg3 → Reg4 (epigenetic); Mem2 → Mem3 (eukaryogenesis).
Load-bearing compositions: 5 firm triangles, 1 borderline quad. No firm quads.
Hasse paths: 2 — fewest of any analyzed domain.
Self-referential bootstrap: Reg → P → R → T → G → Reg cycle. Biology's instance of the structural pattern shared with all substrate domains.
Position in topology: Substrate of biology arrangement. Bridges down to chemistry; bridges up to organism architecture. Provides En in the SSA.
Cross-domain mapping: Structural homomorphism with entity system. Same six functional roles, same dependency order, same core triad structure, same filter stringency (~13%). Differences: biology bundles E+I into G (sequence IS identity); biology's evaluator (R) is genome-encoded; biology's regulation is analog while entity system's is discrete; biology's boundary is physical while entity system's is logical; evolution operates on biology, design selects entity system primitives.
Open work for biology arrangement:
- Chemistry-to-biology bridge (
analysis-chemistry-to-biology-bridge.md) — captures the En → Vr separation event - Biology-to-organism bridge (deferred) — captures Mc (~12 developmental mechanisms)
- Organism architecture (deferred) — Sf surface
- Environment context root and ecosystem — Cx and Cm
- Selection mechanism (within ecosystem) — Se
Referenced by the model
Cited as a source by 10 model records (browse the model census):
- biology-substrate —
domainbiology/sc1 - ecoli —
manifestationbiology/sc3/ecoli - halobacterium —
manifestationbiology/sc3/halobacterium - human —
manifestationbiology/sc3/human - methanococcus —
manifestationbiology/sc3/methanococcus - tetrahymena —
manifestationbiology/sc3/tetrahymena - yeast —
manifestationbiology/sc3/yeast - phylogenesis-stem —
trajectorybiology/sc3/phylogenesis-stem-to-human - phylogenesis-stem-calibrated —
ratebiology/sc2 - phylogenesis-prokaryote-to-eukaryote —
population_contextbiology/sc2/ecoli