The Architecture of the
Epistemic Zero
π Read the Preprint
"The Einstein Test and Beyond:
The Architecture of the Semantic Zero

The AI industry is currently locked in a multi-billion-dollar race to hoard compute power, treating intelligence as a discrete asset to be trapped inside massive, centralized data centers. In our newly published preprint, "The Einstein Test and Beyond: The Architecture of the Semantic Zero", we argue that this scale-maximalist paradigm is mathematically flawedβbuilding ever-larger "Roman calculators" that are thermodynamically forced to hallucinate because their underlying notation structurally lacks a semantic zero. To dissolve this performance asymptote, we introduce the Semiotic Web: a federated architecture that utilizes cryptographically anchored Canonical Concept Identities (CCI) and Contextual Tokum Instances (CTI) to ground machine reasoning in verifiable reality, permanently shifting the future of AI from hoarded computation to a dynamic, transparent flow of verified gap-closure.
π Explore the Slides
π§ Listen to the Deep Dive
πΊ Watch the Video
The Piaget Test: What AI Got Half-Right β and Why That Half Is Not Enough
| Dimension | PART 1 β β...how much we know how to DOβ (Execution under certainty) | PART 2 β β...how we BEHAVE when we DON'T KNOWβ (Behaviour under genuine uncertainty) | What Current AI Does β and What Is Missing |
|---|---|---|---|
| Core epistemic mode | Retrieval & interpolation. The system matches a prompt to its training distribution and returns the statistically most likely completion. | Genuine gap-sensing. The system detects that no grounded answer exists, registers the absence as a first-class fact, and acts accordingly. | βCurrent AI has only Part 1. Softmax forces every output into a probability distribution that sums to 1 β there is no coordinate for honest absence. |
| Architectural prerequisite | A vast vocabulary + attention weights over trained data. βAttention is all you needβ (Vaswani et al., 2017) β the Part 1 machine. | A Semantic Zero: a stable, addressable coordinate where verified absence can live before inference begins. (CCI β Canonical Concept Identity) | βWithout a Semantic Zero, the architecture is mechanically forced to guess. Hallucination is not a bug β it is the spec. |
| Processing sequence (correct order) | STEP 2 β Reason on grounded elements. Allocate attention to what is known, verified, and addressable. | STEP 1 β Define the gap FIRST. Map what is not known; register absence as a structural address before reasoning begins. | βAI inverts this sequence: it reasons first (Step 2) without ever completing Step 1. The Notation Inversion restores the correct order. |
| Response to unknown input | Confident output regardless of grounding. Softmax redistributes probability β uncertainty is only a ranking problem among guesses. | Structurally Bounded Refusal. The system returns NIL, signals the gap via CCI, then either seeks evidence (System 2) or switches to explicit stochastic mode (System 1 / CMP). | βAI cannot refuse structurally. NIL-Token Injection Ablation (Β§3.2) is the proposed falsification test: isolating forced guessing as root cause. |
| What Chollet's ARC-AGI measures | Part 1 capacity β rate of skill acquisition from prior training data. State-of-the-art LLMs: ~95%+ on standard benchmarks. | Part 2 capacity β fluid intelligence on genuinely novel tasks, withheld from training distribution. State-of-the-art LLMs: 0.26% on ARC-AGI-3. | βThe 186Γ collapse (100% β 0.26%) is the exact measurable cost of operating without a Semantic Zero when confronted with gap-sensitive tasks. |
| Knowledge model | Closed, sealed manifold. All knowledge is encoded at training time; inference redistributes statistical mass within a fixed surface. | Open, verifiable graph. CCI provides a platonic address space; CTI (Contextual Tokum Instance) anchors real-world observations as cryptographic proofs. | βA sealed manifold is GΓΆdel-incomplete by construction (Β§8.3). Only an externally anchored CTI punctures the topology β changing the manifold class from closed to open. |
| Definition of intelligence | Ptolemaic (current AI assumption): Intelligence = a discrete stock accumulated inside a single, isolated agent (more compute β more intelligence). | Copernican (Semiotic Web): Intelligence = a FLOW that reduces systemic stress through gap-closure β a property of distributed, holonic federation. | The Notation Inversion abandons the hoarding premise. Intelligence is not a stock inside a machine; it is a flow across verified, bounded agents (Semantic Light Cone of Care). |
| Analogy | The greatest Roman calculator, given every surviving text, unlimited clay tablets, and maximum motivation. β Cannot compute a derivative. Not a cleverness deficit β a notation deficit. | Hindu-Arabic positional notation: the zero is a structural address where emptiness is a manipulable mathematical object. β Enables calculus, algebra, and all that follows. | Adding more parameters to a Stochastic Guessing Engine will never yield a Semantic Zero β just as adding more tally marks never yields the concept of zero. Only a Notation Inversion achieves this. |
| Conclusion: are both parts necessary? | βYES β Part 1 is essential. Current LLMs are the most powerful formal accelerators in history: sublime cartographers mapping known semantic territory at superhuman speed. | βYES β Part 2 must come FIRST. Without the gap-defining step, Part 1 reasoning is structurally ungrounded. The machine cannot be the compass β only the map. | β The one innovation needed (Nadella): Not a new scaling law β a Notation Inversion: introducing a Semantic Zero and a cryptographically verified Observer's Mark into the epistemic substrate. |
β ELEVATOR SPEECH
Every system we call intelligent today β from the smallest chatbot to the largest frontier model β excels at the first half of Piaget's test: it knows an extraordinary amount, and it executes with confidence. What it structurally cannot do is the second half: behave honestly when it doesn't know. That is not a data problem or a scale problem β it is a notation problem. Softmax, the terminal function of every modern transformer, must always sum to 1; there is no coordinate in the architecture where verified absence can live. The result is a system mechanically compelled to guess, dress the guess in fluent language, and call it knowledge. The Semantic Zero β a stable, cryptographically anchored address for βI don't knowβ β is the structural missing term. Once it exists, the correct processing sequence is restored: first define the gap (Part 2), then reason on grounded elements (Part 1). Intelligence stops being a stock hoarded inside a sealed machine and becomes what it always was in nature: a flow β the rate at which a network of verified, bounded agents closes the distance between what is known and what is not.
Source: Blaettler & McCarey, βThe Einstein Test and Beyond: The Architecture of the Semantic Zeroβ, https://doi.org/10.5281/zenodo.21108888, 2026 | tokum.ai/Semantic-Zero
The Piaget Test: What AI Got Half-Right β and Why That Half Is Not Enough
| Dimension | PART 1 β β...how much we know how to DOβ (Execution under certainty) | PART 2 β β...how we BEHAVE when we DON'T KNOWβ (Behaviour under genuine uncertainty) | What Current AI Does β and What Is Missing |
|---|---|---|---|
| Core epistemic mode | Retrieval & interpolation. The system matches a prompt to its training distribution and returns the statistically most likely completion. | Genuine gap-sensing. The system detects that no grounded answer exists, registers the absence as a first-class fact, and acts accordingly. | βCurrent AI has only Part 1. Softmax forces every output into a probability distribution that sums to 1 β there is no coordinate for honest absence. |
| Architectural prerequisite | A vast vocabulary + attention weights over trained data. βAttention is all you needβ (Vaswani et al., 2017) β the Part 1 machine. | A Semantic Zero: a stable, addressable coordinate where verified absence can live before inference begins. (CCI β Canonical Concept Identity) | βWithout a Semantic Zero, the architecture is mechanically forced to guess. Hallucination is not a bug β it is the spec. |
| Processing sequence (correct order) | STEP 2 β Reason on grounded elements. Allocate attention to what is known, verified, and addressable. | STEP 1 β Define the gap FIRST. Map what is not known; register absence as a structural address before reasoning begins. | βAI inverts this sequence: it reasons first (Step 2) without ever completing Step 1. The Notation Inversion restores the correct order. |
| Response to unknown input | Confident output regardless of grounding. Softmax redistributes probability β uncertainty is only a ranking problem among guesses. | Structurally Bounded Refusal. The system returns NIL, signals the gap via CCI, then either seeks evidence (System 2) or switches to explicit stochastic mode (System 1 / CMP). | βAI cannot refuse structurally. NIL-Token Injection Ablation (Β§3.2) is the proposed falsification test: isolating forced guessing as root cause. |
| What Chollet's ARC-AGI measures | Part 1 capacity β rate of skill acquisition from prior training data. State-of-the-art LLMs: ~95%+ on standard benchmarks. | Part 2 capacity β fluid intelligence on genuinely novel tasks, withheld from training distribution. State-of-the-art LLMs: 0.26% on ARC-AGI-3. | βThe 186Γ collapse (100% β 0.26%) is the exact measurable cost of operating without a Semantic Zero when confronted with gap-sensitive tasks. |
| Knowledge model | Closed, sealed manifold. All knowledge is encoded at training time; inference redistributes statistical mass within a fixed surface. | Open, verifiable graph. CCI provides a platonic address space; CTI (Contextual Tokum Instance) anchors real-world observations as cryptographic proofs. | βA sealed manifold is GΓΆdel-incomplete by construction (Β§8.3). Only an externally anchored CTI punctures the topology β changing the manifold class from closed to open. |
| Definition of intelligence | Ptolemaic (current AI assumption): Intelligence = a discrete stock accumulated inside a single, isolated agent (more compute β more intelligence). | Copernican (Semiotic Web): Intelligence = a FLOW that reduces systemic stress through gap-closure β a property of distributed, holonic federation. | The Notation Inversion abandons the hoarding premise. Intelligence is not a stock inside a machine; it is a flow across verified, bounded agents (Semantic Light Cone of Care). |
| Analogy | The greatest Roman calculator, given every surviving text, unlimited clay tablets, and maximum motivation. β Cannot compute a derivative. Not a cleverness deficit β a notation deficit. | Hindu-Arabic positional notation: the zero is a structural address where emptiness is a manipulable mathematical object. β Enables calculus, algebra, and all that follows. | Adding more parameters to a Stochastic Guessing Engine will never yield a Semantic Zero β just as adding more tally marks never yields the concept of zero. Only a Notation Inversion achieves this. |
| Conclusion: are both parts necessary? | βYES β Part 1 is essential. Current LLMs are the most powerful formal accelerators in history: sublime cartographers mapping known semantic territory at superhuman speed. | βYES β Part 2 must come FIRST. Without the gap-defining step, Part 1 reasoning is structurally ungrounded. The machine cannot be the compass β only the map. | β The one innovation needed (Nadella): Not a new scaling law β a Notation Inversion: introducing a Semantic Zero and a cryptographically verified Observer's Mark into the epistemic substrate. |
β ELEVATOR SPEECH
Every system we call intelligent today β from the smallest chatbot to the largest frontier model β excels at the first half of Piaget's test: it knows an extraordinary amount, and it executes with confidence. What it structurally cannot do is the second half: behave honestly when it doesn't know. That is not a data problem or a scale problem β it is a notation problem. Softmax, the terminal function of every modern transformer, must always sum to 1; there is no coordinate in the architecture where verified absence can live. The result is a system mechanically compelled to guess, dress the guess in fluent language, and call it knowledge. The Semantic Zero β a stable, cryptographically anchored address for βI don't knowβ β is the structural missing term. Once it exists, the correct processing sequence is restored: first define the gap (Part 2), then reason on grounded elements (Part 1). Intelligence stops being a stock hoarded inside a sealed machine and becomes what it always was in nature: a flow β the rate at which a network of verified, bounded agents closes the distance between what is known and what is not.
Source: Blaettler & McCarey, βThe Einstein Test and Beyond: The Architecture of the Semantic Zeroβ, https://doi.org/10.5281/zenodo.21108888, 2026 | tokum.ai/Semantic-Zero