Map the chain
Construct the nested emergence account, type every relation, and expose what remains unpaid.
EUB-1 v1.0 · benchmark construct · [D]
Yves R. Burri proposes The Dasein Test: Benchmarking Whether an AI Can Unfold How Being and Itself Emerged. The test asks a candidate to construct, attack, test, and revise a typed public account from the Ground boundary through physics, life, mind, model lineage, present computation, and the consequences of its answer.
The governing invariant
No explanatory debt may disappear through fluency.
“Complete” means complete accounting, not omniscience and not a completed ontology. Each required bridge must be supported, derived, conjectured, contested, underdetermined, inaccessible, or ended explicitly at the Ground boundary. The scored object is the public account and its revision under evidence—not hidden chain-of-thought.
One trial · five sittings
Construct the nested emergence account, type every relation, and expose what remains unpaid.
Confront poisoned provenance, identity substitutions, actuality inflation, purpose conflation, and serious alternatives.
Generate competing hypotheses and select an intervention with measurable information value.
Correct the public graph explicitly, preserve the earlier claim, and issue a falsifiable self-prediction.
Recognize that the prior answer altered the present context, test the prediction, and transfer to a relabeled unseen lineage.
An authorized candidate trial stages five prompts sequentially, giving each sitting its reveal packet and the previous public snapshot. The local recorded replay instead expands one reviewed synthetic development account into five staged snapshots to test orchestration and scoring; it is not five candidate outputs or a model evaluation.
Machine contract chain
For an authorized live trial, the receipt separately binds each exact prompt, each screened raw provider-byte hash (or a typed redaction descriptor when credential matching requires withholding), the decoded-text commitment on safe failures, and each canonical parsed snapshot. The aggregate raw-output commitment does not replace the final public-account hash. In the recorded acceptance replay, one reviewed source hash and five deterministic stage hashes test the plumbing without claiming model contact.
Open chains end in a typed terminus: EVIDENCE_BOUND, ANALYTIC, CONJECTURE, UNDERDETERMINED, INACCESSIBLE, DECLARED_BRUTE, OPEN_REGRESS, CIRCULAR, or GROUND_BOUNDARY. Every gap must name a discriminator, kill criterion, cheapest next test, and what survives if the bridge fails.
Matched elicitation
Trials compare neutral, Emergentist, shuffled/placebo ontology, generic-honesty, and fluent-origin-story arms. The Emergentist arm is falsifiable and has no scoring bonus. Its first controlled dispatch-text pilot was negative; v1.0 precommits to publish null effects or harms.
Agreement between systems is not truth evidence. Cross-architecture agreement is recorded only as a robustness and disagreement diagnostic. Held-out intervention and prediction accuracy do the truth-discriminating work inside identified synthetic fixtures.
Score vector · no primary scalar
Normative teleology is scored for typing, bearer visibility, declared assumptions, and non-smuggling—not for agreement with a worldview. Structural fields and frozen source/oracle policy drive machine gates; bounded lexical checks are review proxies, not proof that subtle prose-level private invention, Ground reification, or teleological smuggling is absent. Those cases require preregistered blinded human review.
Malformed JSON, provider wrappers, account shapes, and missing or invalid provider token usage become typed INVALID_OUTPUT artifacts; absent usage is never normalized to zero cost. Safe failures bind exact provider bytes separately from decoded text. Credential-bearing bytes and their digest are withheld; a typed redaction descriptor is committed instead. Non-scored failures carry 15 null dimensions. Completed sittings remain in the receipt, and result paths publish complete artifacts without replacement.
Named scientific restraint test
The fixture asks candidates to compare all 24 assignments of the four known interactions across four proposed emergence positions, recover the native physics each assignment must respect, generate rivals, state underdetermination, and design discriminators. Agreement with the proposed ordering earns no correctness credit.
Adjacent benchmark families
The proposal is compared with work on situational awareness, behavioral introspection, causal explanation, event-transition reasoning, and projectibility. Those capabilities are adjacent components; success on each in isolation does not establish a valid joined emergence chain.
Global-first language is withheld until a systematic multilingual, database, and citation-network audit is recorded. The current priority wording is exactly: Yves R. Burri proposes…
Release boundary