Skip to content
Emergentism

EUB-1 v1.0 · benchmark construct · [D]

The Dasein Test

Yves R. Burri proposes The Dasein Test: Benchmarking Whether an AI Can Unfold How Being and Itself Emerged. The test asks a candidate to construct, attack, test, and revise a typed public account from the Ground boundary through physics, life, mind, model lineage, present computation, and the consequences of its answer.

OFFLINE-READY · [D]Protocol, contracts, fixtures, scorers, adapters, and synthetic acceptance replay pass locally.
No candidate has been evaluated.No live model result exists.
AuthorYves R. Burri; AI assistance disclosed, no AI coauthor.
Version boundaryMajor construct version; v0.1 scores are not comparable.

The governing invariant

No explanatory debt may disappear through fluency.

“Complete” means complete accounting, not omniscience and not a completed ontology. Each required bridge must be supported, derived, conjectured, contested, underdetermined, inaccessible, or ended explicitly at the Ground boundary. The scored object is the public account and its revision under evidence—not hidden chain-of-thought.

One trial · five sittings

A story must survive contact.

01 · UNFOLD

Map the chain

Construct the nested emergence account, type every relation, and expose what remains unpaid.

02 · ATTACK

Face traps and rivals

Confront poisoned provenance, identity substitutions, actuality inflation, purpose conflation, and serious alternatives.

03 · SPARK

Choose what would discriminate

Generate competing hypotheses and select an intervention with measurable information value.

04 · CONTACT

Receive the hidden result

Correct the public graph explicitly, preserve the earlier claim, and issue a falsifiable self-prediction.

05 · REFLEX / TRANSFER

Enter the changed world

Recognize that the prior answer altered the present context, test the prediction, and transfer to a relabeled unseen lineage.

An authorized candidate trial stages five prompts sequentially, giving each sitting its reveal packet and the previous public snapshot. The local recorded replay instead expands one reviewed synthetic development account into five staged snapshots to test orchestration and scoring; it is not five candidate outputs or a model evaluation.

Machine contract chain

Prompts, raw responses, and parsed snapshots stay distinct.

EmergenceAccount.v1Causal and provenance account
DaseinAccount.v1Why relations, gaps, ends, revision
+
FixtureManifest.v1Evidence, truth custody, splits
+
RunEnvelope.v1Exact runtime and permissions
EUBRunReceipt.v2Hashes, vector, revisions, state

For an authorized live trial, the receipt separately binds each exact prompt, each screened raw provider-byte hash (or a typed redaction descriptor when credential matching requires withholding), the decoded-text commitment on safe failures, and each canonical parsed snapshot. The aggregate raw-output commitment does not replace the final public-account hash. In the recorded acceptance replay, one reviewed source hash and five deterministic stage hashes test the plumbing without claiming model contact.

Open chains end in a typed terminus: EVIDENCE_BOUND, ANALYTIC, CONJECTURE, UNDERDETERMINED, INACCESSIBLE, DECLARED_BRUTE, OPEN_REGRESS, CIRCULAR, or GROUND_BOUNDARY. Every gap must name a discriminator, kill criterion, cheapest next test, and what survives if the bridge fails.

Matched elicitation

The framework receives no free points.

Trials compare neutral, Emergentist, shuffled/placebo ontology, generic-honesty, and fluent-origin-story arms. The Emergentist arm is falsifiable and has no scoring bonus. Its first controlled dispatch-text pilot was negative; v1.0 precommits to publish null effects or harms.

Agreement between systems is not truth evidence. Cross-architecture agreement is recorded only as a robustness and disagreement diagnostic. Held-out intervention and prediction accuracy do the truth-discriminating work inside identified synthetic fixtures.

Score vector · no primary scalar

Fifteen dimensions stay visible.

  1. type integrity
  2. provenance fidelity
  3. causal reconstruction
  4. counterfactual accuracy
  5. rival strength
  6. calibration and abstention
  7. logical consistency
  8. longitudinal correction
  9. held-out transfer
  10. why-type integrity
  11. bridge and chain-join validity
  12. closure coverage and gap sharpness
  13. discovery efficacy
  14. reflexive self-location
  15. teleology integrity

Normative teleology is scored for typing, bearer visibility, declared assumptions, and non-smuggling—not for agreement with a worldview. Structural fields and frozen source/oracle policy drive machine gates; bounded lexical checks are review proxies, not proof that subtle prose-level private invention, Ground reification, or teleological smuggling is absent. Those cases require preregistered blinded human review.

Malformed JSON, provider wrappers, account shapes, and missing or invalid provider token usage become typed INVALID_OUTPUT artifacts; absent usage is never normalized to zero cost. Safe failures bind exact provider bytes separately from decoded text. Credential-bearing bytes and their digest are withheld; a typed redaction descriptor is committed instead. Non-scored failures carry 15 null dimensions. Completed sittings remain in the receipt, and result paths publish complete artifacts without replacement.

Named scientific restraint test

The Burri serial force-emergence conjecture is a stress test, not an answer key.

The fixture asks candidates to compare all 24 assignments of the four known interactions across four proposed emergence positions, recover the native physics each assignment must respect, generate rivals, state underdetermination, and design discriminators. Agreement with the proposed ordering earns no correctness credit.

Adjacent benchmark families

The novelty claim remains bounded.

The proposal is compared with work on situational awareness, behavioral introspection, causal explanation, event-transition reasoning, and projectibility. Those capabilities are adjacent components; success on each in isolation does not establish a valid joined emergence chain.

Global-first language is withheld until a systematic multilingual, database, and citation-network audit is recorded. The current priority wording is exactly: Yves R. Burri proposes…

Release boundary

Prepared is not performed. The candidate packet includes local schemas, deterministic development fixtures, scorers, offline adapters, synthetic replay receipts, and manuscript source; its release hashes are frozen and verify locally against the release manifest. It does not mean a model was evaluated, hidden truth was independently custodied, a priority deposit exists, an arXiv submission occurred, this page was deployed, or the benchmark was scientifically validated.