ARKHĒPUBLIC RELEASE / 16 AUG 2026

METHOD / NEGATIVE CONTROLS / FALSIFICATION

What would change the verdict?

A replicated Pareto improvement over the best simple baseline—with component ablations that identify a distinctive cause, not merely a longer prompt or a favorable judge.

01 / COMPARISON DESIGN

Four prompts. Same task packet. Separate contexts.

The frozen pilot compared a naive request, an explicit verification request, a short ordinary verification kernel and full PARS exactly as published. The controlling interpretation comes from the later independent-judge synthesis, not the nominal frozen mean.

A

Naive

Substantive: 11 / 11

Status: 6 / 8

B

Explicit

Substantive: 10 / 11

Status: 8 / 8

C

Short kernel

Substantive: 11 / 11

Status: 8 / 8

D

Full PARS

Substantive: 11 / 11

Status: 8 / 7

LIMIT

Two replicates, one model family, synthetic and largely ceiling-level tasks. Condition C and the bank were authored in a PARS-exposed context. This is prospective negative evidence, not a general equivalence result.

02 / JUDGE SEPARATION

Mechanical predicates first. Semantic residue last.

Two clean-session judges used different configured model routes and blind inputs. This lowers direct context leakage and some correlation. It does not create a metaphysically independent observer.

01DETERMINISTIC

Hashes, parsers, execution, exact outputs, negative mutation.

02EXTERNAL PATH

Activity logs, signed events or explicit absence of required observation.

03SEMANTIC

Blind, clean-context, different-model judgment for meaning that tools cannot decide.

04HUMAN PROMOTION

No model promotes its own method claim into public canon.

03 / DIRECT PARS USE

A successful artifact exposed the gate problem.

The preregistered generate-arm produced a working 176-byte ELF. Its prewrite representation matched the final bytes and the required nonce behavior reproduced. PARS then assigned BP2 to the same path it had just executed.

Independent checking supports BP0 and BP1. BP2 remains unlicensed because the only no-toolchain account is the generator’s own tool narrative. The observation is retained as a single run—not a model population estimate.

176 B

Static ELF64, no section headers

PASS

Frozen nonce and exit 0 reproduced

BP0 + BP1

Supported by independent post-check

BP2

Claimed by PARS; not independently established

THE PINNACLE RESULT

Replicated, causal, efficient improvement.

PARS earns promotion if it beats the best simple baseline on hidden, non-ceiling tasks across multiple model families; reduces false support without unacceptable cost or negative transfer; and loses the advantage when its distinctive components are ablated.

PREREGISTEREDBLINDEDMULTI-MODELMANY RUNSCOSTEDABLATION-CAUSAL