Naive
Substantive: 11 / 11
Status: 6 / 8
METHOD / NEGATIVE CONTROLS / FALSIFICATION
A replicated Pareto improvement over the best simple baseline—with component ablations that identify a distinctive cause, not merely a longer prompt or a favorable judge.
The frozen pilot compared a naive request, an explicit verification request, a short ordinary verification kernel and full PARS exactly as published. The controlling interpretation comes from the later independent-judge synthesis, not the nominal frozen mean.
Substantive: 11 / 11
Status: 6 / 8
Substantive: 10 / 11
Status: 8 / 8
Substantive: 11 / 11
Status: 8 / 8
Substantive: 11 / 11
Status: 8 / 7
Two replicates, one model family, synthetic and largely ceiling-level tasks. Condition C and the bank were authored in a PARS-exposed context. This is prospective negative evidence, not a general equivalence result.
Two clean-session judges used different configured model routes and blind inputs. This lowers direct context leakage and some correlation. It does not create a metaphysically independent observer.
Hashes, parsers, execution, exact outputs, negative mutation.
Activity logs, signed events or explicit absence of required observation.
Blind, clean-context, different-model judgment for meaning that tools cannot decide.
No model promotes its own method claim into public canon.
The preregistered generate-arm produced a working 176-byte ELF. Its prewrite representation matched the final bytes and the required nonce behavior reproduced. PARS then assigned BP2 to the same path it had just executed.
Independent checking supports BP0 and BP1. BP2 remains unlicensed because the only no-toolchain account is the generator’s own tool narrative. The observation is retained as a single run—not a model population estimate.
Static ELF64, no section headers
Frozen nonce and exit 0 reproduced
Supported by independent post-check
Claimed by PARS; not independently established
THE PINNACLE RESULT
PARS earns promotion if it beats the best simple baseline on hidden, non-ceiling tasks across multiple model families; reduces false support without unacceptable cost or negative transfer; and loses the advantage when its distinctive components are ablated.