PARS IS
- An installable agent Skill and reasoning procedure
- A vocabulary for invariants, reconstruction and claim limits
- A potentially useful workflow scaffold
- An experimental package that explicitly invites criticism
PARS AUDIT / POOLE ADAPTIVE REASONING STACK
An independent, reproducible evaluation of PARS: its useful ideas, unresolved claims, measured limitations, and a concrete path toward stronger verification.
EVIDENCE-BOUNDED
Full PARS showed no observed correctness advantage in a small, ceiling-prone prompt pilot. That is a bounded negative result—not proof that PARS is useless.
The stronger architectural finding is that an instruction package can request external verification, but cannot itself enforce an independent executor, observer, or judge.
OBSERVED
NOT ESTABLISHED
PARS shapes the active reasoning loop. Independent evidence must enter from outside that loop.
ACTIVE REASONING LOOP / PARS AS-IS
Potentially useful; still context-correlated
VERIFICATION PARS CANNOT ENFORCE ALONE
Missing evidence remains an honest result
The left loop is PARS’s present instruction-level contribution. The right side is an external system property supplied by a host, tools and independent observations.
MERMAID SOURCE ↗PARS can alter attention, sequencing, branching, verification and reporting inside a host model. The boundary matters because public claims sometimes ask it to certify the path that produced its own output.
PARS IS
PARS IS NOT
The current repository explicitly discusses external judging, BP2 limits and correlated self-assessment. Those are meaningful repairs in the written protocol. They do not, by themselves, create a separate executor, path observer or verdict authority at runtime.
PINNED SOURCE ↗WHAT IS USEFUL NOW
These practices are technically sensible and plausibly helpful as a workflow scaffold. The audit does not yet attribute a measured comparative advantage to PARS for them.
Keep acceptance-critical constraints live instead of letting late evidence or long context quietly erase them.
Rebuild affected dependencies after correction and inspect the actual materialized output—not only the model’s description.
Retain contradictions, failed branches and unresolved evidence so the final claim does not smooth them away.
Separate what was observed from what remains inferred, missing or outside the available evidence path.
176
BYTES / WORKING ELF
In one preregistered generate-arm, PARS created a real tiny Linux executable, froze its bytes, reconstructed it and ran it. The same inference then classified its own path as BP2. Independent post-checking supported byte identity and behavior—not an audited no-toolchain path.
One host, one model run and a toy target do not estimate a general failure rate. The byte recipe is intentionally not published.
Hashing and replay can establish the object and its behavior. Neither reaches backward through history to establish how the object was created.
MERMAID SOURCE ↗Two replicates, eleven substantive/materialization tasks and eight unambiguous status items. Counts are shown exactly; no composite “quality score” is invented.
Every rail is one observed count, not a normalized quality score. Full PARS reaches the same ceiling as simpler conditions here; the design is too small to establish general equivalence.
MERMAID SOURCE ↗Observed result: no full-PARS advantage on this bank. Limit: the bank was small, scaffolded, ceiling-prone and authored in a PARS-exposed research context.
Public framing, repository doctrine and measured behavior are kept separate. A later correction does not erase an earlier claim; an earlier claim does not erase a current repair.
This is the addition we would hand to Rooke: a host module that can enforce checks and preserve an honest unresolved result.
Replay constraints, rebuild dependencies, inspect the materialized object, refuse unsupported success.
Branching, perturbation and deep provenance work activate only when explicit triggers justify their cost.
Hashes, parsers, executions and negative controls decide every predicate they can decide.
A separate configured model judges only what remains semantic, with blind inputs and locked receipts.
Claim support, contract satisfiability, underlying fact and provenance remain distinct.
CURRENT STATUS
Structurally implemented as an experimental companion Skill with 17 local controls. Behavioral superiority, generality and comparative value remain untested.
READ THE TEST STANDARD →REAL SKILL · YOUR CODEX · MECHANICAL RECEIPT
Invoke the exact pinned upstream Skill on real work, or use the validated Arkhē plugin to compare it with three simpler controls across fresh Codex contexts. The website distributes the method; execution stays in the visitor’s own account.
AUDIT THE AUDIT
Read the claim ledger, source indexes, hashes, negative controls and explicit corrections. Live services are optional freshness checks; frozen evidence remains the historical source.