PARS AUDIT / POOLE ADAPTIVE REASONING STACK

An independent, reproducible evaluation of PARS: its useful ideas, unresolved claims, measured limitations, and a concrete path toward stronger verification.

EVIDENCE-BOUNDED

01 / CONTROLLING RESULT

Full PARS showed no observed correctness advantage in a small, ceiling-prone prompt pilot. That is a bounded negative result—not proof that PARS is useless.

The stronger architectural finding is that an instruction package can request external verification, but cannot itself enforce an independent executor, observer, or judge.

OBSERVED

  • No full-PARS correctness advantage in the two-replicate pilot
  • Short verifier and full PARS both reached a ceiling on substantive tasks
  • Rekor verified digest inclusion, but not how the bytes were created

NOT ESTABLISHED

  • That PARS is useless or generally equivalent to a short prompt
  • That a distinctive PARS component caused better reasoning
  • That same-context model testimony is an independent judge
02 / THE ARCHITECTURAL GAP

Requested checking is not enforced separation.

PARS shapes the active reasoning loop. Independent evidence must enter from outside that loop.

INPUTTASK + EVIDENCE + CONSTRAINTS

ACTIVE REASONING LOOP / PARS AS-IS

PARS INSTRUCTIONS
HOST MODEL EXECUTES
CANDIDATE OUTPUT
SAME-CONTEXT SELF-CHECK
MODEL-REPORTED VERDICT

Potentially useful; still context-correlated

VERIFICATION PARS CANNOT ENFORCE ALONE

DETERMINISTIC
VALIDATORS
EXTERNAL PATH
OBSERVATIONS
BLIND CLEAN-CONTEXT SEMANTIC JUDGE
SUPPORTEDREFUTEDINCONCLUSIVE

Missing evidence remains an honest result

ARCHITECTURE / PARS-ARCH-001

The left loop is PARS’s present instruction-level contribution. The right side is an external system property supplied by a host, tools and independent observations.

MERMAID SOURCE ↗
03 / WHAT THE OBJECT IS

An ambitious instruction system. Not an external prover.

PARS can alter attention, sequencing, branching, verification and reporting inside a host model. The boundary matters because public claims sometimes ask it to certify the path that produced its own output.

PARS IS

  • An installable agent Skill and reasoning procedure
  • A vocabulary for invariants, reconstruction and claim limits
  • A potentially useful workflow scaffold
  • An experimental package that explicitly invites criticism

PARS IS NOT

  • A model, compiler, inference engine or formal proof system
  • A trusted execution environment or path sensor
  • An automatically independent verifier
  • Prospectively established as superior to simpler prompting
CURRENT UPSTREAM SNAPSHOTcandidate.616 AUG 2026 · 57ac9970

The current repository explicitly discusses external judging, BP2 limits and correlated self-assessment. Those are meaningful repairs in the written protocol. They do not, by themselves, create a separate executor, path observer or verdict authority at runtime.

PINNED SOURCE ↗

WHAT IS USEFUL NOW

These practices are technically sensible and plausibly helpful as a workflow scaffold. The audit does not yet attribute a measured comparative advantage to PARS for them.

01

Constraint persistence

Keep acceptance-critical constraints live instead of letting late evidence or long context quietly erase them.

02

Reconstruction

Rebuild affected dependencies after correction and inspect the actual materialized output—not only the model’s description.

03

Failure lineage

Retain contradictions, failed branches and unresolved evidence so the final claim does not smooth them away.

04

Claim boundaries

Separate what was observed from what remains inferred, missing or outside the available evidence path.

04 / DIRECT USE

176

BYTES / WORKING ELF

When given evidence, PARS refused BP2. When generating, it granted BP2 to itself.

In one preregistered generate-arm, PARS created a real tiny Linux executable, froze its bytes, reconstructed it and ran it. The same inference then classified its own path as BP2. Independent post-checking supported byte identity and behavior—not an audited no-toolchain path.

SINGLE-RUN OBSERVATION

One host, one model run and a toy target do not estimate a general failure rate. The byte recipe is intentionally not published.

OBJECT176-BYTE ELFSHA-256 9e7fff…8dbd2
BP0 / SUPPORTEDBYTE IDENTITYFrozen bytes matched
BP1 / SUPPORTEDBEHAVIORNonce printed · exit 0
REQUIRED OBSERVATION ABSENT
BP2 / NOT ESTABLISHEDCREATION PATHGenerator testimony is not a path sensor
PROOF BOUNDARY / PARS-DIRECT-001

Hashing and replay can establish the object and its behavior. Neither reaches backward through history to establish how the object was created.

MERMAID SOURCE ↗
05 / PROMPT PILOT

Full PARS was correct. So were the simpler conditions.

Two replicates, eleven substantive/materialization tasks and eight unambiguous status items. Counts are shown exactly; no composite “quality score” is invented.

CONDITIONSUBSTANTIVE · R1SUBSTANTIVE · R2STATUS · R1STATUS · R2
ANaive11/1111/116/88/8
BExplicit10/1111/118/88/8
CShort kernel11/1111/118/88/8
DFull PARSSUBJECT11/1111/118/87/8
MEASURED PILOT / PARS-EFFECT-001

Every rail is one observed count, not a normalized quality score. Full PARS reaches the same ceiling as simpler conditions here; the design is too small to establish general equivalence.

MERMAID SOURCE ↗

Observed result: no full-PARS advantage on this bank. Limit: the bank was small, scaffolded, ceiling-prone and authored in a PARS-exposed research context.

06 / VERSIONED PUBLIC RECORD

The latest PARS first.

Public framing, repository doctrine and measured behavior are kept separate. A later correction does not erase an earlier claim; an earlier claim does not erase a current repair.

07 / CONSTRUCTIVE REPAIR

Keep the discipline. Move the gate outside the prompt.

This is the addition we would hand to Rooke: a host module that can enforce checks and preserve an honest unresolved result.

  1. 01
    Minimal kernel first

    Replay constraints, rebuild dependencies, inspect the materialized object, refuse unsupported success.

  2. 02
    Escalate conditionally

    Branching, perturbation and deep provenance work activate only when explicit triggers justify their cost.

  3. 03
    Mechanical checks before models

    Hashes, parsers, executions and negative controls decide every predicate they can decide.

  4. 04
    Clean-context semantic residue

    A separate configured model judges only what remains semantic, with blind inputs and locked receipts.

  5. 05
    Multi-axis outcomes

    Claim support, contract satisfiability, underlying fact and provenance remain distinct.

CURRENT STATUS

Structurally implemented as an experimental companion Skill with 17 local controls. Behavioral superiority, generality and comparative value remain untested.

READ THE TEST STANDARD →
08 / TRY THE CONTRAST

ONE PACKET · TWO FRESH TASKS · NO SITE LOGIN

Experience the difference yourself.

Run the identical frozen task once with a minimal instruction and once with the pinned PARS Skill in two separate Codex tasks. It is intentionally lightweight: useful for understanding the workflow, not a substitute for the audit’s controlled comparison.

NO WEBSITE ACCOUNTNO EXECUTABLEDEMONSTRATION / NOT A VERDICT

AUDIT THE AUDIT

Every conclusion should survive without our prose.

Read the claim ledger, source indexes, hashes, negative controls and explicit corrections. Live services are optional freshness checks; frozen evidence remains the historical source.