VISITOR-OWNED EXPERIMENT / ALPHA
Compare the protocol.
Keep the evidence.
Four frozen instructions, one task packet, fresh contexts and a mechanical scorer. Your Codex installation runs every selected arm; this page never sends results to Arkhē.
01 / READINESS
Local execution boundary
The runner is looking for a ChatGPT-authenticated installation.
- Inference
- Your ChatGPT/Codex allowance
- API keys
- Removed from every child process
- Transport
- 127.0.0.1 only; no upload endpoint
- Persistence
- Local logs and exportable JSON receipt
02 / FROZEN COMPARISON
Select the arms to run
A full four-arm demonstration may consume substantial Codex usage. A trivial local canary reported 23,617 aggregate tokens; that is a route-overhead observation, not a prediction for each arm.
———03 / RUN
Fresh contexts in randomized order
QUEUED04 / LOCAL RECEIPT
Observed outputs
These scores describe this local run only. Compare correctness, failure modes, tokens and latency; do not promote a one-run winner.
| Arm | Condition | Score | Tokens | Time | State |
|---|
PROOF BOUNDARY
The runner can prove which frozen inputs were used, which local files materialized and how deterministic scoring was applied. It cannot make a same-run semantic judge independent, prove global superiority or guarantee that every visitor has an identical model route.