Proof Desk

Paste the “it’s done” report and the output behind it — get a per-claim verdict.

Back to SkillSafe
Or pick a file: it is read locally, nothing uploads until you run.
Long logs are cut tail-heavy before sending — a traceback’s diagnosis is in its last lines.
Context — the project, the real commands, what is already known flaky
How it works

Nothing to hand? Load the — a coding agent’s confident summary with a lint pass standing in for a test run, a regression test that never went red and a failing test buried in the log — or the , where the correct verdict is that the claims hold and the useful output is what is still unproven before a deploy.

1

Paste both sides — the prescan is free

No upload, no AI: the prescan reads the report and the log in your browser. It pulls out every sentence that makes a claim and sorts it into a kind — tests, lint, build, regression, bugfix, delegated agent, requirements, deploy, performance, manual, or a plain “it’s done” — then detects what the log actually contains: test summaries, build results, exit codes, diffs, coverage, red-green cycles, CI conclusions, deploy health checks. It shows you, before you spend anything, how many claims have no evidence of the kind they need, and it flags the language the skill this app is built on names as red: hedging, celebration before verification, stale evidence, partial checks, the linter cited for a build, an agent’s self-report taken as proof, “just this once”.

2

The gate rules on each claim — this is the metered part

A release engineer’s pass: every claim gets a status — verified, partially verified, unverified, contradicted — a risk rating scaled to what happens next, the verbatim line of your log that supports it, and, when it is open, the exact command that would settle it and what output would count as proof. A failure signal in the log outranks any sentence in the report. Nothing is invented: an unsupported claim gets an empty evidence field rather than a plausible-sounding quote. Pricing is honest — a worst-case amount is reserved before the run and only what the run uses is charged; the meter next to the button shows both.

3

Run the commands, paste the output, re-check for free

The gate hands back a runnable verify.sh — one step per open claim, cheapest-first, exiting non-zero if any step fails — plus a tickable pre-merge checklist, the ledger as CSV, an honest status message that separates what is proven from what is merely believed, and, if the report came from a coding agent, a follow-up addressed straight back to it: every open claim, the command that would settle it, and an instruction to paste the raw output rather than another summary. Each proof command has its own copy button, without the shell prompt. Then go and run them. When you come back with fresh output, the prescan re-runs as you type and the strip above the button counts the flags you cleared, the ones still open and any you have just introduced — in the browser, for free, before you pay for a second gate. Gates are saved to your SkillSafe account when you are signed in, and restoring one puts both the report and the evidence back in the form.

Derived from the @obra/verification-before-completion skill (obra/superpowers, MIT) — evidence before claims, always.