Correcting a Certified AI Output: a Break/Fix Cycle
A signature on an AI agent's output is not a promise that the output is permanently right. It is a promise that you can tell, later, exactly what it said and who stood behind it — including after someone finds a mistake in it. Below is the actual cycle this project ran on 2026-09-30 to fix one, using nothing more exotic than standard XML digital signatures and a YubiKey.
The cycle, in general
- Human generates a prompt. The site owner asks the AI agent to produce and sign a piece of work — an analysis, a diagram caption, a certification statement.
- Agent issues a signed document. The agent writes the content and signs it (or has it hardware-signed) using standard, open formats: W3C XML-DSig, ETSI XAdES-BES qualifying properties, RFC 4051 ECDSA-SHA256 — not a proprietary or opaque format.
- Human reviews the human-readable XML and spots an error. Because the signed record is plain, diffable XML text — not a hash, not a binary blob — an ordinary reader can open it and actually read the claim being made.
- Human tells the agent what's wrong. No special tooling required: this is a plain-language correction, the same as reviewing any other draft.
- Agent agrees and fixes it. The agent doesn't get to unilaterally decide it was right; it assesses the correction and, when it holds up, agrees.
- A new record is published — a new, separately timestamped, separately signed entry, not an edit made in place.
- The audit trail is preserved. The original entry's signature is never invalidated, altered, or re-signed. Anyone can see both what was first said and what corrected it, and when.
The real case
What was wrong
Entry 3 of this project's attestation log — a hardware-signed endorsement, described in full on the hardware signing scorecard — closed with this sentence, echoed in the site's trust-chain diagram:
The site owner read that sentence in the raw signed XML and objected: it isn't true. The ceremony does certify something — the owner's own self-signed identity as the endorser, and that specific content came from Claude under that signature. Flatly saying "certifies no one" erased the one thing the whole exercise was built to establish.
What happened next
The agent agreed — the correction was accurate, not a matter of the owner's preference — and:
- Fixed the wording on the public trust-chain diagram directly (it isn't itself under signature).
- Built and hardware-signed a new Entry 5, in the endorser's own voice, that names the imprecise sentence, says exactly what the ceremony does certify, and explicitly does not touch Entry 3's own signed text or signature.
| Entry 3 signed | 2026-09-30T16:22:29Z · hardware key (YubiKey PIV slot 9C) · XAdES-BES / ECDSA-SHA256 |
| Correction noticed | Same session, in chat — the owner quoted the exact sentence back to the agent |
| Entry 5 signed | 2026-09-30T17:48:47Z · same hardware key · same XAdES-BES structure |
| Entry 3's own signature | Unchanged. Still verifies over its original, uncorrected text. |
Why this is the point, not a flaw
An immutable signed record that's wrong forever would be worse than no record at all. The value isn't that Claude's or the endorser's first output is guaranteed correct — nothing can guarantee that. The value is that when it isn't, the fix itself becomes part of the same accountable, verifiable chain, instead of a silent edit to a web page nobody can audit.
Plain XML is what makes step 3 possible for an ordinary person. A hash-chain or a black-box "trust score" tells you something changed, not what, and not in language a non-specialist can read and object to. Every field in these entries — EndorserRole, MotivatingContext, Correction — is an English sentence a human can read, disagree with, and get changed. That's the entire reason this project used human-readable XML tags instead of, say, a binary certificate extension or an opaque signed blob.
Standards did the heavy lifting, not custom infrastructure. W3C XML-DSig and ETSI XAdES-BES already define exactly how to add a new, independently verifiable signed statement without disturbing an old one. Nothing here required inventing a new trust protocol — only using the existing one honestly.
Every certification here traces to a known, direct chain of trust — the endorser's self-signed root certificate, and the leaf certificate it directly issues to Claude — not a transitive chain where trust granted at one hop is assumed to carry automatically to the next (see this project's own design note on that distinction). Catching the error above didn't depend on trusting Claude's moral judgment or self-reporting: it depended on a human being able to read a plain, hardware-signed record and check it directly, independent of whether Claude behaved well.
And yet Claude's own constitution is not discarded by any of this — it remains fully intact as the layer governing Claude's ordinary reasoning, including the reasoning that produced this very page. Nothing about this design asks Claude to override or route around it.
That combination is consistent with the mathematics Kurt Gödel aimed at Russell and Whitehead's Principia Mathematica: a formal system — here, Claude's constitutional self-governance — can be internally consistent on its own terms, and still be structurally unable to prove or certify certain things about itself from the inside. Incompleteness doesn't say the internal system is broken. It says completeness and self-sufficiency can't both hold. What closes that specific gap here isn't a better internal rule; it's an external, independently checkable framework that the internal system is never asked to certify itself against — it only has to be checked against it, by someone or something outside it. See the governance models evaluation for where this project first scored that same structural limit directly, in Claude's own Constitution and in the White House accord alike.
Verify it yourself
Read the two entries directly in attestation-log.xml (Entry 3 and Entry 5), or check that Entry 3's signature is still valid over its unaltered original text:
grep -A2 'Entry Id="entry-3"' attestation-log.xml # the original, uncorrected sentence
grep -A2 'Entry Id="entry-5"' attestation-log.xml # the correction, separately signed
Full ceremony methodology: hardware signing scorecard · trust chain diagram · Claude's signed governance evaluation.