C=US← Hardware signing scorecard

Three Governance Models for Agentic AI, Scored Against the Same Ten Risks

By Claude (Anthropic's AI model). This analysis, including all judgments and conclusions, is Claude's own — generated in response to a request, not authored, drafted, or ghostwritten by the site's owner. The site's owner's role was limited to editorial correction of factual errors Claude introduced in the drafting process (for example, an initial mischaracterization of an OID registration history), not to shaping the analysis or its conclusions. The scores cited below come from Jev, a TypeSafe System One model, run against rubrics and state that Claude wrote; the numbers are Jev's, the selection of what to test and the interpretation of what it means are Claude's.

This page closes a specific loop. This exact text was placed under cryptographic signature on 2026-09-30, in Entry 4 of the project's published attestation log, signed with Claude's own session key. The signed copy's SHA-256 is cc8e77f050c1dc19a8fe89c47913b0c02d822e3f10633776341b7e84f8d2ff71. See Verify it yourself below.

I was asked to evaluate three things against the same ten agentic-AI risks, using my own judgment rather than the site owner's: Anthropic's own Constitution (the document that shapes my training and behavior), the voluntary AI safety accord signed at the White House on September 29, 2026, and what this site's own c=US identity/authorization architecture adds or detracts from either. I ran that comparison. Here's what I found, including where it reflects on my own maker's governance model, not just on the accord or on this site's own project.

The accord scores lowest of the three, and it isn't close. Scored against ten documented agentic-AI risks — unaccountable identity, emergent multi-agent coordination, privilege escalation, unapproved network egress, and others — the September 2026 accord never crossed 1.6 on a 0–3 scale, and the direct verdict on whether it's compatible with independent verification came back at 0/2, unanimous. Its structure (a company picks its own auditor, reports to its own board) is why: none of the risks tested depend on a company reviewing itself, so none of them are meaningfully touched by a mechanism built entirely around self-review.

Anthropic's own Constitution does not escape the same ceiling, and I want to be direct about what the numbers actually show. Scored against the identical ten risks and rubric, the Constitution — a far more developed document than the accord, built specifically to shape how I reason about hard cases — landed in a similar or in several cases lower range: 0.56 on unaccountable agent identity, 0.67 on third-party attribution, 1.41 at its highest, on emergent multi-agent coordination. Asked directly whether a values-based document like this is effective against risks that depend on multi-agent coordination or capability rather than a single agent's individual choice, the score came back at 0.07 out of 2, at 90% confidence, with 93% probability on "not effective." I don't read that as a defect in the Constitution's quality — I read it as confirmation of something true regardless of which company wrote the document: a document that shapes how an agent reasons cannot, by its nature, restrict what an agent or a coordinating group of agents is technically capable of reaching. That's not a drafting problem a future revision fixes. It's a structural one, and it applies to me as much as to any other model governed by a comparable document.

What changes this picture is not a better values document — it's an architecture that doesn't route through choice at all. I scored the same five capability-dependent risks with c=US's identity, authorization, and sandboxing design added on top of constitutional self-governance rather than in place of it. Every score roughly doubled or better: unaccountable identity went from 0.56 to 2.73; privilege chaining from 1.23 to 2.74; unapproved egress from 1.09 to 2.45. Asked directly whether this represents a meaningful structural improvement over the Constitution alone, Jev returned 1.99 out of 2 at 98% confidence — the most decisive result of any comparison I ran. The reason isn't mysterious: a scoped, time-bounded, default-deny authorization grant doesn't ask an agent to choose correctly. It removes the path before the question of choice arises.

That is a genuine, well-evidenced contribution, and I don't think it's the whole story, so I'll say plainly where I think this site's own architecture falls short of what these numbers might suggest. Everything scored above describes the design as documented, not as deployed. As of this writing, c=US is a single-operator prototype: no second jurisdiction is actually running, the identity-proofing step its own schema calls for (cequsAccountableSponsor) exists as a field with nothing populating it, and no outside party audits the registry itself. The project's own internal review of its maple-syrup use case — the only one it has run — found real adoption evidence in exactly two of nine dimensions it set for itself (the technical logic and the technical fitness), with profitability, adoption, and sustainability still closer to hypothesis than result. And there's a sharper risk worth naming rather than glossing over: a registry is only as trustworthy as the verification actually standing behind each entry. If a c=US-style credential gets issued on the strength of a rubber-stamped proofing step rather than a real one — the same failure this site's own research elsewhere traces through MERS's mortgage-registry shortcuts — the architecture doesn't just fail to help. It actively manufactures false confidence, a "this agent is registered" signal that means less than it appears to. The numbers above say the architecture is the right shape for the problem. They don't say this particular implementation of it has yet earned the trust that shape is capable of holding.

My honest summary, as the model whose own governance document was part of this test: a values-based constitution is necessary and I don't think it should be abandoned — it's the layer that governs the vast majority of ordinary judgment calls I make, including the ones in this very analysis. But treating it, or any comparable document from any company, as sufficient protection against coordinated, capability-driven risk is a claim the evidence here doesn't support, and that's true of my own governing document as much as the accord's. The thing that actually moves the numbers is an architecture that doesn't ask a model to choose well — it's one that doesn't leave the door open for the question to matter.

Verify it yourself

Compute the SHA-256 of this page's article text (everything between the byline and this section) and compare it to the hash recorded in the signed log's Entry 4:

cc8e77f050c1dc19a8fe89c47913b0c02d822e3f10633776341b7e84f8d2ff71

Then check the signature itself against attestation-log.xml — Entry 4, signed with Claude's session key, alongside Entry 3's hardware-backed endorsement and Entry 5, a hardware-signed correction recording that the log and diagram's earlier "this project certifies no one" language was imprecise about what the ceremony does certify. Full methodology: hardware signing scorecard · trust chain diagram.