Writing · No. 3

The Evidence Standard

Every other transaction an auditor examines must be supported by independent evidence. AI spend should not be the exception.

Ask any auditor what makes a transaction supportable and the answer is boring on purpose: evidence produced or confirmed by someone other than the party who benefits from the number. A payroll figure is checked against time records the payee didn't write alone. An inventory balance is checked by counting the shelf, not by reading the warehouse's own summary. A bank balance is confirmed with the bank, not with the bookkeeper. The whole discipline rests on one dull, load-bearing rule: the measured party is not the sole source of the measurement.

Then the invoice is for AI, and the rule quietly disappears.

A metered AI service today is often one continuous party from end to end: it runs the workload, counts the tokens, applies the rate, computes the charge, and produces the dashboard offered as the record of all of it. The number on the invoice and the evidence for the number come from the same hand. In any other line of the ledger, an auditor would call that what it is — a self-generated record — and go looking for something independent to corroborate it. For AI spend, in most organizations, there is nothing independent to find.

This is not an accusation of dishonesty. Most vendors count carefully. It is a statement about evidence. A careful count that only the counter can vouch for is still, as a matter of audit practice, unsupported. Honesty is not the standard; verifiability is. That distinction is the entire reason audit standards exist — they were not written for the dishonest vendor, they were written so nobody has to guess which vendor that is.

The stakes are no longer small

While the evidence rule lapsed, the line item grew. AI spend is moving from experiment to obligation — a budgeted, recurring, board-visible figure, and in government, an appropriated one. Growth alone would argue for scrutiny. But the structure of the spend argues harder: the waste lives in quantities, not rates. Context resent on every turn. Retrieval fetched and never read. Retries billed alongside the calls they retried. A claimed optimization is a claim about quantities — and a quantity claim graded by its claimant is not a finding, it is a press release.

Savings claims make the gap sharpest. "We cut your AI spend thirty percent" is an assertion about a counterfactual — what you would have paid — measured by the party being paid to move the number. No CFO would accept that construction from a freight auditor or an energy-savings contractor without independent measurement. It arrives from AI vendors daily, and it is accepted, because there is no standard that says it shouldn't be.

What the standard should say

The fix is not a product. It is a criterion — a test any conforming implementation can pass, stated in the same vocabulary auditors already use: segregation of duties, tamper-evidence, and the basis of evidence.

Independent attestation of AI usage records

Usage and spend records should be countersigned by a party that:

  1. Does not generate or route the workload it measures — separation of generation and attestation;
  2. Cannot alter the record after signing — tamper-evident, integrity-protected structures;
  3. States the basis of every quantity — independently verified, or provider-asserted;
  4. Can be checked before deployment — by architecture review, not post-hoc assurances.

A record signed solely by the party being measured should not, by itself, constitute audit evidence.

Each element earns its place. Separation is the substance — without it, the countersignature is a second signature from the same interest. Integrity protection makes the record durable against revision, which is what converts a log into evidence. Basis designation is the honesty mechanism: real deployments contain quantities the attester could verify and quantities it could only receive, and a record that refuses to say which is which is asking for trust it hasn't earned. And verifiability before deployment is what distinguishes an architectural property from a policy promise — a control you can inspect on day zero, rather than a paragraph in someone's terms of service.

Note what the criterion does not require: content. The attested artifacts are quantities, rates, and cryptographic digests. Nothing in the standard asks anyone to read, retain, or disclose a prompt, a document, or a line of a workload. The evidence question and the privacy question have the same answer — attest the count, never the content — which is how a standard earns adoption in environments where both auditors and privacy officers hold a veto.

The record will be requested

Provenance thinking has already arrived at the content layer: the industry has begun cryptographically signing what AI produces. The spend layer is next, because the auditors are next. Financial statement audits, inspector-general reviews, improper-payment examinations — each runs on document requests, and the request for AI usage support is coming to organizations that have never had to answer it. The ones who fare well will not be the ones with the most persuasive dashboard. They will be the ones holding a record someone else signed.

We publish this criterion in the open, stated so that anyone — including our competitors — can be measured against it. A test only one vendor can pass is not a standard; it is an advertisement. This one is neither. It is the same rule every other line of the ledger has always lived under, arriving late to the newest line item on the books.

Interest disclosure: TokenMark™ is a commercial AI spend assurance product built to satisfy the criterion described here, and its developers hold pending U.S. patent applications in this field. Versions of this criterion have been offered, with this interest disclosed, for consideration in public standards processes. The criterion has not been adopted by NIST, GAO, OMB, or any government body, and no affiliation or endorsement is implied.

Programs & participation
NVIDIA InceptionMember
NIST Zero DraftsSubmissions filed
NIST AI 300-1Public comment
NIST NCCoEPost-Quantum Cryptography
Community of Interest
DOE Genesis MissionConsortium participant
Congressional Internet CaucusAdvisory Group — former member

Participation in an open public process is not endorsement. No agency, standards body, consortium or company listed here endorses Atom Works™, its products or its claims.