AI SPEND ASSURANCE · POSITION PAPER · WP-SG-001

The Self-Graded Stack

Artificial intelligence has a scoring problem, and it is the same problem at every layer.

1 · The pattern

In 2026 the industry named a phenomenon it had tolerated for years: benchmaxing. Models were being tuned to score well on a small set of public benchmarks, and the scores were being presented as evidence of capability. Palantir's chief technologist put the objection plainly — the benchmarks measure what the benchmarks measure, and that has almost nothing to do with the task a customer is trying to accomplish. His prescription was to move evaluation onto an empirical basis: build the measurement that reflects your reality, then see which model actually wins.

He was describing one instance of a general defect. Once you see its shape, you find it everywhere in the stack:

LayerThe claimWho produces the number
Model"This model is the most capable."The lab that trained it, on benchmarks it optimized against.
Gateway / optimization"We cut your AI spend 30%."The vendor paid a share of the savings it reports.
Agent"The task completed successfully."The agent that performed the task, writing its own log.

Three layers, one structure: the party that acts is the party that scores the action. Every incentive points the same direction, and no amount of good faith changes the geometry. This is not a claim about anyone's honesty. It is a claim about who holds the pen.

Benchmaxing is self-grading at the model layer. Spend reporting is self-grading at the cost layer. The industry has agreed the first one is a problem. The second one is larger, arrives monthly, and gets paid. Every vendor claims savings. Only one can be recomputed.

2 · The industry has started on half of it

In August 2026 the Linux Foundation launched the Tokenomics Foundation with thirty member organizations — among them JPMorganChase, BNY, IBM, Oracle, SAP, ServiceNow, and Accenture — to build open, vendor-neutral standards for the economics of AI. Its published roadmap covers definitions of token value and density, a reference model for the full cost of AI, a standard method for cost to serve expressed per call rather than per token, a framework relating spend to outcomes, and token cost telemetry carried in the FOCUS billing specification.

This is the right work and it was overdue. It is also, at present, entirely concerned with what to measure. No published line of that roadmap addresses who certifies a figure and whether anyone else can recompute it — which is the defect this paper describes. A shared unit of measure inherits the trustworthiness of whoever does the measuring.

A founding member said so at launch. The president of Cast AI, describing where the industry's attention is going: it is standardizing the meter in the middle and ignoring both ends.

A specification can carry this property, and it is far cheaper to include than to retrofit. A per-figure field recording whether a value was measured against a pre-pinned baseline, recomputed in a separate trust domain, and verified offline by a third party would be useful regardless of which vendor implements the underlying control — including vendors who license nothing from anyone. The count is only as good as who counted it.

3 · Why the cost layer is worse

4 · What the fix looks like

The prescription for benchmaxing was empirical measurement against your own reality. The prescription at the cost layer is the same idea carried to its conclusion: the entity that produces a number must not be the entity that certifies it.

In practice that means four properties, none of which is exotic. They are ordinary separation of duties, applied to software:

5 · The category

We call this AI Spend Assurance, and we use the term generically on purpose. It names a layer inside tokenomics rather than a rival to it. It is not a product name. It is a description of a class of controls that AI buyers will require — the same way transaction assurance became unremarkable in payments once the volume got large enough that faith stopped being a control.

Assurance is not optimization. An optimizer's job is to make the number smaller. An assurer's job is to make the number true. When the same vendor does both without separation, the second job quietly loses.

6 · The standard we hold ourselves to

A company arguing this cannot exempt itself. TokenMark™ is built so that the component performing the optimization is structurally incapable of signing the savings record — the attester runs in a separate privilege domain and the operating system denies the optimizer access to its key. We publish savings figures only when they are attester-signed and independently verified by the customer. And when a pilot returns a number lower than we hoped, we publish that number.

If the argument in this paper is correct, it applies to us first.

Atom Works™, Inc. builds attestation infrastructure for AI systems. We hold pending patents in this area and therefore have a commercial interest in the criteria described above. We have tried to state them so that any implementation satisfying them — including one that is not ours — would count.

Programs & participation
NVIDIA InceptionMember
NIST Zero DraftsSubmissions filed
NIST AI 300-1Public comment
NIST NCCoEPost-Quantum Cryptography
Community of Interest
DOE Genesis MissionConsortium participant
Congressional Internet CaucusAdvisory Group — former member

Participation in an open public process is not endorsement. No agency, standards body, consortium or company listed here endorses Atom Works™, its products or its claims.