Artificial intelligence has a scoring problem, and it is the same problem at every layer.
In 2026 the industry named a phenomenon it had tolerated for years: benchmaxing. Models were being tuned to score well on a small set of public benchmarks, and the scores were being presented as evidence of capability. Palantir's chief technologist put the objection plainly — the benchmarks measure what the benchmarks measure, and that has almost nothing to do with the task a customer is trying to accomplish. His prescription was to move evaluation onto an empirical basis: build the measurement that reflects your reality, then see which model actually wins.
He was describing one instance of a general defect. Once you see its shape, you find it everywhere in the stack:
| Layer | The claim | Who produces the number |
|---|---|---|
| Model | "This model is the most capable." | The lab that trained it, on benchmarks it optimized against. |
| Gateway / optimization | "We cut your AI spend 30%." | The vendor paid a share of the savings it reports. |
| Agent | "The task completed successfully." | The agent that performed the task, writing its own log. |
Three layers, one structure: the party that acts is the party that scores the action. Every incentive points the same direction, and no amount of good faith changes the geometry. This is not a claim about anyone's honesty. It is a claim about who holds the pen.
In August 2026 the Linux Foundation launched the Tokenomics Foundation with thirty member organizations — among them JPMorganChase, BNY, IBM, Oracle, SAP, ServiceNow, and Accenture — to build open, vendor-neutral standards for the economics of AI. Its published roadmap covers definitions of token value and density, a reference model for the full cost of AI, a standard method for cost to serve expressed per call rather than per token, a framework relating spend to outcomes, and token cost telemetry carried in the FOCUS billing specification.
This is the right work and it was overdue. It is also, at present, entirely concerned with what to measure. No published line of that roadmap addresses who certifies a figure and whether anyone else can recompute it — which is the defect this paper describes. A shared unit of measure inherits the trustworthiness of whoever does the measuring.
A founding member said so at launch. The president of Cast AI, describing where the industry's attention is going: it is standardizing the meter in the middle and ignoring both ends.
The prescription for benchmaxing was empirical measurement against your own reality. The prescription at the cost layer is the same idea carried to its conclusion: the entity that produces a number must not be the entity that certifies it.
In practice that means four properties, none of which is exotic. They are ordinary separation of duties, applied to software:
We call this AI Spend Assurance, and we use the term generically on purpose. It names a layer inside tokenomics rather than a rival to it. It is not a product name. It is a description of a class of controls that AI buyers will require — the same way transaction assurance became unremarkable in payments once the volume got large enough that faith stopped being a control.
Assurance is not optimization. An optimizer's job is to make the number smaller. An assurer's job is to make the number true. When the same vendor does both without separation, the second job quietly loses.
A company arguing this cannot exempt itself. TokenMark™ is built so that the component performing the optimization is structurally incapable of signing the savings record — the attester runs in a separate privilege domain and the operating system denies the optimizer access to its key. We publish savings figures only when they are attester-signed and independently verified by the customer. And when a pilot returns a number lower than we hoped, we publish that number.
If the argument in this paper is correct, it applies to us first.
Atom Works™, Inc. builds attestation infrastructure for AI systems. We hold pending patents in this area and therefore have a commercial interest in the criteria described above. We have tried to state them so that any implementation satisfying them — including one that is not ours — would count.
Participation in an open public process is not endorsement. No agency, standards body, consortium or company listed here endorses Atom Works™, its products or its claims.