Kratos Labs · Field note

Stop Measuring AI Cost in Tokens

Tokens measure one component of an AI system. The useful metric is the total lifecycle cost of producing an acceptable result at the required quality, reliability, risk, and speed.

Tokens are not a business outcome.

They are an accounting unit for one component of an AI system. They tell you how much text a model processed and produced. They do not tell you whether the work was correct, whether it arrived on time, whether a human had to repair it, or whether the system exposed the business to unacceptable risk.

Yet token usage is often treated as the cost of the system.

That is like measuring a delivery operation by fuel consumption while ignoring failed deliveries, damaged goods, driver time, vehicle maintenance, and refunds. Fuel matters. It is not the whole cost, and it is definitely not the outcome.

The useful unit is cost per acceptable business outcome.

The metric appeared while defining a research system

The distinction became clear while we were drafting a research proposal for a model-independent control system for multi-agent AI operations.

The proposed system would let models plan and act, while an external controller enforced permissions, approval gates, recovery rules, audit records, and resource limits. Models could change. The rules governing the operation could not.

That created a practical evaluation problem. Comparing token counts would reveal almost nothing about whether one configuration was better than another. A cheaper model might fail more often. A stronger model might need fewer retries. An extra verification step might increase the cost of a successful run while sharply reducing the expected cost of a bad one.

The proposal therefore used a harder metric:

Codetext
cost per verified successful operation=(model calls + tools + retries + checks + human review + delay + rework)/successful rule-compliant operations

The word verified matters. So does rule-compliant.

A system does not earn credit merely because it produced an answer. It earns credit when the result clears the required standard without violating the rules that make the operation safe and useful.

Tokens price a model call, not a system

Imagine two configurations performing the same business task.

Configuration A uses a cheap model. The first output often needs another attempt, and every result requires five minutes of human review.

Configuration B uses a more capable model. Its calls cost more, but it usually produces an acceptable result on the first attempt and needs only a quick check.

If you compare API spend, A wins.

If you compare the total cost of accepted work, B may be cheaper.

The same problem appears in more complex systems. Token cost ignores:

  • hosting and runtime infrastructure
  • third-party tools and data providers
  • retrieval, storage, and orchestration
  • automated checks and independent evaluation
  • retries, fallbacks, and recovery work
  • human review and correction
  • maintenance when models, prompts, or tools change
  • delay caused by failed or slow runs
  • the expected cost of failures that escape into the business

A low token bill can coexist with an expensive operation. It can also hide a system that appears efficient only because humans absorb the missing work.

An acceptable outcome has several thresholds

Business value is not simply "the model returned something."

An outcome is acceptable only when it reaches the required level across the dimensions that matter for that operation:

  • Quality: Is the result correct and useful enough for its purpose?
  • Reliability: Does the system clear that standard consistently?
  • Risk: Did it stay inside its permissions and avoid unacceptable harm?
  • Speed: Did it arrive while the result was still useful?
  • Review burden: How much human attention was needed before acceptance?

The threshold changes with the job.

A first-pass internal summary can tolerate more uncertainty than a payment approval. A draft support response can use a different review path from an infrastructure repair. A reversible low-impact action should not inherit the same model, checks, and human gates as a destructive production change.

This is why there is no universally cheapest AI model. There is only a cheapest configuration for a defined outcome and risk class.

The system has a minimum useful capability

Below the acceptance threshold, the system fails. A configuration that is cheap but produces unusable work has no economic advantage.

Above the threshold, more capability remains valuable only while the additional outcome justifies the additional cost.

That gives AI system design a lower and an upper boundary:

Codetext
below the threshold -> unacceptable outcomeat the threshold    -> cheapest reliable candidateabove the threshold -> buy more capability only when marginal value exceeds marginal cost

This is not an argument for always choosing a smaller model. It is an argument against choosing models by habit, prestige, or benchmark position.

The best configuration might use a small model for classification, deterministic code for rules, a stronger model for ambiguous cases, independent checks for important outputs, and a human gate for high-impact actions. It might also use the frontier model immediately when the cost of a failed first attempt exceeds the model premium.

The unit of optimization is the entire configuration, not the model in isolation.

Measure total lifecycle cost

A useful operating equation is:

Codetext
system efficiency=expected business value delivered/total lifecycle cost

The numerator asks what the system delivered. The denominator asks what the business spent to get it, keep it working, and absorb its failures.

For a repeatable workflow, we can also use the more concrete inverse:

Codetext
cost per acceptable outcome=total lifecycle cost/number of accepted outcomes

Both equations need a stable definition of acceptance. Otherwise the metric can be gamed by lowering the standard, skipping checks, or counting fluent output as successful work.

This is where an operation contract becomes useful. Before a run, define:

  • the result the business needs
  • the data and tools the system may use
  • the actions it may take
  • the rules that must never be violated
  • the checks required before acceptance
  • the retry and resource limits
  • the conditions that require human approval
  • the failure and recovery path

The model may propose the work. It should not be allowed to redefine what counts as success while doing it.

Optimize the cheapest reliable path

Once the outcome and its acceptance contract are clear, compare complete configurations rather than isolated calls.

  1. Establish a baseline. Run the task with the current model, tools, review process, and retry policy.
  2. Measure accepted outcomes. Record success, violations, retries, latency, review time, and recovery work.
  3. Test alternatives. Change one meaningful part of the configuration, such as model tier, reasoning budget, verification method, or retry limit.
  4. Include failure cases. Test prompt injection, missing data, tool errors, permission boundaries, and cases that previously failed.
  5. Route by risk and uncertainty. Give routine work the lightest configuration that clears the bar. Escalate ambiguous or high-impact work.
  6. Stop when added capability stops paying. More intelligence is not automatically more efficient.

This approach can reveal counterintuitive results.

More verification can lower total cost when it prevents expensive failures. A stronger model can lower total cost when it removes retries and review. A deterministic rule can beat every model when the task does not need interpretation. A human gate can be the efficient choice when the downside is large and the action is rare.

Reliability makes the metric real

Cost optimization without reliability is simply cost cutting.

If the system becomes cheaper by producing weaker work, hiding failures, or transferring more effort to reviewers, the business did not gain efficiency. The cost moved somewhere less visible.

The same applies to safety. A cheaper system that sometimes exceeds its authority is not efficient because the rare failure can dominate the savings from thousands of ordinary runs.

That is why our broader reliability work separates builders from evaluators and carries decisions through inspectable artifacts. The evaluator should test the real output against a frozen standard, not accept the builder's explanation of why the result is probably fine. We describe that pattern in Never Let an AI Coding Agent Loop Over Its Own Reasoning.

Verification is part of the product. Its cost belongs in the denominator, and the failures it prevents belong in the value calculation.

Every unit of compute should justify itself

There is also a wider engineering principle here.

AI systems consume real compute, hardware capacity, energy, and operator attention. Using more capability than the outcome requires is waste, even when the waste is hidden inside a subscription or absorbed by a platform provider.

The goal is not to build the smartest system the budget can afford.

The goal is to build the least wasteful system that reliably does the job.

That means using deterministic software where uncertainty adds nothing. It means giving models only the permissions and context they need. It means escalating to more capable models when the work earns it. It means counting retries, review, maintenance, and failure instead of celebrating a cheap API call.

Engineering AI systems is not automation for its own sake.

It is delivering reliable business outcomes at the lowest honest total cost.

Kratos Labs builds governed AI systems for organizations moving from pilots into daily operations.

Continue reading