AI market · 12 Aug 2026 · 7 min read

Model prices have fallen; the cost of dependable automation has not

Cheaper tokens have not produced cheaper automation. Consumption rises as the unit price falls, and most of what an accepted task costs sits outside the model bill entirely.

The procurement paper

Ninety percent off the model bill and a longer exception queue

A procurement paper recommends swapping the model inside a document-processing service for a cheaper alternative. The inference invoice duly falls. At the next operating review the picture is less flattering. The exception queue is longer, quality sampling has been widened to cover it, and two engineers have spent a month rebuilding a retrieval feature the previous provider supplied as standard. Buyers of language technology will recognise the shape of this.

Nothing in the paper was arithmetically wrong. The unit was wrong. A model supplies tokens; a business pays for work that somebody is willing to accept. Everything between those two events is where the saving went.

The contention

Suppliers should be compared on the fully loaded cost of one task accepted into operations. Open weights are worth buying where they resolve a constraint that a managed service cannot.

01

Token prices collapsed and AI bills grew anyway

The fall in the price of machine intelligence is one of the few facts in this market nobody disputes. Stanford records the cost of querying a model above a GPT-3.5-level MMLU threshold falling from about $20.00 per million tokens in November 2022 to $0.07 by October 2024. That threshold is historical and says nothing about whether either model would hold up in a live workflow. As a statement about the price of a unit of capability, it stands.Sources: Stanford HAI, AI Index 2025

Evidence

Querying at a fixed capability threshold became far cheaper

Each point is the lowest-priced model exceeding a GPT-3.5-level MMLU threshold in that period, per million tokens.

Established. The price of querying a model above one historical benchmark threshold fell by more than two orders of magnitude between November 2022 and October 2024.

Not established. A shared MMLU threshold does not imply equivalent quality, reliability or total cost on a business workflow.

Management action. Spend the saving on broader benchmarking. Hold the acceptance standard where it is.

Source: Stanford HAI, AI Index 2025

Budgets did not follow prices down. Reporting through 2026 has traced the opposite movement, with enterprise AI spending climbing while the price of a token keeps falling, and economists reaching for Jevons paradox to account for it. A cheaper input is an invitation to consume more of it. Cheaper tokens made whole categories of work worth attempting that had never justified the meter before.Sources: Fortune, on cheaper tokens and rising AI spending, June 2026

Consumption is where the compounding happens. Retrieval widened the average prompt. Reasoning models generate long internal traces the buyer pays for and never reads. Agentic patterns loop, retry and call tools, so a request that cost a fraction of a penny as a single completion costs considerably more once it is rebuilt as a sequence of steps. Independent evaluation has already adjusted to this: Artificial Analysis prices a model on the total tokens it consumes across a task, reasoning tokens included, divided by the number of tasks.Sources: Artificial Analysis, language model benchmarking methodology

So the headline price per million tokens has stopped being a decision variable. It sets a floor under a bill whose size is settled elsewhere: by how much work the organisation sends, how many attempts each piece of work takes, and how much of the output survives review. A procurement process that optimises the floor is optimising the part of the bill it controls least.

If the token price sets only the floor, the comparison that decides a supplier has to be assembled from the lines above it.

02

Cost per accepted task exposes false bargains

One task accepted into operations costs more than an inference call. Add retrieval and orchestration. Add the attempts that failed and had to be retried. Add the share of work a person finished by hand, and the review effort spent checking the share a person did not. Add evaluation, monitoring, and the engineering that falls due again every time a prompt, a provider or a model version changes.

Those lines do not move together. Inference scales with volume. Review scales with the error rate multiplied by the consequence of an error, which is why a small quality regression in a regulated workflow is expensive and the same regression in an internal draft is tolerable. Integration and control are close to fixed, and they are incurred before a single task has been accepted. A comparison that holds only the first of these constant has compared very little.

Assembling the full figure requires a harness the buyer controls: a set of representative cases including the difficult and the adversarial ones, a rubric that does not move between candidates, and a cost model that attaches review and exception effort to each candidate’s output. The comparison is then between complete services at a fixed acceptance standard. It can be rerun in the week a new model appears, which is the only way a sourcing decision survives a market moving at this speed.

Implementation pattern

A portable evaluation harness protects model choice

Control flow
  1. 01cases = dataset.load("representative")
  2. 02for model in candidates:
  3. 03 outputs = model.run(cases)
  4. 04 score = evaluate(outputs, rubric)
  5. 05select(score.quality, score.totalCost)
Components
  1. 01Test dataset
  2. 02Model adapters
  3. 03Common rubric
  4. 04Cost model
  5. 05Release decision

Run that way, the comparison strips out false bargains quickly. A model at half the price that doubles the rejection rate has raised the cost of an accepted task. A very accurate model can be uneconomic too, if its rare failures oblige every output to pass an expensive specialist. Customer support has already reached this conclusion commercially. Several vendors now charge per resolution and bill nothing for conversations the system fails to resolve, and their published arithmetic shows a platform that resolves a smaller share of contacts at a lower unit price losing to a dearer one once the human-handled remainder is priced at a human cost.Sources: Lorikeet, AI support resolution rates and unit economics, 2026 (vendor material)

Open-weight models deserve the same test and now survive it more often. Stanford tracked the performance gap between the best closed and the best open model on the Chatbot Arena leaderboard narrowing from 8.0 percent in January 2024 to roughly 4.2 percent by the middle of that year and 1.7 percent by February 2025. Proximity on a leaderboard is no guarantee of proximity on a particular document set, and any candidate still has to be benchmarked on the actual work. Automatic dependence on a single frontier supplier now needs a reason.Sources: Stanford HAI, AI Index 2025

Evidence

Open weights closed most of the leaderboard gap in one year

Reported difference between the best closed and the best open-weight model on the Chatbot Arena leaderboard, at three points.

Established. On one public leaderboard, the best open-weight model moved close to the best closed model between January 2024 and February 2025.

Not established. Leaderboard proximity says nothing about performance on a specific document set, and nothing about the cost of operating an open-weight service.

Management action. Treat open weights as a credible candidate and benchmark them on the real work at a fixed acceptance standard.

Source: Stanford HAI, AI Index 2025

That reason cannot be the invoice alone. Self-hosting removes a usage fee and adds licensing, provenance, patching, capacity planning, security monitoring and service continuity, all of which land on the buyer or its infrastructure partner. It also exposes utilisation. An accelerator held against bursty traffic sits idle most of the time, and an idle accelerator is an expensive way to serve a token. NIST treats generative AI risk as a lifecycle matter for a related reason, since a release that passed on the day it shipped degrades as models, data and usage patterns change. A data boundary, a latency requirement or a resilience obligation is a good reason to run a model locally. Avoiding a metered fee, alone, will not survive the first quarter of operating one.Sources: NIST, Generative AI Profile

Since the burden of running a model transfers to whoever runs it, the case for staying with a managed supplier deserves to be put at its strongest.

03

Paying a premium supplier can lower total cost

A demand for portability has a price, and it is usually paid in engineering. Managed platforms arrive with retrieval, evaluation tooling, safety filters, observability, identity integration and a support contract that answers at three in the morning. A buyer who insists on an abstraction layer thin enough to swap providers at will gives up most of that and funds the replacement internally. Where the bundled services materially improve the cost of an accepted task, the resulting dependency is a commercial judgement that pays for itself.

Operating capacity is the harder constraint. Gartner has forecast that more than 40 percent of agentic AI projects will be cancelled by the end of 2027, attributing the cancellations to escalating costs, unclear business value and inadequate risk controls. Those are the failure modes of organisations that acquired more system than they could run. A firm without the people to operate a model service will find that the open weights it downloaded were the cheapest part of the undertaking.Sources: Gartner, over 40 percent of agentic AI projects will be cancelled by end of 2027, June 2025

The concession has a boundary. The buyer must keep the evaluation set, the rubric, the acceptance thresholds and the business rules, because those are what make a second supplier assessable at all. Execution can sit inside proprietary features. The cost of leaving should be calculated before signing, written down and revisited at renewal, so that the dependency remains a decision the firm has taken at a known price.

Because the managed route and the self-hosted route can each be defended independently, the choice has to be settled by a measure that applies to both.

04

Set one denominator before the next model switch

Procurement, engineering and the manager answerable for the workflow should agree a single denominator before any of them looks at a price list. Tasks accepted into operations, at the required standard of quality and consequence. Every candidate is then quoted in one currency, and arguments about model families become arguments about numbers.

With the denominator in place the sourcing rules are short. Reserve expensive capability for the tasks where it changes whether work is accepted. Use a smaller or open-weight model wherever the harness shows it clears the same bar. Deploy locally where a constraint requires it and the firm can staff the consequences. Repeat the comparison on a schedule, because the price and the capability of every candidate will have moved before the contract ends.

This price war is real, and it is being fought over the smallest line in the bill. Taken as a discount, the saving reappears somewhere else on the same page. Spent on evaluation, on connecting the workflow properly and on staffing the exceptions, it buys the dependable automation the lower price was supposed to make affordable.

What follows

Decisions arising from the analysis

  1. Choose one high-volume, reviewable task
  2. Create normal, difficult and adversarial examples
  3. Compare at least two model families
  4. Report cost per accepted output and the causes of rejection

Written by Quiet Gears. If your operating data contradicts the argument above, that is worth more than a defence of the piece.

Send a challenge