TL;DR
AI pricing models are being rebuilt in 2026 because AI, unlike classical software, burns real compute on every single use: AI product vendors now target median gross margins around 50 percent, well below the 80 percent SaaS is used toGrowth Unhinged: The 2026 State of B2B SaaS and AI Monetization (230 Unternehmen, April/Mai 2026). The AI price ladder sorts the field into five steps, from price per token to cost per solved task to paying for outcomes. List prices won't carry you far up that ladder, since in 32 percent of studied model pairs the cheaper-listed model ends up costing more overallChen et al.: The Price Reversal Phenomenon (arXiv, März 2026). Hybrid AI pricing models jumped from 25 to 37 percent of vendors within a yearGrowth Unhinged: The 2026 State of B2B SaaS and AI Monetization (230 Unternehmen, April/Mai 2026). Outcome pricing takes the payment risk off the buyer's shoulders; whether the result is any good is still mostly measured by the vendor itself. If you're buying, negotiate the counting unit first and the price second.
I maintain a cost calculator that tracks 75 AI models, and it keeps humbling me. Twice in two weeks my price data went stale without a single test firing: one model listed at five times its actual price, another vendor had added a second price tier for long inputs, noted only in a footnote under its price table. Then, on July 30, OpenAI cut the price of a three-week-old model by 80 percentCNBC: OpenAI cuts API prices, 30.07.2026.
Anyone buying or pricing AI is operating in a market where the price tags spin faster than any budget cycle. It looks chaotic, but there's an order to it, and it fits on five steps. I call it the AI price ladder: price per token, tariff mechanics, price per capability, cost per solved task, price per outcome. With every step, payment risk shifts a little further from buyer to vendor.
Every AI answer costs money, which is why every pricing model is wobbling
A sold software license cost its vendor almost nothing to serve. Twenty years of subscription pricing rested on that marginal-cost logic. AI breaks it, because every answer consumes compute on expensive hardware, and that bill lands on the vendor every month.
The survey data now shows the fallout. AI product vendors target median gross margins around 50 percent, and only 12 percent believe the classic 80 percent is achievableGrowth Unhinged: The 2026 State of B2B SaaS and AI Monetization (230 Unternehmen, April/Mai 2026). A second, independently collected sample projects 52 percent for 2026ICONIQ Growth: State of AI 2026, Bi-Annual Snapshot (rund 300 Führungskräfte). On top of that, 70 percent of vendors say their customers pay for AI out of existing software budgetsGrowth Unhinged: The 2026 State of B2B SaaS and AI Monetization (230 Unternehmen, April/Mai 2026). The money itself is real enough: enterprises roughly tripled their generative AI spending to an estimated 37 billion dollars in 2025, according to Menlo VenturesMenlo Ventures: 2025 The State of Generative AI in the Enterprise (November 2025). The room for pricing mistakes is small, yet the industry is openly searching: 76 percent of software vendors have launched AI features, and most report revenue impact below 10 percent so farSimon-Kucher: Global Software Study 2025.
One concrete case will carry us up the ladder: a company with 10,000 customer inquiries a month that it wants AI to answer. Every step of the AI price ladder can be priced against that single workload.
What does a token really cost?
The bottom step sounds reassuringly simple. Language models bill in tokens, small chunks of text, with list prices per million. Our 10,000 inquiries at roughly 3,700 tokens each cost about 37 dollars a month on a budget model, a figure straight from Anthropic's own pricing documentationAnthropic: Preisdokumentation (Cache, Batch, Referenzrechnung; Abruf 10.08.2026). So far, any intern can do this math.
The tariff behind it, step two of the ladder, has more moving parts. Recurring inputs can be cached; reading from that cache costs a tenth of the input price at Anthropic, but writing to it carries a surcharge of 25 to 100 percentAnthropic: Preisdokumentation (Cache, Batch, Referenzrechnung; Abruf 10.08.2026). At DeepSeek, a cache hit costs around one percentDeepSeek: API-Preisdokumentation (Cache-Treffer; Abruf 10.08.2026). If you can wait, batch processing halves the billAnthropic: Preisdokumentation (Cache, Batch, Referenzrechnung; Abruf 10.08.2026). For long inputs, several vendors add a second price tier, often noted only in a footnote under the price table: Google doubles the rate on Gemini 2.5 Pro above 200,000 tokens, retroactively for every token in the callGoogle: Gemini-API-Preisseite (Kontextstaffel ab 200.000 Token; Abruf 10.08.2026), and OpenAI raises rates on several models above 272,000 tokensOpenAI: API-Preisseite (Kontextstaffeln ab 272.000 Token; Abruf 10.08.2026). Then there's thinking: modern models generate invisible reasoning tokens that show up on the bill like any others, and whose volume swings by up to 9.7 times on the identical taskChen et al.: The Price Reversal Phenomenon (arXiv, März 2026). The output side of your bill has become, quite literally, incalculable.

The third step of the ladder is price per capability, and it carries a price dynamic that runs in both directions at once. The price of a fixed capability level falls 5 to 10 times per yearGundlach et al.: The Price of Progress (MIT, arXiv, März 2026), and measured at specific performance thresholds, the median decline is 50 times per yearEpoch AI: LLM inference prices have fallen rapidly but unequally across tasks (Abruf 08/2026). Meanwhile the price of the frontier, the best available model, rises 3 to 18 times per yearGundlach et al.: The Price of Progress (MIT, arXiv, März 2026). That's why companies experience both at once: yesterday's task keeps getting cheaper while the newest ambition keeps getting pricier.

The honest unit is the solved task
On the fourth step, the question changes: what does it cost to get a task done? Research has had a plain formula for it since 2025: cost per solved task equals price per attempt divided by success rateErol et al.: Cost-of-Pass, An Economic Framework for Evaluating Language Models (arXiv). A model that costs half as much but succeeds a third as often is the expensive one on this step.
A study published in March quantified how often list prices mislead: in 32 percent of model pairs, the cheaper-listed model causes higher total cost, by up to 28 times in the extremeChen et al.: The Price Reversal Phenomenon (arXiv, März 2026). One model listed 80 percent below its competitor came out 38 percent more expensive across all tasksChen et al.: The Price Reversal Phenomenon (arXiv, März 2026). The study attributes this to those swinging reasoning tokens plus extra working stepsChen et al.: The Price Reversal Phenomenon (arXiv, März 2026). The list price fails as a price signal exactly where it promises the most.

For our 10,000 inquiries, the 37 dollars from earlier are only a good number if the budget model actually resolves them. If every fifth answer is unusable and a human cleans up, the real bill lives somewhere else. There's even a number for when the strongest model wins despite its price: once a single error costs around 10 cents, the frontier model almost always comes out ahead, per a Caltech analysis at mid-2025 pricesZellinger, Thomson: Economic Evaluation of LLMs (Caltech, arXiv, Juli 2025). That sounds low, but it follows from the proportions: what the frontier model costs extra per single request is small against what one avoided error saves. With customer inquiries, one lost customer clears that bar easily.
This step is why I built my cost calculator around exactly this unit, and I took the cost levers behind it apart in cutting AI costs through model choice.
Which AI pricing models are winning in 2026?
Above the buying side, the vendor side of the AI price ladder begins, and it's been on the move for two years. In 2024, most software vendors priced AI features the way they'd always priced software, as a per-seat add-on. Then the margin math arrived, and the market has been testing new counting units in public ever since.
The American developer-tools market shows the migration in fast motion. GitHub moved Copilot to usage-based billing on June 1, 2026, replacing flat request allowances with token-based credits while keeping plan prices unchangedGitHub Blog: Copilot is moving to usage-based billing (Umstellung zum 01.06.2026). Cursor's June 2025 switch from request counts to a monthly compute budget went badly enough that the company publicly apologized and refunded surprise chargesCursor Blog: June 2025 pricing (Umstellung, Entschuldigung, Rückerstattungen). Salesforce runs two consumption models for its service agent side by side, 2 dollars per conversation or credits at 10 cents per standard action, and buyers have to do the math on which one is cheaperSalesforce-Hilfe: Agentforce-Preismodelle (Conversations und Flex Credits).

The fifth rung, the top of the ladder, is already on the rate card. Intercom charges 99 cents per outcome, counted as a confirmed resolution or a completed workflowIntercom: Preisseite Fin (0,99 USD je Outcome; Abruf 10.08.2026). HubSpot cut its price in April from 1 dollar per conversation to 50 cents per resolved conversationHubSpot Company News: Now you pay when the task is complete (14.04.2026). Zendesk goes furthest: a second AI model verifies which resolutions actually count before they're billedZendesk Newsroom: Relate 2026, verifizierte Resolutions (19.05.2026). In Europe, SAP built a corporate answer, a prepaid currency called AI Units that AI features draw on across its whole product line; one Joule consultant package bundles 22,900 requests per user and month, pooled across all usersSAP Learning: Evaluating the Commercial Model (Joule, AI Units). For our 10,000 inquiries, two worlds now coexist: a two-digit dollar figure as a token bill, or several thousand dollars as an outcome bill once you scale 50 to 99 cents per resolved case (own calculation). The spread between them is the vendor's margin, risk and verification cost. The ladder sorts the logic; in the market it condenses into four basic types of pricing models.
| Criterion | Per seat | Usage (tokens, credits) | Per outcome | Hybrid (base plus component) |
|---|---|---|---|---|
| Counting unit | Workplace | Tokens, actions, credits | Resolved case | Base fee plus usage or outcome |
| Cost risk sits with | Vendor | Buyer | Vendor | Shared |
| Predictability for the buyer | High | Low | Medium | Medium to high |
| Typical dispute | Unused licenses | Usage spikes | Who verifies the outcome? | Contract complexity |
| 2026 example | Microsoft 365 Copilot | GitHub Copilot, SAP AI Units, Salesforce Agentforce | Intercom, HubSpot, Zendesk | Most common type: 37% of B2B SaaS vendors |
The adoption numbers match the story. Among roughly 300 surveyed AI companies, 58 percent bill by subscription, 35 percent by usage and 18 percent by outcome, with multiple answers allowedICONIQ Growth: State of AI 2026, Bi-Annual Snapshot (rund 300 Führungskräfte). Hybrid pricing jumped from 25 to 37 percent of 230 B2B software vendors within twelve months, making it the most common modelGrowth Unhinged: The 2026 State of B2B SaaS and AI Monetization (230 Unternehmen, April/Mai 2026). German-speaking Europe shows the same break: classic SaaS there sells 92 percent by subscription, but for AI products the share of usage-based pricing leaps to 69 percenthy × OMR Reviews: SaaS & AI Pricing Report 2026 (4.000+ Profile, 180 Befragte). And 45 percent of software vendors already run two or more pricing models in parallelSimon-Kucher: Global Software Study 2025. Behind all of it sits an old pricing rule with new teeth: put the price on the unit that grows with customer value.

I traced what this does to software contracts, including the double-payment trap of licenses plus credits, in SaaS pricing in the agent era, and the professional-services side in consulting fees in the AI era.
Payment risk moves, the information problem stays
Only resolved cases get billed, and that does move the payment risk to the vendor. What stays with the buyer is the information problem: knowing what that outcome is actually worth. The research of the past twelve months explains why that weighs so heavily.
Start with a question that sounds paranoid: is the model you pay for actually the model that answers? With software alone, buyers can't reliably tell. Statistical tests on outputs fail against subtle substitutions, log-probability checks are defeated by inference nondeterminism, and dependable proof requires trusted hardware or cryptographic attestation, per current researchCai et al.: Auditing Model Substitution in LLM APIs (arXiv).
That nondeterminism deserves a number. Send the same model the same prompt a thousand times at temperature zero, with randomness switched off, and today's production setups return dozens of different answers; one documented run counted 80 distinct completions out of 1,000Thinking Machines Lab: Defeating Nondeterminism in LLM Inference (Lab-Report, 09/2025). Server load from other customers changes the batch composition, which changes the arithmetic, which changes the outputThinking Machines Lab: Defeating Nondeterminism in LLM Inference (Lab-Report, 09/2025). Serving configuration alone, batch size, GPU type and count, can shift an open model's measured accuracy by up to 9 percentage pointsYuan et al.: Numerical Sources of Nondeterminism in LLM Inference (arXiv). It's fixable in principle, at a cost in serving efficiencyThinking Machines Lab: Defeating Nondeterminism in LLM Inference (Lab-Report, 09/2025), yet it's the normal mode you're buying today. Behavior also drifts over time: the same tasks GPT-4 solved at 84 percent in March 2023 dropped to 51 percent by JuneChen, Zaharia, Zou: How is ChatGPT's behavior changing over time? (arXiv/HDSR), and software engineering research now treats silent provider updates as a governance problem in its own rightChishti et al.: Governing Updates in the LLM Supply Chain (LLMSC @ FSE 2026).
On the incentive side, economists have formally modeled the temptation to serve the cheapest just-good-enough model behind a fixed outcome price as moral hazardSaig et al.: Incentivizing Quality Text Generation via Statistical Contracts (NeurIPS 2024); contract theory actually designed outcome-based contracts as the remedy for that temptation under token pricingSaig et al.: Incentivizing Quality Text Generation via Statistical Contracts (NeurIPS 2024), and there's no field evidence of providers secretly downgrading. What is peer-reviewed: when one AI judges another AI's work, the judgment is systematically biased, up to favoring related modelsLi et al.: Preference Leakage, Contamination in LLM-as-a-judge (ICLR 2026). Economically, an AI service is a credence good, like a car repair, where the seller knows more about the adequate service than the buyer, and experiments show that asymmetry works against buyersErlei: Generative AI in Credence Goods Markets (arXiv).
For our 10,000 inquiries, the several thousand dollars on an outcome bill are a fair trade only if the measurement holds. Who verifies what resolved means, which model does the work, and how quality holds up over the term, that's a contract question.
Who are you actually signing with?
An AI vendor is a useless category at the negotiating table, because three very different counterparties hide behind it. Large enterprises negotiate framework agreements directly with the model providers; smaller companies buy the same technology off the shelf. Anthropic sells its Enterprise plan self-serve from 20 seats with usage billed on top at API ratesAnthropic: Enterprise-Plan-Dokumentation (Self-Serve ab 20 Plätzen), and Microsoft built a dedicated Copilot tier for businesses below 300 seatsMicrosoft 365 Blog: Copilot Business für Unternehmen unter 300 Nutzern (12/2025). In between sit two product worlds: established software that bolts AI onto its existing price model, and AI-native products that wouldn't exist without it. Our toolbox carries exactly this distinction as a dimension of every tool profile ("Established", "Established, AI retrofitted", "AI-native"), because it decides which contract questions matter.
Signing with a model provider directly, the number that matters most is the model's lifespan. OpenAI commits to at least six months' deprecation notice for mature models, Anthropic to at least 60 days, and Microsoft's Azure catalog carries the sentence worth memorizing: "Retirement dates aren't extendable"OpenAI: Deprecations-Dokumentation (Abkündigungsfristen)Anthropic: Model-Deprecations-DokumentationMicrosoft Learn: Azure OpenAI model retirements. Prices and terms can be adjusted unilaterally, material changes with roughly 30 days' notice at OpenAI and Anthropic alikeOpenAI: Dienstvereinbarung für Geschäftskunden (gültig ab 01.01.2026, Änderungs- und Kündigungsfristen)Anthropic: Commercial Terms of Service (Preisänderungsfrist); if an update materially reduces functionality, OpenAI grants a termination right that has to be exercised within five business daysOpenAI: Dienstvereinbarung für Geschäftskunden (gültig ab 01.01.2026, Änderungs- und Kündigungsfristen). Dated snapshots can be pinned, but pinning fixes behavior; it doesn't keep the model aliveAnthropic: Model-Deprecations-Dokumentation.
With retrofitted incumbent software, three payment layers stack: the base license, an AI surcharge per user, and consumption credits. The credits expire, at HubSpot verbatim at the end of each usage period, at Salesforce at the end of the subscription term, at SAP after twelve monthsHubSpot Knowledge Base: Credits und Abrechnung (Verfallsregel)Salesforce-Hilfe: Agentforce-Preismodelle (Conversations und Flex Credits)SAP Learning: Evaluating the Commercial Model (Joule, AI Units). And nobody promises which model powers the feature: Salesforce describes its default as a "managed mix of trusted models (currently including GPT-4o)"Salesforce Developer-Dokumentation: Supported models (managed mix), and Microsoft has been orchestrating Anthropic models inside Copilot since fall 2025, hosted outside its own termsMicrosoft 365 Blog: Expanding model choice in Microsoft 365 Copilot (09/2025). The contract is hiding in the word "currently".
AI-native products are the most fluid group, in price and in plumbing. Replit prices per unit of effort since 2025, from six cents to several dollars for the same checkpointReplit Blog: Effort-Based Pricing; Cursor's terms allow changing or discontinuing features "at any time … without notice"Cursor: Terms of Service (Änderungsvorbehalt). And the model behind the product can flip overnight: Windsurf lost most of its access to its supplier's Claude models in June 2025 with less than five days' warningTechCrunch: Anthropic kappt Windsurfs Claude-Zugang (06/2025). Buying an AI-native product means buying a bet on its supply chain.
| Criterion | Model provider directly | Established, AI retrofitted | AI-native |
|---|---|---|---|
| Typical purchase | Enterprise framework agreement or self-serve API | Add-on to the existing license plus credits | Subscription or credits, sometimes per outcome |
| What can change | Model retirement (6 months to 60 days notice), prices and terms (roughly 30 days) | Model behind the feature ("managed mix"), credit rules | The pricing model itself, the model supply chain, features |
| Credit expiry | None (pay-as-you-go) | Monthly to end of term | Varies by product, often variable per task |
| Key contract question | Model lifespan and snapshot path | Total cost per case across all three layers | Grandfathering on price and model changes |
Three out of four requests don't need a frontier model
One lever runs across every step of the AI price ladder, and it only becomes visible once companies combine models. The average enterprise already uses 3.1 model providersICONIQ Growth: State of AI 2026, Bi-Annual Snapshot (rund 300 Führungskräfte), mostly without a system. Yet the sorting logic works like a mailroom: nobody sends every letter by express courier, and of our 10,000 inquiries, most are routine and few are delicate.
Berkeley's RouteLLM work measured what that sorting is worth: a trained router sent only about one in five requests to the expensive model, versus every second one under random assignment, while holding 95 percent of the top model's qualityOng et al.: RouteLLM, Learning to Route LLMs with Preference Data (ICLR 2025). In a production workload, a calibrated router cut raw model costs by 31 percent at stable qualityUCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing (arXiv, Mai 2026). As an order of magnitude: a third to two thirds of cost, always relative to a declared quality bar.

Two caveats keep this honest. First, many off-the-shelf routers, commercial ones included, fail to beat a simple baseline, as a benchmark across 400,000 test cases showedLLMRouterBench (arXiv, Januar 2026). Second, the error-cost threshold from the previous step applies here too: once a mistake costs real money, the single frontier model beats any combinationZellinger, Thomson: Economic Evaluation of LLMs (Caltech, arXiv, Juli 2025). And for agents that work autonomously across multiple steps, model choice already happens per step, which turns combination into an architecture question.
The example can now be priced end to end, every figure an own calculation based on the list prices cited, as of August 10, 2026:
| Step | Calculation | Result |
|---|---|---|
| 1 · Token price | 10,000 inquiries × roughly 3,700 tokens on a budget model [6] | ~$37 per month |
| 2 · Tariff mechanics | Cache reads at 0.1× and batch at 0.5× lower it, context tiers and reasoning tokens push it up [6][8][9] | roughly $20 to $75, depending on setup |
| 3 · Price per capability | the same capability class costs 5 to 10 times less a year later [10] | the budget holds, ambitions grow into it |
| 4 · Cost per solved task | $37 ÷ 0.8 success rate, a human reworks the rest [12] | ~$46 plus rework time |
| 5 · Price per outcome | 8,000 resolved cases × $0.50 to $0.99 at a vendor [17][18] | $4,000 to $7,920 per month |
How do you bring order into your AI pricing?
- Date every price snapshot. Write the retrieval date of vendor price pages onto every AI calculation, and treat any undated number as stale. My own calculator went stale twice in two weeks despite daily use, so this catches people who do it for a living.
- Measure cost per solved task for one week. Run a real workload on a cheap and a strong model in parallel and count how many results are usable without rework. The AI cost calculator does the math behind it, including retries, caching and human cleanup.
- Negotiate the counting unit before the price. As a buyer, ask what exactly the vendor counts, who verifies a resolved outcome, and what happens above the quota. As a vendor, pick the unit that grows with your customer's value, and enter hybrid, base fee plus variable component; the hard cut can come later.
- Define your error-cost threshold. Put a number on what a wrong result costs per task type. Above the threshold, the best available model does the work; below it, the cheapest model that holds your quality bar.
One word on the switch itself, because it's communication work. Cursor showed in 2025 how a defensible price change turns into reputation damage when customers first notice it on the invoiceCursor Blog: June 2025 pricing (Umstellung, Entschuldigung, Rückerstattungen). If you change your AI pricing model, announce it, publish worked examples, and give existing customers time.
My Take
The number I keep coming back to is the 37 percent of AI companies planning to change their pricing model within twelve monthsICONIQ Growth: State of AI 2026, Bi-Annual Snapshot (rund 300 Führungskräfte). That's a market admitting, in public, that every current price tag is a draft. If I were signing an AI contract this quarter, I'd treat it that way: short term, defined exit, an adjustment clause both sides can live with.
The next battleground I see is the meter itself. Once vendors charge per verified outcome, whoever controls the verification controls the revenue, which is why a second model counting resolutions strikes me as more consequential than any single price cut. I'd expect buyers to demand audit rights on that meter the way they once demanded audit rights on license counts. Maybe that takes longer than I think, procurement habits are slow to move.
What I can report first-hand: in late July I benchmarked five popular AI cost calculators (own survey), and none of them could show cost per solved task. Everyone counts tokens. The market argues about the top rung of the AI price ladder with tools that still stand on the bottom one, and whether we'll even talk about tokens in five years... I'll leave that open.
