AI software does not have one dominant pricing unit. Products charge by seat, token, credit, minute, task, run, resolution, or combinations of those units. The only fair comparison starts with a defined workload, translates every plan into that workload, and records what happens at the limit.
What this guide covers
- the main AI software pricing models;
- the risk each unit transfers to the buyer or vendor;
- a normalization worksheet for real workloads;
- questions that expose hidden multipliers and plan boundaries.
EXVIV corpus snapshot, August 4, 2026: the public record covered nine AI software markets, 1,035 companies, 1,334 monitored pages, and 219 reviewed events. The markets use materially different billable units, which is why this guide normalizes workloads instead of manufacturing one universal price ranking.
Why the monthly headline is not the comparison
Traditional SaaS often made a seat price the visible center of the offer. AI vendors face variable inference, speech, search, compute, and human-review costs. Their packaging increasingly combines a predictable subscription with a variable meter.
Across the nine AI markets EXVIV tracks, the billable unit often reflects the product's cost structure or claimed value:
- model APIs use input/output tokens, cached input, batches, or compute;
- voice agents use minutes plus speech, model, telephony, and add-ons;
- app builders and sales tools use credits whose consumption varies by action;
- customer-support agents may use seats plus successful resolutions;
- workflow tools use executions, steps, tasks, credits, or compute;
- security and governance platforms often expose a base unit but quote enterprise controls.
Two plans with the same monthly price can produce very different total costs and very different failure modes.
The eight common models
1. Per seat
The buyer pays for each authorized user. It is predictable when team size is stable and use is broadly similar. It becomes awkward when a small number of agents do most of the expensive work or automation reduces the number of human users.
Ask whether viewer, admin, service, and occasional seats are billed differently.
2. Per token or compute unit
The bill follows model input, output, cached context, or compute. OpenAI's model documentation, like other model-provider pages, exposes separate model characteristics and prices rather than one universal request fee.
Token rates are precise but not sufficient. Prompt length, response length, retry rate, tool calls, cache behavior, batch eligibility, and successful task rate determine effective cost.
3. Credit
A vendor-defined credit can simplify a mixed product surface, but the reader must inspect the conversion table. Clay's pricing and its AI pricing documentation distinguish fixed and variable credit behavior for different operations.
The core question is not “How many credits?” It is “How many credits does our median and worst-case workflow consume, and can that conversion change?”
4. Per minute
Voice-agent platforms often show a platform or agent rate per minute. The real stack may also include speech recognition, text-to-speech, model usage, telephony, recording, transfer, and concurrency. Retell's pricing calculator illustrates why model and voice selections matter to the result.
Model connected minutes and completed business outcomes separately. A cheap minute is expensive when calls are long, repeatedly transferred, or unsuccessful.
5. Per task, run, or execution
Workflow and agent products may charge for a top-level run, every internal step, or the AI work inside the run. Branches, retries, loops, polling, and subflows can turn one business task into many billable events.
Draw the workflow and label every metered boundary before estimating cost.
6. Per resolution or outcome
Outcome pricing aligns the bill with claimed value only when the outcome is well-defined and auditable. Intercom's pricing FAQ explains that the plan context and definition of a Fin resolution matter.
Ask how abandonment, partial answers, repeated contacts, escalation, refunds, and disputed outcomes are treated. Also account for human work on interactions that do not resolve.
7. Flat subscription with allowances
A base subscription includes a quantity of requests, credits, minutes, runs, or traces. This model is easy to budget until the allowance or throttle boundary becomes important.
Record whether excess use is blocked, slowed, rolled over, billed, or requires a plan upgrade.
8. Hybrid
Many AI offers combine seats, platform fees, included usage, overage, add-ons, and enterprise commitments. Hybrid is not inherently bad; it can match several kinds of value. It is hard to compare when the buyer collapses the stack into one visible number.
Normalize around a workload
Define a unit of useful work before opening pricing pages.
Examples:
- one accepted coding task;
- one deployed app iteration;
- one million input and output tokens at a known ratio;
- one completed voice call with a known duration and transfer rate;
- one qualified sales record;
- one verified support resolution;
- one production workflow with its retries and branches;
- one million retained observability spans.
Then fill in this worksheet:
| Input | Expected | High case | Notes |
|---|---|---|---|
| Active users/seats | include roles | ||
| Useful tasks per month | define success | ||
| Metered units per task | tokens, credits, steps, minutes | ||
| Retry/failure multiplier | unsuccessful work still costs | ||
| Included allowance | per workspace or seat | ||
| Overage/unit rate | tiered or flat | ||
| Required add-ons | security, retention, support | ||
| Human overlap cost | review, escalation, correction | ||
| Commitment/term | monthly, annual, prepaid |
Calculate three totals: base commitment, expected variable spend, and high-case spend. Also calculate cost per successful unit—not merely cost per attempted unit.
Evaluate who carries the risk
| Model | Buyer gets | Buyer risks | Vendor risks |
|---|---|---|---|
| Seat | budget predictability | paying for unused access | heavy users with high variable cost |
| Token/compute | precise consumption | volatile and hard-to-forecast workflows | price competition and efficiency gains |
| Credit | simplified bundle | opaque or changing conversion | mismatch between credits and cost |
| Minute | intuitive voice volume | layered components and failed calls | expensive calls or providers |
| Outcome | payment tied to value | disputed definitions and quality | nonpayment for expensive attempts |
| Execution/step | workflow visibility | multiplier from branches/retries | complex workloads hidden in one run |
The right model depends on the job, margins, and observability available to both sides. Stripe's pricing and packaging guide recommends selecting a value metric that tracks customer value; AI products add the practical requirement that the metric must also be measurable and understandable.
Questions to ask before comparing vendors
- What exact event creates a billable unit?
- What is included in the base plan?
- Which actions consume different amounts?
- Are retries, failures, tests, and previews billed?
- Does unused allowance roll over?
- What happens at the limit: stop, throttle, overage, or upgrade?
- Which security, retention, support, and governance features require another tier?
- Can the unit definition or credit conversion change during the term?
- Can the customer independently reconcile usage?
- What workload makes the apparent low-cost option cross over?
Use EXVIV Evidence Compare to inspect offers in context, and keep the appropriate market page beside the workload model. A current pricing page is an input; the comparison is the normalized decision model you build from it.
Sources and further reading
- Stripe: SaaS pricing and packaging strategy
- OpenAI API models
- Retell AI pricing
- Intercom pricing FAQ
- Clay pricing
Method note
This is a comparison framework, not a current price list. Vendor units and plans are volatile and may differ by region, workload, or contract. Recheck the linked official pages and model the buyer's actual workload before making a purchasing or pricing decision.