AI software does not have one dominant pricing unit. Products charge by seat, token, credit, minute, task, run, resolution, or combinations of those units. The only fair comparison starts with a defined workload, translates every plan into that workload, and records what happens at the limit.

What this guide covers

  • the main AI software pricing models;
  • the risk each unit transfers to the buyer or vendor;
  • a normalization worksheet for real workloads;
  • questions that expose hidden multipliers and plan boundaries.

EXVIV corpus snapshot, August 4, 2026: the public record covered nine AI software markets, 1,035 companies, 1,334 monitored pages, and 219 reviewed events. The markets use materially different billable units, which is why this guide normalizes workloads instead of manufacturing one universal price ranking.

Why the monthly headline is not the comparison

Traditional SaaS often made a seat price the visible center of the offer. AI vendors face variable inference, speech, search, compute, and human-review costs. Their packaging increasingly combines a predictable subscription with a variable meter.

Across the nine AI markets EXVIV tracks, the billable unit often reflects the product's cost structure or claimed value:

  • model APIs use input/output tokens, cached input, batches, or compute;
  • voice agents use minutes plus speech, model, telephony, and add-ons;
  • app builders and sales tools use credits whose consumption varies by action;
  • customer-support agents may use seats plus successful resolutions;
  • workflow tools use executions, steps, tasks, credits, or compute;
  • security and governance platforms often expose a base unit but quote enterprise controls.

Two plans with the same monthly price can produce very different total costs and very different failure modes.

The eight common models

1. Per seat

The buyer pays for each authorized user. It is predictable when team size is stable and use is broadly similar. It becomes awkward when a small number of agents do most of the expensive work or automation reduces the number of human users.

Ask whether viewer, admin, service, and occasional seats are billed differently.

2. Per token or compute unit

The bill follows model input, output, cached context, or compute. OpenAI's model documentation, like other model-provider pages, exposes separate model characteristics and prices rather than one universal request fee.

Token rates are precise but not sufficient. Prompt length, response length, retry rate, tool calls, cache behavior, batch eligibility, and successful task rate determine effective cost.

3. Credit

A vendor-defined credit can simplify a mixed product surface, but the reader must inspect the conversion table. Clay's pricing and its AI pricing documentation distinguish fixed and variable credit behavior for different operations.

The core question is not “How many credits?” It is “How many credits does our median and worst-case workflow consume, and can that conversion change?”

4. Per minute

Voice-agent platforms often show a platform or agent rate per minute. The real stack may also include speech recognition, text-to-speech, model usage, telephony, recording, transfer, and concurrency. Retell's pricing calculator illustrates why model and voice selections matter to the result.

Model connected minutes and completed business outcomes separately. A cheap minute is expensive when calls are long, repeatedly transferred, or unsuccessful.

5. Per task, run, or execution

Workflow and agent products may charge for a top-level run, every internal step, or the AI work inside the run. Branches, retries, loops, polling, and subflows can turn one business task into many billable events.

Draw the workflow and label every metered boundary before estimating cost.

6. Per resolution or outcome

Outcome pricing aligns the bill with claimed value only when the outcome is well-defined and auditable. Intercom's pricing FAQ explains that the plan context and definition of a Fin resolution matter.

Ask how abandonment, partial answers, repeated contacts, escalation, refunds, and disputed outcomes are treated. Also account for human work on interactions that do not resolve.

7. Flat subscription with allowances

A base subscription includes a quantity of requests, credits, minutes, runs, or traces. This model is easy to budget until the allowance or throttle boundary becomes important.

Record whether excess use is blocked, slowed, rolled over, billed, or requires a plan upgrade.

8. Hybrid

Many AI offers combine seats, platform fees, included usage, overage, add-ons, and enterprise commitments. Hybrid is not inherently bad; it can match several kinds of value. It is hard to compare when the buyer collapses the stack into one visible number.

Normalize around a workload

Define a unit of useful work before opening pricing pages.

Examples:

  • one accepted coding task;
  • one deployed app iteration;
  • one million input and output tokens at a known ratio;
  • one completed voice call with a known duration and transfer rate;
  • one qualified sales record;
  • one verified support resolution;
  • one production workflow with its retries and branches;
  • one million retained observability spans.

Then fill in this worksheet:

InputExpectedHigh caseNotes
Active users/seatsinclude roles
Useful tasks per monthdefine success
Metered units per tasktokens, credits, steps, minutes
Retry/failure multiplierunsuccessful work still costs
Included allowanceper workspace or seat
Overage/unit ratetiered or flat
Required add-onssecurity, retention, support
Human overlap costreview, escalation, correction
Commitment/termmonthly, annual, prepaid

Calculate three totals: base commitment, expected variable spend, and high-case spend. Also calculate cost per successful unit—not merely cost per attempted unit.

Evaluate who carries the risk

ModelBuyer getsBuyer risksVendor risks
Seatbudget predictabilitypaying for unused accessheavy users with high variable cost
Token/computeprecise consumptionvolatile and hard-to-forecast workflowsprice competition and efficiency gains
Creditsimplified bundleopaque or changing conversionmismatch between credits and cost
Minuteintuitive voice volumelayered components and failed callsexpensive calls or providers
Outcomepayment tied to valuedisputed definitions and qualitynonpayment for expensive attempts
Execution/stepworkflow visibilitymultiplier from branches/retriescomplex workloads hidden in one run

The right model depends on the job, margins, and observability available to both sides. Stripe's pricing and packaging guide recommends selecting a value metric that tracks customer value; AI products add the practical requirement that the metric must also be measurable and understandable.

Questions to ask before comparing vendors

  1. What exact event creates a billable unit?
  2. What is included in the base plan?
  3. Which actions consume different amounts?
  4. Are retries, failures, tests, and previews billed?
  5. Does unused allowance roll over?
  6. What happens at the limit: stop, throttle, overage, or upgrade?
  7. Which security, retention, support, and governance features require another tier?
  8. Can the unit definition or credit conversion change during the term?
  9. Can the customer independently reconcile usage?
  10. What workload makes the apparent low-cost option cross over?

Use EXVIV Evidence Compare to inspect offers in context, and keep the appropriate market page beside the workload model. A current pricing page is an input; the comparison is the normalized decision model you build from it.

Sources and further reading

Method note

This is a comparison framework, not a current price list. Vendor units and plans are volatile and may differ by region, workload, or contract. Recheck the linked official pages and model the buyer's actual workload before making a purchasing or pricing decision.