The invoice must follow customer value
An AI agent can send messages, retrieve records, call tools, and consume tokens without producing a result a buyer would pay for. Those are inputs. A billable outcome is a completed, inspectable, accepted business event: a support case resolved without reopening, a document posted with validated fields, or a qualified meeting that occurred.
Outcome-based pricing shifts the question from system effort to work the customer no longer performs. The price is defensible only with a firm result boundary. If the vendor cannot show what happened, who or what accepted it, and when it can be reversed, the model becomes a dispute mechanism with an AI label.
The test: can a finance manager reconcile an invoice line to an event log and reach the vendor's answer? If not, the outcome is not ready to bill.
A billable result needs boundaries
Teams often sell aspirations—better support, faster operations, more pipeline—not invoice units. A chargeable result needs a start condition, completed state, evidence, exclusions, and correction limit. Put these in a short outcome contract before pricing.
| Contract field | Decision to make | Example for a support agent |
|---|---|---|
| Eligible item | Which work enters the meter? | Customer-submitted tickets in selected queues |
| Start event | When does responsibility begin? | Ticket receives an agent-assigned status |
| Accepted outcome | What proves value was delivered? | Issue resolved and ticket closed with required case notes |
| Evidence | Which records settle a dispute? | Ticket ID, transcript, action log, closure code, timestamps |
| Exclusions | Which events never count? | Spam, duplicate tickets, policy-restricted issues, human-only queues |
| Reversal rule | What invalidates the charge? | Same issue reopens within 14 days or fails quality review |
| Attribution window | How long can the result be credited? | From assignment through the 14-day reopen period |
This forces an early choice: is the customer buying a completed task, verified quality state, or downstream business event? The farther downstream the outcome, the more appealing it sounds and the less control the agent has. Billing a sales agent per closed deal, for example, bundles product fit, sales skill, budget cycles, and buyer timing into a result software did not create alone.
Three outcomes that survive audit
The right outcome depends on the workflow. The useful pattern is an event chain reflecting the agent's contribution and giving both sides evidence to inspect.
| Workflow | Billable outcome | Acceptance and reversal rule | Sensible attribution window |
|---|---|---|---|
| Customer support | Resolved eligible ticket | Ticket is closed with a valid resolution code; reverse if the same issue reopens or an auditor finds a policy breach | Assignment to 14 days after closure |
| Document processing | Validated document posted to the target system | Required fields pass validation and the document is posted; reverse for a correction caused by extraction or routing error | Upload to 7 days after posting |
| Sales qualification | Qualified meeting held and accepted in CRM | Meeting occurs, contact meets agreed criteria, and CRM owner accepts the record; reverse for duplicate, no-show, or misrepresented qualification | First agent contact to 30 days after meeting |
Support resolution is often the cleanest start because the item, action history, and reopen event are in one system. Charging for every closed ticket invites premature closure. Require a closure reason, eligible queue, and reopen or quality-failure credit.
Document processing should not charge for a file merely read. Value arrives when usable data reaches a governed destination. A posted invoice missing a tax field may appear complete to the agent but create downstream rework. Define field-level validation, duplicate handling, and ownership of exception queues before billing.
For sales qualification, charge for a verified meeting or accepted opportunity signal, not an email sent or lead scored. The commercial team must agree on fit criteria, disqualifiers, and the CRM acceptance field. A vendor defining qualification alone gets paid for pipeline noise.
Acceptance rules protect both parties
Acceptance should be operational, not a vague promise of later customer quality judgment. Combine automated checks with a narrow human dispute path. Automation determines eligibility and completion; a named customer role can challenge an item with event-log evidence. The vendor credits a confirmed reversal under the published rule.
Set the attribution window to the workflow's natural failure cycle. A support issue may recur quickly; a procurement document may take a week to surface a correction; a sales meeting needs time to confirm attendance and qualification. Too short transfers risk to the buyer; too long delays invoicing and makes collection unpredictable.
State which agent version performed the work. An accepted outcome from version 1.4 must remain traceable after version 1.5 changes prompts, tools, routing, or approval thresholds. Preserve original evidence sufficiently for billing review, while restricting personal-data access and retaining only records needed for commercial and operational purposes.
Choose outcome pricing only when proof is cheap
Per-seat pricing works when access has value and user usage varies little. Usage pricing fits when costs scale with consumption but the customer cannot fairly promise a result, such as a research assistant exploring uncertain questions. Outcome pricing merits its complexity when the result is observable, repeatable, materially influenced by the agent, and cheaper to verify than argue about.
A hybrid model often suits early deployments: a fixed platform fee for integration, governance, reporting, and support, plus a charge per accepted outcome for variable value. This avoids burying fixed delivery cost in a high unit price and lets the buyer budget for a minimum service level while the vendor shares in proven volume.
For charging for activity rather than results, see usage-based pricing economics, metering, and margins. Metering records consumption; outcome pricing must establish that consumption produced an accepted result.
Do not force outcome pricing on ambiguous ownership. If a human reviewer changes most agent output, charge for assisted processing or use staged pricing until handoff is measurable. Calling it outcome pricing does not remove the need to account for human labor.
Model margin before promising price
Price per accepted outcome must cover every workflow-triggered cost, including work never billable. An agent may process an item, call paid tools, and send it to review before failing acceptance. Start unit-margin modeling from eligible attempts, not completed outcomes.
A practical formula is:
Expected contribution = net outcome revenue - model and tool cost - review cost - support cost - reversal remediation - payment cost
This illustrative, non-benchmark invoice assumes an agent receives 1,000 eligible support cases in one billing period, produces 720 accepted resolutions, and has 45 reversed within the 14-day window. The agreed price is $6 per net accepted resolution; model and tool cost is $0.70 per eligible case, human review $1.20 per accepted resolution, support administration $0.25 per eligible case, and reversal remediation $1 per reversed case.
Net billable outcomes = 720 accepted - 45 reversals = 675
Invoice amount = 675 × $6 = $4,050
Model and tools = 1,000 × $0.70 = $700
Review = 720 × $1.20 = $864
Support administration = 1,000 × $0.25 = $250
Reversal remediation = 45 × $1 = $45
Contribution before payment fees and fixed costs = $2,191
Dividing variable cost only by 675 billed items hides that inference and support were incurred for every eligible case. If reversals rise after a model change, the invoice falls while remediation rises; track both. Price floors, volume tiers, and caps on exceptional manual review can protect margin, but must not obscure legitimate customer credits.
Close gaming paths before launch
Every billable definition changes behavior. A support agent can close easy tickets early; a document agent can receive duplicate uploads that inflate volume; a sales agent can schedule low-quality meetings when attendance is rewarded without acceptance criteria. Buyers can game the model by routing work outside agreed queues, delaying disputes, or disputing valid outcomes after the window.
Treat these as design risks, not trust problems. Before launch, test duplicate identifiers, reopened tickets with a new subject line, documents split into partial files, meetings rescheduled after initial acceptance, and customer-admin overrides. Add deduplication rules, reason codes, account-level anomaly review, and a defined escalation owner.
Avoid incentives that punish honest reporting. Pressure on quality reviewers to keep reversals low destroys evidence credibility. Separate agent operators from the person or process adjudicating disputed outcomes. Sampling supports quality control, but cannot replace a deterministic acceptance rule when money changes hands.
Scale turns definitions into operations
At low volume, a founder can inspect disputes manually. At scale, every outcome needs a durable event schema: account ID, work-item ID, eligibility decision, agent version, timestamps, acceptance status, reversal reason, and invoice period. Without these fields, finance rebuilds the meter from exports monthly and product teams cannot diagnose changed outcomes.
Segmentation is also necessary. A support-resolution price may be profitable for password resets but not complex integrations. One blended price can work while routing and case mix are stable. When enterprise accounts, languages, channels, or approval requirements sharply shift cost, create outcome classes or exclude high-touch work until its economics are understood.
Version changes require release discipline. Run a limited cohort, compare acceptance and reversal behavior with the prior version, and preserve rollback. A release increasing completions while increasing reversals has not improved the commercial outcome.
Run a monthly proof cycle
A reliable billing rhythm links product telemetry, customer operations, and finance. Review whether the agent created accepted value, the meter remained trustworthy, and the unit still pays for delivery.
- Reconcile eligible items, accepted items, reversals, credits, and net billable outcomes by account.
- Review acceptance and reversal rates by workflow, agent version, and customer segment.
- Investigate sudden changes in case mix, duplicate rates, human-review load, or dispute volume.
- Share disputed examples with the customer and record rule changes before the next invoice period.
- Recalculate contribution using attempted volume, not only billed volume.
Assign every field an owner. Product owns event definition and release effects; engineering owns instrumentation reliability; operations owns exception handling; finance owns invoice reconciliation; the customer designates the acceptance contact. Without this map, a dashboard is numbers nobody can settle.
A price is a product promise
The strongest outcome price does not claim an AI agent controls the full business result. It identifies the workflow portion the agent can complete, proves completion, and credits failures without drama. Start with one narrow, auditable outcome and a reversal window reflecting real customer risk. After the evidence survives several billing cycles, expand confidently rather than optimistically.