The draining side of running a business

The promise and limits of outcome-based pricing

Ernesto Spruyt
28 September 2026 · 7 min read
260928 Hero - The promise and limits of outcome-based pricing

The promise of outcome-based pricing is that you stop paying for effort. You only pay when it works: per resolution, per completed process, per merged pull request.

Running a company, my first thought is that I like this. My second is whether it works that way in practice, and whether we should be offering something like it at Tunga. So I spent a week reading how it works in the places where it already runs.

What I came away with: the rate is the easiest part of an outcome-based deal to compare, and the least important of the things that decide what you end up paying.

Where it already runs

AI customer service agents have been priced this way for a while now, and most of the rates are public:

  • Intercom's Fin agent: \$0.99 per outcome
  • Zendesk: \$1.50 on an annual commitment, \$2.00 pay as you go
  • HubSpot's Breeze agent: \$0.50, down from \$1.00 this spring
  • Gorgias, the support platform many Shopify stores run on: \$1.00 per automated resolution

Then you read what counts as a billable outcome.

In Intercom's own description, a billable outcome includes any conversation where the customer does not ask for further help after the agent's last answer. Configured handoffs count too. So you are billed the same whether the customer got a good answer or gave up and emailed you instead.

260928 Echo 1 - The definition is the invoice

At Gorgias an automated resolution also uses up a ticket in your plan, so the same conversation is billed twice. Unless a human steps in within 72 hours, in which case it counts as a ticket only.

In May 2026 Zendesk changed its model so that only verified resolutions are billed. Somebody there decided the gap between "resolved" and "the customer stopped typing" was worth giving up revenue over.

Comparing the rates is easy. Working out how each supplier's definition will affect your actual bill is not.

What actually gets negotiated

Two things decide an outcome-based bill: the rate, and the definition of the event you are paying for. My guess is that most negotiations are mainly about the first.

A rate is a number you can put in a spreadsheet next to three competitors. A definition takes a conversation with the people who will use the thing every day, and they are usually not in the room when the contract is signed. So it goes through as written.

Whether it is open to negotiation at all depends on who you are buying from. A large platform hands you both fixed and you take it or leave it. A boutique supplier will usually open both. And if a supplier will not discuss the definition at all, you have found out what kind of relationship this is going to be while you can still walk away.

So the first thing to establish is whether the definition is negotiable, because that tells you what kind of deal you are in.

But a definition you did negotiate can still work against you, depending on what it counts.

What the unit is actually counting

The pattern here is older than AI. Per seat pays for headcount: more people on the account, more revenue, whether they use it or not. Per hour and per ticket pay for volume, including the volume that should never have existed.

Per outcome is meant to fix both, and it does when the outcome is something made. Pay a marketing agency per lead that converts and both sides want the same number to grow.

A lot of what gets sold as outcome pricing counts trouble handled instead: a resolved ticket, a replaced developer, a closed incident. That is only a problem when the supplier is in a position to influence how often the trouble happens. An incident response firm cleaning up outages you caused has no say in how many there are, and paying them per incident is fine. A support vendor whose own product quality drives your ticket volume does have a say.

Alex Turnbull, who founded the support software company Groove and sold it, says per ticket rewards the volume you should have prevented, and per resolution punishes you for having a good product. He now runs Helply and prices it with neither seats nor resolutions, just credits.

So work out whether your supplier can influence how often the billable event happens. If they can, and they are paid per event, your interests point in opposite directions, and nobody has to act in bad faith for that to cost you.

260928 Echo 2 - Every pricing unit rewards something

All of that is easier to act on when you can see the count for yourself.

Whether you can count it yourself

On 16 September Sourcegraph started charging for its coding agent per merged pull request. You pay when the PR goes in, not per token.

The interesting part is where the counting happens. A merge is in your own repository, under your own review process, and you would notice if the number looked wrong. A resolution happens in the supplier's system, by the supplier's rule, and reaches you as a total.

If the count happens on your side, you can check the bill, notice a definition that is drifting, and argue from your own numbers rather than theirs. That is worth more than a better rate.

So I put the same questions to my own industry.

What staffing does with this

Staffing wants this too and mostly cannot do it. In a September 2026 study by HFS Research and Capgemini among 101 financial services firms, 67% had introduced business outcome terms into their AI contracts, and 82% were still paying per FTE or per licence, because outcomes are hard to baseline, measure and attribute.

There is one outcome here that can be measured cleanly: whether the match works. Most suppliers put a guarantee on it. If the person does not work out, you get a replacement or you are not billed. Toptal's window runs up to two weeks. On the recruitment side of the industry, 90 days has been the standard for years.

It is worth knowing which model sits underneath, because that decides whose interest the guarantee serves. Where the supplier is paid once for making the placement, the guarantee is a refund window on money already earned. Where the supplier bills for as long as the person works, which is what Toptal does and what we do, a replacement is already in their own interest, because no replacement means no revenue.

Then there is what our own numbers say. Across four years of terminations, not one ending caused by performance or fit surfaced within two weeks. About a fifth surfaced within three months. The median was 151 days.

So whether the window is two weeks or 90 days, what it guarantees is the stretch in which almost nothing goes visibly wrong.

260928 In-article - What a two-week guarantee covers

What we do about it at Tunga

For our clients, long-term continuity is the thing. Talent leaving early is one of their bigger worries, which is why, long before outcome-based anything was trending, we made a point of making sure that the team stays staffed.

It works. We replace, with no limit on how long into an engagement that holds. And it is rarely needed: of the placements we ran in 2025, roughly 1 in 56 ended because the developer left us for another job. So far in 2026, none have.

The harder number is the one where the match itself was not working and the client did not want anyone new. That is about 1 in 40 placements. Although in 80% of those cases the note in our log also records a budget cut, a project winding down or a restructure, so on the strict reading, where nothing but the match went wrong, it is closer to 1 in 200. Which of those two you believe depends entirely on how you define it, and that is rather the point of everything above. We would prefer both numbers to be zero, but hey, we're human.

260928 Echo 3 - How often continuity actually breaks

What we have not worked out is how to turn any of it into an outcome-based way of earning. Like the rest of our industry, we still charge by the hour.

Work with us

Stop searching. Start building.

Tell us what you're looking for. We'll come back within one working day.