AI Economics

Why the Cheapest AI Model Per Token Can Cost More Per Finished Task

Bill Cava/

In a benchmark of 2,400 AI agent runs published this summer, Google's Gemini 3.5 Flash cost about 29 cents per attempt, the cheapest of the four models compared here. It was also the most expensive per finished task, at $1.23, because it passed only 23 percent of the time.[1]

This post is about why the sticker price of a model and the price of getting work done can point in opposite directions, and how to measure the second one.

What is cost per successful task?

Cost per successful task is everything a model spent across every attempt, including the runs that failed, retried or timed out, divided by the number of tasks it actually completed. Price per token and price per attempt cannot see failures. Cost per successful task counts them, so it can rank the same models in a different order.

The benchmark behind that definition comes from Arize, an AI monitoring company, and Fireworks, which serves models. Laurie Voss ran ten models through 40 agent tasks, six times each.[1] The table has two price columns, and they disagree.

A slope chart with two columns. Price per attempt: Claude Sonnet 5 $0.49, Kimi K3 $0.44, GPT-5.5 $0.42, Gemini 3.5 Flash $0.29. Price per successful task: Gemini 3.5 Flash $1.23 at a 23 percent pass rate, Claude Sonnet 5 $1.01 at 49 percent, Kimi K3 $0.67 at 66 percent, GPT-5.5 $0.64 at 67 percent. The Gemini line crosses from cheapest to most expensive.
The cheapest model per attempt becomes the most expensive per finished task once its failures are counted. Data: Arize and Fireworks, 2026.

The multiplier between the two columns is what Arize calls the retry tax. For GPT-5.5 and Kimi K3 it was 1.5 times. For Gemini 3.5 Flash it was 4.3.

We made the case for this metric in why OpenAI's price cut would not move agent bills. The data here shows how far the two numbers can drift apart.

Is Kimi K3 cheaper than GPT-5.6 Sol?

Per attempt, yes, by a wide margin. On a coding benchmark, Together AI measured Kimi K3 at $4.65 per run against $8.37 for GPT-5.6 Sol, and 2.8 times more solved tasks per dollar. But Sol solved more tasks on all four tries. Which one is cheaper for your team depends on whether a retry is acceptable.

The Together comparison, by Zain Hasan and Shobhit Dixit, ran both models four times on each task of a software engineering benchmark called DeepSWE.[2] Together sells hosting for Kimi K3, so read it as a vendor's test. The numbers still show the split clearly:

  • One try: Sol solved 72.7 percent, Kimi K3 68.5.
  • Four tries: Kimi K3 solved 89.4 percent, Sol 85.8.
  • Every try: Sol solved 61 tasks on all four attempts, Kimi K3 45.

If your workflow lets an agent retry until a test passes, Kimi K3's lower price wins. If a task has to be right the first time, because a person is waiting or a wrong answer ships, Sol's reliability is what you are paying for.

Why do models that tie on average win at opposite ends?

Because an average mixes easy and hard work together, and cheaper models are often strongest on the easy part. Two models with the same overall pass rate and the same price can split the work between them, one clearing every simple task and the other doing better on the hardest ones.

Arize found exactly that. GPT-5.5 and Kimi K3 finished in a near tie: 67 and 66 percent success, at 64 and 67 cents per successful task. Kimi K3 solved every easy task in all six trials, where GPT-5.5 slipped to 69 percent. On the hardest tasks they swapped, 51 percent for GPT-5.5 against 32 for Kimi K3.[1]

An average price per task tells you almost nothing until you know the difficulty of your own work.

Hard tasks are also where retrying stops helping. A model either has the capability or it does not.

A model that cannot do a task does not learn it on the fourth attempt, it just bills you four times.

Laurie Voss, Arize AI, Cost per successful task, 2026

Does routing between models save money?

Often, when two models fail on different tasks. Together found that a cascade sending work to Kimi K3 first and escalating failures to GPT-5.6 Sol covered 108 of 113 tasks, beating either model alone. Routing works on easy and mixed work. It cannot rescue a task that no cheap model can do.

The two models in Together's test succeeded on different tasks, with only a loose 0.46 correlation between which ones each got right.[2] That is what makes a cascade worth building: the second model catches what the first one misses, and the expensive call only happens when it is needed.

Arize's data shows the limit. The cheapest model per success in the whole study, an open model called gpt-oss-120b, reliably solved only 8 of the 40 tasks.[1] Its price looked good because it only won the tasks it could do.

For the rest, you pay for a stronger model or the work does not get done.

Your own code base moves this number too. In one published test, refactoring a large file cut the tokens an agent read for the same change by 83 percent, with the same model at the same price.

How should a team compare AI model costs?

On its own work. Decide what counts as a finished task, run each candidate model several times on a sample of real jobs, and divide total spend by successful results. Split the results by difficulty, because a tie on average often hides a split underneath. Published benchmarks are a starting point, not a verdict.

A practical version:

  1. Pick 20 to 40 real tasks from recent work, easy and hard.
  2. Write down what "done" means for each one before running anything.
  3. Run each model three to six times per task and log every attempt's cost.
  4. Divide total spend by passes, then split the result by difficulty.
  5. Decide on retries. If your work allows them, cheaper models gain. If it does not, reliability is the number to buy.

Step 2 is the one that decides the answer, and it is the same step we described for checking a replacement model before the old one retires. The person who knows what the work is for has to write it down.

Where your seat fits in

All of this is about API work, where each call is billed. Developers on flat-rate coding subscriptions mostly do not see per-call prices at all, which we covered in how Claude Code pricing depends on your seat.

The moment a workload moves to the API, the number that matters is cost per successful task, and it is only knowable on your own work.

References

Frequently asked

What is cost per successful task for an AI model?
›It is everything a model spent across every attempt, including runs that failed, retried or timed out, divided by the number of times it actually finished the task.
⌄It is everything a model spent across every attempt, including runs that failed, retried or timed out, divided by the number of times it actually finished the task. Per-token and per-attempt prices cannot see retries or failures. Cost per successful task counts all of them, which is why it can rank models in a different order.
Is a cheaper AI model per token actually cheaper to run?
›Not necessarily. In one benchmark of 2,400 agent runs, the cheapest of four comparable models per attempt cost the most per finished task, because it passed only 23 percent of the time.
⌄Not necessarily. In one benchmark of 2,400 agent runs, the cheapest of four comparable models per attempt cost the most per finished task, because it passed only 23 percent of the time. The failed attempts are part of the bill.
Is Kimi K3 cheaper than GPT-5.6 Sol?
›Per attempt, clearly. Together AI measured Kimi K3 at a little over half the cost of Sol per coding run, with nearly three times more solved tasks per dollar.
⌄Per attempt, clearly. Together AI measured Kimi K3 at a little over half the cost of Sol per coding run, with nearly three times more solved tasks per dollar. But Sol solved more tasks on all four tries, 61 against 45. Which is cheaper for you depends on whether a retry is acceptable in your work.
How should a team compare AI model costs?
›On its own tasks. Write down what counts as a finished task, run each candidate model several times on a sample of real work, and divide total spend by successful results.
⌄On its own tasks. Write down what counts as a finished task, run each candidate model several times on a sample of real work, and divide total spend by successful results. Split the results by difficulty, because models that tie on average often win at opposite ends.
Does routing between AI models save money?
›Often, when two models fail on different tasks. Together AI found that trying Kimi K3 first and escalating its failures to OpenAI's Sol model covered 108 of 113 coding tasks.
⌄Often, when two models fail on different tasks. Together AI found that trying Kimi K3 first and escalating its failures to OpenAI's Sol model covered 108 of 113 coding tasks. Routing cannot rescue a task that no cheap model can do, though. There the only option is to pay for the stronger model.
Work with us

Let’s build it together.

We turn clever prototypes into production systems people can rely on. If you’re building with agents and want a hand making it real, leave your email and we’ll be in touch.

Straight to the team. No spam.