Sep 2026


AI providers keep raising prices, though mostly not by touching the price. They quietly cut the limits instead. I work in Codex, and from what I can tell its limits have dropped by at least 2–3x over the last three months.

I'm doing a lot of indie development right now, so my costs are essentially token costs. Running every task on Sol High or Opus High stopped being an option.

For reviews and genuinely hard tasks, the answer is simple: use whatever your budget allows. I stay on Sol High or Opus High, and they crack something like 95% of what I throw at them.

Routine work is the harder question, because it needs a balance. The model has to be much cheaper than a big model at high effort, otherwise there's no point switching. But it can't be too dumb either — then it starts making mistakes even on small, decomposed pieces, creates more work for the reviewer, and wipes out the savings.

So I wanted a way to compare cost against quality instead of going on impressions. And not just across models — across providers, and across effort levels of the same model.

The internet is full of benchmark charts ranking models by intelligence. But when price is what you care about, the smartest kid on the block isn't the question. The ratio is.

That's how I ended up on Artificial Analysis. One chart there does what I need: Cost per Task against Intelligence, for coding agents specifically.

Artificial Analysis Coding Agent Index vs. Cost per Task
Artificial Analysis Coding Agent Index vs. Cost per Task. Source: Artificial Analysis

The numbers were more interesting than I expected. On the left is the average cost of completing a task for this model and effort, on the right is its level of intelligence:

  • Luna Max — $1.51 · 57
  • Terra High — $1.56 · 55
  • Sol Low — $1.66 · 55
  • Terra xHigh — $1.88 · 56
  • Sol Medium — $2.81 · 62
  • Sol High — $3.85 · 64

Terra turns out to be a useless middle. For roughly the same money you can take Luna Max, which is both slightly smarter and slightly cheaper, or move up to Sol Low. And Sol Low has an extra advantage if you already run Sol High on the hard tasks: switching models resets the cache, so every switch also costs you the savings on cached tokens.

There's one more number, though — how long a model takes to finish a task.

  • Sol Low — 3.5 min per task
  • Luna Max — 8 min per task

Here Sol Low wins outright, and everyone has to decide for themselves which they'd rather have: the couple of extra intelligence points, or the time back.

I've spent some time staring at these charts. It's the only way I've found so far to get a real sense of what the models can do and how to work with them.

I expect subagents will eventually pick models and effort levels themselves, faster and better than I can, and then none of this will matter. For now we read charts.


I'm Nick, a designer with 15 years of experience, formerly at Yandex. Now building Leaflo, a mental health app for self-help and reflection.

LinkedIn