AI Trends / Opinion Articles

Choosing AI Models by Value for Money

In the past, people compared who had the smartest model. I've become increasingly clear that now, what's really being compared is the CP value of completing a task: time, computational power, and money combined, who is the most cost-effective.

Recently, I've started doing something I never did before: calculating AI's CP value.

In the past one or two years, what were people comparing? Who had the smartest model. Who had the highest score, who had the best ranking, and who broke some record. But I've become increasingly clear that this way of comparing has become outdated. Now, what's really being compared is who has the highest CP value for completing a task. And this CP value includes three things: time, computational power, and money.

Who is this for
  • Those who have several models available and are used to directly opening the strongest one
  • Those who start to feel that API bills or subscription fees are going up
  • Those who want to know which level of model should be used for a task
What you can take away
  • Three models' actual cost comparison, understanding why cheaper might be more expensive
  • A judgment sentence for selecting a model: first ask how many points are needed this time
  • A more preceding question: which problems are worth being solved

Three models' analogy

Assume you have three models.

The first is the flagship model, with a strength of 100 points. You ask it a question, and it gives you a full score answer in one go. The disadvantage is that it is very expensive, costing 100 dollars per question, and taking 10 minutes to run.

The second is the mid-level model, with only 80 points per question. But it's cheap, costing only 10 yuan per use. You can let it iterate repeatedly: the second time it scores 90, the third time 95, and by the fifth time it reaches 100. Five times together cost 50 yuan, which is still a full score, but only half the price of the flagship model. The cost is one session takes 10 minutes, five sessions take 50 minutes, which is more time-consuming.

The third one is the cheapest model, costing 5 yuan per use, fast, but only scores 60 points per use. It sounds like the best value for money, right? The problem is its upper limit is low, and repeated iterations can lead to distortion, making it less effective over time. To push it to 100 points, you might need to ask 30 times, totaling 150 yuan, and it would take two and a half hours.

Have you noticed? Using the cheapest model to achieve a full score is actually the most expensive and slowest.

A table to understand clearly

The basics of the three models

ModelSingle-use qualitySingle-use priceSingle-use time
Flagship100 points100 yuan10 minutes
Mid-range80 points10 yuan10 minutes
Budget60 points5 yuanabout 5 minutes

the cost of reaching 100 points individually

Modelhow many iterations to reach full markstotal costtotal time spent
Flagship1 time100 yuan10 minutes
Mid-range5 times (80→90→95→100)50 yuan50 minutes
Budget30 times (60→80→90→100)150 yuan2.5 hours

See how many points you need, and choose different models

Your requirementsBest choice
100 points, need the fastestFlagship (10 minutes, done in one go)
100 points, need the cheapestMid-range (50 yuan, half the cost of flagship)
Only 80 points neededMid-range (10 yuan per call is enough)
Only 60 points neededBudget (5 yuan per call, 5 minutes)
These numbers are metaphors Scores, prices, and time are all examples to illustrate trade-offs, not actual pricing from any model. You should replace these numbers with your own for your task.

So the real question is always 'How many points does your task actually need?' Asking which model is the strongest comes later.

To get 100 points as fast as possible, use the flagship, which finishes in 10 minutes. To get 100 points as cheaply as possible, use the mid-range, letting it work slowly. If the task actually only needs 60 points, the cheapest model is enough for one call, so why use the flagship?

We're all becoming cost calculators

That's why I say models aren't necessarily the smarter the better. We, who use AI every day, are becoming a new kind of role: cost calculators.

Why has cost calculation become so important? Because two things are happening at the same time. On one hand, the tasks we need to handle are becoming more complex. On the other hand, computing power is getting more expensive, whether you're using a cloud API or local models with hardware, the cost is going up.

When every call costs money and time, you can't just blindly say 'the strongest model is always best.' You need to choose the right tool for each task like a person who knows how to calculate.

But the bigger question is: which problems are worth solving?

What I want to talk about is actually not just the matter of selecting a model.

When new generations of models are becoming increasingly capable of solving problems, AI has already been used for tasks such as early cancer prevention. I think the questions humans should consider have quietly shifted.

Previously, we asked: Can AI solve this problem? Now this question is almost obsolete, because the answer is increasingly often 'yes'. Thus, the real question has surfaced:

Two sorting questions Which problems are worth solving? Which problems should be prioritized?

Computing power is limited. The number of problems to solve is overflowing. When the resources you have are no longer unlimited intelligence, but a grid of costs to be calculated, you are forced to sort.

This is actually the same thing as the previous CP value table, just scaled up to the human level: first think clearly whether 'this thing is worth doing to perfection', and then decide whether to spend effort on 'whether it can be done to perfection'.

Being able to calculate CP value is the basic skill of this era. Being able to sort which problems are worth solving is the real dividing line.

AI Trends Model selection Cost calculation Decision support

Use AI as a tool to help you accumulate judgment

I host two free online seminars every month, discussing how to organize workflows, judgments, and experience into prompts, skill packages, and knowledge bases that AI can flexibly use. If you want to receive notifications about the seminars or want to discuss which model level is suitable for your task, feel free to start from the community.

Join LINE Community

Free seminar sessions will be announced here first, and you can also bring your actual usage data to calculate together.

Join the community ↗