Recently, I've started doing something I never did before: calculating AI's CP value.
In the past one or two years, what were people comparing? Who had the smartest model. Who had the highest score, who had the best ranking, and who broke some record. But I've become increasingly clear that this way of comparing has become outdated. Now, what's really being compared is who has the highest CP value for completing a task. And this CP value includes three things: time, computational power, and money.
- Those who have several models available and are used to directly opening the strongest one
- Those who start to feel that API bills or subscription fees are going up
- Those who want to know which level of model should be used for a task
- Three models' actual cost comparison, understanding why cheaper might be more expensive
- A judgment sentence for selecting a model: first ask how many points are needed this time
- A more preceding question: which problems are worth being solved
Three models' analogy
Assume you have three models.
The first is the flagship model, with a strength of 100 points. You ask it a question, and it gives you a full score answer in one go. The disadvantage is that it is very expensive, costing 100 dollars per question, and taking 10 minutes to run.
The second is the mid-level model, with only 80 points per question. But it's cheap, costing only 10 yuan per use. You can let it iterate repeatedly: the second time it scores 90, the third time 95, and by the fifth time it reaches 100. Five times together cost 50 yuan, which is still a full score, but only half the price of the flagship model. The cost is one session takes 10 minutes, five sessions take 50 minutes, which is more time-consuming.
The third one is the cheapest model, costing 5 yuan per use, fast, but only scores 60 points per use. It sounds like the best value for money, right? The problem is its upper limit is low, and repeated iterations can lead to distortion, making it less effective over time. To push it to 100 points, you might need to ask 30 times, totaling 150 yuan, and it would take two and a half hours.
A table to understand clearly
The basics of the three models
| Model | Single-use quality | Single-use price | Single-use time |
|---|---|---|---|
| Flagship | 100 points | 100 yuan | 10 minutes |
| Mid-range | 80 points | 10 yuan | 10 minutes |
| Budget | 60 points | 5 yuan | about 5 minutes |
the cost of reaching 100 points individually
| Model | how many iterations to reach full marks | total cost | total time spent |
|---|---|---|---|
| Flagship | 1 time | 100 yuan | 10 minutes |
| Mid-range | 5 times (80→90→95→100) | 50 yuan | 50 minutes |
| Budget | 30 times (60→80→90→100) | 150 yuan | 2.5 hours |
See how many points you need, and choose different models
| Your requirements | Best choice |
|---|---|
| 100 points, need the fastest | Flagship (10 minutes, done in one go) |
| 100 points, need the cheapest | Mid-range (50 yuan, half the cost of flagship) |
| Only 80 points needed | Mid-range (10 yuan per call is enough) |
| Only 60 points needed | Budget (5 yuan per call, 5 minutes) |
So the real question is always 'How many points does your task actually need?' Asking which model is the strongest comes later.
To get 100 points as fast as possible, use the flagship, which finishes in 10 minutes. To get 100 points as cheaply as possible, use the mid-range, letting it work slowly. If the task actually only needs 60 points, the cheapest model is enough for one call, so why use the flagship?
We're all becoming cost calculators
That's why I say models aren't necessarily the smarter the better. We, who use AI every day, are becoming a new kind of role: cost calculators.
Why has cost calculation become so important? Because two things are happening at the same time. On one hand, the tasks we need to handle are becoming more complex. On the other hand, computing power is getting more expensive, whether you're using a cloud API or local models with hardware, the cost is going up.
When every call costs money and time, you can't just blindly say 'the strongest model is always best.' You need to choose the right tool for each task like a person who knows how to calculate.
But the bigger question is: which problems are worth solving?
What I want to talk about is actually not just the matter of selecting a model.
When new generations of models are becoming increasingly capable of solving problems, AI has already been used for tasks such as early cancer prevention. I think the questions humans should consider have quietly shifted.
Previously, we asked: Can AI solve this problem? Now this question is almost obsolete, because the answer is increasingly often 'yes'. Thus, the real question has surfaced:
Computing power is limited. The number of problems to solve is overflowing. When the resources you have are no longer unlimited intelligence, but a grid of costs to be calculated, you are forced to sort.
This is actually the same thing as the previous CP value table, just scaled up to the human level: first think clearly whether 'this thing is worth doing to perfection', and then decide whether to spend effort on 'whether it can be done to perfection'.
Being able to calculate CP value is the basic skill of this era. Being able to sort which problems are worth solving is the real dividing line.
Use AI as a tool to help you accumulate judgment
I host two free online seminars every month, discussing how to organize workflows, judgments, and experience into prompts, skill packages, and knowledge bases that AI can flexibly use. If you want to receive notifications about the seminars or want to discuss which model level is suitable for your task, feel free to start from the community.
Free seminar sessions will be announced here first, and you can also bring your actual usage data to calculate together.
Join the community ↗