

Why models cost wildly different amounts
Behind every model sits different hardware doing different amounts of work. Frontier models from OpenAI, Anthropic, Google, and xAI are enormous. Every token they produce takes far more computation than a small model needs.
Reasoning adds another layer. Some models think through a problem step by step before they answer, and that hidden work is billed in tokens too. A two-line answer can carry minutes of invisible reasoning behind it.
So the spread is real. Per token, a frontier reasoning model can cost a hundred times more than a budget model. Price tracks computation. Whether your question needs that computation is a different matter, and the intermediate track is about exactly that.
Go deeper
Parameters are the amplifier size
A model’s size is measured in parameters, the internal weights it learned during training. Small models have a few billion. Frontier models have hundreds of billions or more. Every token generated passes through all of them. A model ten times larger does roughly ten times the arithmetic per token, on ten times the hardware, and that hardware is the most expensive computing equipment in the world right now.
The hi-fi parallel is a fair one. A 30 W integrated and a 300 W monoblock both play music. The bigger one costs more to build, more to run, and most of the time you are not using what you paid for. The difference is that with amplifiers you buy once. With models you pay per token, so the size of the tool shows up on every single answer.
Practical takeaway: read the model group in the picker as a size class, the same way you read a power rating. Big is for when the job needs it.
Reasoning is billed by the minute, in tokens
Some models are trained to reason before they answer. Given a question, they write out a working process, check it, revise it, and only then produce the reply you see. That working process is generated the same way as any other text, token by token, and every token is billed at the output rate. You see a summary at most. The bill sees all of it.
On a hard question this is worth it. Sizing an amplifier for a difficult speaker load in a large room with a budget constraint has many interacting variables, and a model that thinks through them makes fewer mistakes than one that answers on reflex. On a simple question the same model still reasons, because that is what it does, and you pay for deliberation the answer never needed.
Practical takeaway: reasoning models earn their price on multi-variable problems. Give them those, and give lookups to something that answers on reflex.
The spread is the point
Put the two effects together. Size multiplies the cost per token. Reasoning multiplies the number of tokens. A frontier reasoning model is large and thinks at length, so per question it can land a hundred times above a small model that answers in one pass. The Model costs page shows the actual rates per million tokens, input and output separately, for every model in the picker.
The spread tracks the electricity and hardware behind each answer. A small model is a small model because it does less computation, and for a definition or a unit conversion, less is exactly enough.
Practical takeaway: the picker groups models by class. Start one class below where you think you need to be, and move up if the answer comes back thin.
What the expensive model buys you
Price buys three things in practice. Fewer dropped constraints when a question has many. Better judgment when the question has no clean answer, like which of two well-matched amplifiers suits a room with a bass problem. And more reliable arithmetic when the steps are shown. It does not buy better facts. Every model in the picker reads the same Pure Neo database, so the specs behind a grounded answer are identical whether a small model or a frontier one wrote the sentence around them.
Practical takeaway: pay for reasoning and judgment. The data comes free with every model, and it is the same data.