

Tokens and credits, explained
Models read and write in tokens, small chunks of text. A token is roughly three quarters of a word. A short question might be 20 tokens. A thorough answer, 800.
You pay twice per exchange: once for what goes in, once for what comes out. Output usually costs more. Reasoning models also bill the thinking they do before answering, even though you barely see it.
Credits put all models on one meter. A quick answer from a budget model costs a handful of credits. A long, deeply reasoned answer from a frontier model costs hundreds. Neither is wrong. They’re different tools, and the next units show when each one earns its price.
Go deeper
How text turns into tokens
Before a model reads anything, the text is cut into tokens. Common English words are usually one token each. Longer or rarer words split into pieces: "amplifier" might be two tokens, "Klipschorn" three or four. Numbers, units, and product codes tokenize badly, so a spec sheet full of "8 ohms", "2.83 V", and model codes like "SN-2300" carries more tokens per word than plain prose.
The rule of thumb still holds: three quarters of a word per token, on average. "Which of my speakers has the lowest sensitivity?" is about twelve tokens. The full spec table of a floorstander, pasted in from the manual, is easily 600. The paste is the real cost of that message. The question is almost free.
Practical takeaway: paste the lines that matter, not the whole document. Impedance and sensitivity fit in one line. The twelve-page manual does not.
Three meters, and which one to read first
Every exchange runs on up to three meters. Input tokens are everything you send, including the history of the chat so far. Output tokens are everything the model writes back. Reasoning tokens are the model’s private working through of the problem before it answers. You see a short summary of that thinking at most, but every token of it is billed, at the output rate.
Output costs more than input, on most models several times more. Reasoning adds a second effect: on a hard question, the invisible reasoning can run longer than the visible answer. A two-paragraph reply about matching a tube amp to low-sensitivity speakers may sit on top of thousands of reasoning tokens you never read. That is where the money goes.
Rates are published per one million tokens, the industry convention. A million tokens is roughly 750,000 words. That sounds enormous until you remember that a chat of thirty long messages resends its own history every turn. The Model costs page lists the input and output rate for every model in the picker.
Practical takeaway: when comparing two models on cost, read the output rate first. For a reasoning model, that rate is the price of the thinking too.
What credits do
Every model in the picker has its own price list, in dollars, with separate input and output rates. Nearly every subscription product that offers several AI models solves this the same way, with a credit as the common unit. Credits collapse all of that into one number. Each model’s published rates are converted into a credit cost, and the picker shows the cost units next to every model name, so the comparison happens before you ask, in one unit.
Your plan includes a monthly credit allowance. Every message deducts what that exchange actually used: tokens in, tokens out, reasoning tokens, each at the model’s own rate. There is no flat charge per message, which is why the same question can cost a handful of credits on a small model and a few hundred on a frontier reasoning model. The meter measures work done, and different models do different amounts of work for the same prompt.
The usage overview in the Assistant shows what you have used over the past days and months. Estimating in advance is guesswork, because the cost depends on how much history the chat carries and how long the model decides to think.
Practical takeaway: use the Assistant normally for a week, then read the usage overview. It tells you more about your real consumption than any calculation up front.
Where credits quietly disappear
Four habits burn credits without improving answers. Twenty short follow-ups instead of one complete question, because every follow-up resends the whole chat. Whole manuals pasted when three numbers would do. Definition lookups on a frontier reasoning model, where the thinking costs more than the answer is worth. And one chat left open for weeks, so a question about cartridge alignment drags along last month’s conversation about room treatment.
None of these produces a wrong answer. They are the difference between an allowance that lasts the month and one that runs out on the 15th. The next three units cover the fixes: matching the model to the question, front-loading the prompt, and starting a new chat when the topic changes.
Practical takeaway: one complete question, the relevant numbers only, the smallest model that can answer it, in a fresh chat. That is most of the savings, and the answers get better at the same time.