AI Academy
Artemis drawing her bow beneath a crescent moon, hound at her side, a stag in the distance.
Intermediate

One good prompt beats twenty messages

Treating the Assistant like a texting conversation drains credits fast. Every follow-up resends the entire chat so far. Ten rounds in, each new message carries ten messages of history, and you pay for all of it again.

Front-load instead. Put everything that matters into the first prompt: your gear, your room, your budget, your goal, what you’ve already ruled out.

Weak: "best bookshelf speakers?" Strong: "Bookshelf speakers under $2,000 for a 20 m² room, close to the rear wall, driven by a 60 W amp, mostly jazz and vocals at moderate volume. Shortlist five, with measurements that support each pick."

The detailed prompt costs a few extra tokens up front. It saves hundreds of tokens of clarification, and the answer quality jumps.

Go deeper

Why the meter runs on history

A model has no memory between messages. Each time you send something, the Assistant packages the entire conversation so far and sends it along, because that is the only way the model can know what you said earlier. Message one carries one message. Message ten carries ten. The tokens you pay for on each turn are the whole stack, every time.

A conversation that starts with "best bookshelf speakers?" and then answers six clarifying questions one at a time has resent the growing history six times. The same information delivered in the first message would have cost one pass. The answer would also have been better, because the model saw all the constraints at once instead of revising as they trickled in.

Practical takeaway: everything the model needs to know, it needs to know before it starts. Put it in the first message.

What belongs in the first message

Five things, most of the time. Your gear, at least the parts the question touches: the amplifier’s power and impedance rating, the speakers’ sensitivity, the source. Your room, in numbers: dimensions, where the speakers sit relative to the walls, how much soft furnishing. Your budget, as a number with a currency. Your goal, stated as what you want to hear differently, at what volume, with what music. And what you have already ruled out, so the model does not spend its answer on a path you closed.

If your profile is filled in, the gear and room come for free and you can skip them. The goal, the budget for this specific decision, and the exclusions still belong in the prompt, because they change from question to question.

Practical takeaway: gear, room, budget, goal, exclusions. If one is missing, the model will guess it, and you will pay to correct the guess.

Asking for the answer shape you want

The other half of a good prompt is telling the model what a good answer looks like. "Shortlist five, with the measurement that supports each pick" gets you a table with reasons. "Explain the tradeoff" gets you an essay. "Rank them and say why the top one wins" gets you a verdict. Left unspecified, the model picks a shape at random, and it is usually the longest one, which you pay for.

Constraints on the answer also cut cost directly. "Under 200 words" or "three candidates, one sentence each" is a token budget you set. The model respects it, and the reply is easier to act on.

Practical takeaway: say how long, how many, and in what form. The answer gets shorter, cheaper, and more useful at once.

Weak, strong, and what changed

Weak: "Is this amp enough for my speakers?" The model has to ask which amp, which speakers, how big the room, how loud you listen. Four rounds, four resends, then an answer.

Strong: "I have a 60 W into 8 ohm integrated driving 86 dB speakers with a 4 ohm nominal load, 3 m listening distance, 20 m² room, mostly acoustic jazz at moderate levels. Is that enough headroom, and at what point would it not be? Two paragraphs." One round. The model has the impedance, the sensitivity, the distance, the use case, and the answer format. It can calculate the peak level at your seat and say whether the amp’s 4 ohm behavior matters.

The strong prompt is about 60 tokens longer. It replaces four exchanges and a vaguer result.

Practical takeaway: write the strong version once. It is the cheapest message in the whole conversation.