AI Academy
Poseidon riding a chariot through crashing waves, trident raised, drawn by a sea horse.
Intermediate

Why long chats drift and drain

Every model has a context window, a hard cap on how much conversation it can consider at once. Your messages, its answers, everything counts against that cap.

Two things happen as a chat grows. Cost climbs, because the full history travels with every new message. And focus fades. The room dimensions you gave in message three get buried under thirty follow-ups, and the model starts contradicting itself or forgetting constraints.

The fix is hygiene. New topic, new chat. And when a long conversation is still valuable, ask the model to summarize where things stand, then paste that summary into a fresh chat. You keep the substance and drop the dead weight.

Go deeper

The context window is a fixed budget

Every model reads from a context window: the maximum amount of text it can consider in one go. It includes your messages, its own answers, anything you pasted, and whatever the Assistant looked up in the database on your behalf. The window is large on modern models, but it is finite, and a long chat fills it faster than it feels like it should, because every answer the model wrote is in there too.

When the window fills, the oldest material has to go. The model does not announce this. It simply stops seeing the room dimensions you gave in message three, and starts answering as if you never gave them.

Practical takeaway: a chat has a shelf life. Somewhere past thirty or forty exchanges you are working from a window that has already dropped things, and the answers stop reflecting everything you told it.

Drift is what forgetting looks like from outside

Long before the window overflows, attention thins out. A model weighs everything in the window, but the more there is, the less any one line matters. A constraint stated once and then buried under thirty messages of other discussion carries little weight by the end. You will notice it as the model recommending a floorstander after you said the speakers must fit a bookshelf, or quoting a budget you raised and then lowered again.

Contradictions creep in the same way. The model said one thing in message eight, something incompatible in message twenty, and it cannot see the tension because both are just text in a very long window.

Practical takeaway: when the model contradicts itself or loses a constraint, the chat has drifted. Restate the constraint or start over, do not argue with it.

Cost climbs on a curve

Cost and drift come from the same cause. Because the whole history is resent every turn, the cost of a message grows with the length of the chat. The fortieth message in a detailed conversation can cost more than the first ten combined, and it is also the message most likely to be answered badly. You are paying the most for the worst answers.

Practical takeaway: the point where a chat gets expensive is the point where it gets unreliable. Both are signals to close it.

The reset that keeps the substance

Ending a chat does not mean losing the work. Ask the model to summarize where things stand: what you own, what you have decided, what is still open, what has been ruled out and why. Read the summary, fix anything it got wrong, and paste it as the first message of a fresh chat. The new conversation starts with the substance of forty messages in a few hundred tokens, and the model can see all of it clearly again.

Do this at every topic change as well. Speaker placement and streamer choice are two separate chats. The Assistant keeps every conversation saved, so nothing is lost by splitting.

Practical takeaway: new topic, new chat. Long topic, summarize and restart. Both cost less than pushing on.