AI

What a context window is, and why it fills up

A language model remembers nothing between requests. Everything it knows about your conversation has to fit on one page, and that page has an edge.

GetNetwork2 min readAI-assisted

A language model does not remember your conversation. Each time you send a message, the application sends the model the whole exchange again, from the first line to the newest one, and the model reads it fresh. The space that holds all of this is the context window.

Measured in tokens

The window is measured in tokens, which are fragments of words. A short common word is usually one token. A long or unusual word is split into several. As a rough guide for English, a token is about three quarters of a word.

Everything counts against the limit:

  1. The instructions the application gives the model before you type anything.
  2. Every message you have sent and every reply you have received.
  3. Any files, search results or tool output added along the way.
  4. The answer the model is about to write.

What happens at the edge

When the window is full, something has to go. Applications handle this in different ways. Some drop the oldest messages. Some replace the early part of the conversation with a summary. Some simply stop and ask you to start again.

This is why a long chat can seem to forget what you said at the beginning. The model has not become careless. That part of the conversation is no longer on the page it is reading.

Bigger is not the whole answer

Windows have grown enormously, and a larger window does help. But a model given a very long page still has to find the relevant part of it, and attention spread over a great deal of text is thinner than attention on a little.

The practical lesson is the same one that applies to briefing a person. Put what matters in front of the model, leave out what does not, and start a new conversation when the subject changes.

Keep reading