Tokens and Context Windows: Why AI Chats Forget and Files Get Cut Off
A plain-English guide to tokens and context windows, and why long AI chats lose track of earlier details or reject a long file.
You’ve been chatting with an AI assistant for an hour, planning a family reunion. Early on you told it your aunt is vegetarian. Now it’s suggesting a pig roast. Or you paste in a long report and get a message saying it’s too long.
Both problems come from the same two ideas: tokens and the context window. Once you understand them, a lot of strange AI behavior makes sense, and you’ll know how to work around it.
What a token is
AI language models don’t read text one letter at a time, and they don’t read whole words either. They break text into small pieces called tokens.
Think of tokens as the Lego bricks of text. A short, common word like “the” or “cat” is usually one brick. A longer or rarer word like “unbelievably” might be split into a few bricks, something like “un”, “believ” and “ably”. Punctuation and spaces often take up bricks too.
A few things follow from this:
- Tokens and words aren’t the same. In everyday English, a token is usually a bit smaller than a word on average. So 1,000 words of English comes to somewhat more than 1,000 tokens.
- Different text uses tokens at different rates. Plain English prose tends to be efficient. Unusual names, long numbers, computer code, and many languages other than English often need more tokens to say the same thing.
- Each AI company splits text its own way. The same paragraph can count as a different number of tokens in different assistants.
You almost never have to count tokens yourself. It just helps to know that this is the unit the AI uses to measure how much text it’s handling.
What a context window is
The context window is how much text the AI can keep in view at once, measured in tokens.
Picture a desk. Everything the AI uses to write its next reply has to fit on that desk: your messages, its own earlier replies, any files you attached, and the background instructions the company gives it. The desk is big, but it has edges.
When the desk fills up, something has to come off. That’s the key point.
An accurate version of the analogy: an AI model doesn’t remember your conversation the way a person does. Each time you send a message, the assistant hands the model the conversation so far, or as much of it as fits, and the model writes a reply based on what it’s looking at. If something is no longer on the desk, the model can’t see it. It hasn’t “forgotten” it in the human sense. It simply isn’t there.
Why long chats forget
Back to the reunion chat. Every message you send and every reply you get adds to the pile. In a long enough conversation the pile can outgrow the desk.
Assistants handle this in different ways. Some quietly drop older messages. Some replace older parts with a short summary, and details can get lost in that summary. Some tell you the conversation has reached its limit. If the reunion chat got long enough, the message about your vegetarian aunt may have been dropped or summarized away. But even when it’s still on the desk, the model can overlook it.
That second problem is often the more likely one, because many assistants now have very large desks. With very long inputs, AI models can be less reliable at picking out one detail buried among a lot of other text. It’s a bit like a person skimming a 40-page document: everything is technically in front of them, but a small note on page 12 is easy to miss. How much this matters depends on the model and the task, and researchers are still studying it.
Some assistants also have a separate “memory” feature that saves facts about you between chats. That’s a different system from the context window. It stores selected notes, not the whole conversation.
Why some files are too long
When you attach a file, its text has to fit on the same desk as everything else. A long contract, a book manuscript, or a big spreadsheet can take up a huge number of tokens.
If the file is too big, you might see:
- An error saying the file or message is too long.
- The assistant working from only part of the file, sometimes without telling you clearly.
- Answers that seem to ignore sections near the end, or the middle.
Spreadsheets and PDFs can surprise you. A spreadsheet full of numbers, or a PDF with complicated formatting, can use far more tokens than its page count suggests.
What you can do about it
You don’t need technical skills for any of these:
- Start a fresh chat for a new task. A clean desk works better than a cluttered one.
- Restate what matters. If a detail is important, repeat it near your question: “Remember, my aunt is vegetarian. Suggest a menu.”
- Ask for a recap, then carry it over. Before a long chat gets unwieldy, ask the assistant to summarize the key decisions and facts. Check the summary, fix anything wrong, and paste it into a new chat.
- Split big files. Send one chapter or section at a time, and ask about each piece separately.
- Paste only the relevant part. If your question is about the cancellation clause, paste that clause, not the whole lease. This also means you share less of your information.
- Ask where an answer came from, then check it yourself. “Which section of the document says that?” gives you somewhere to look. But assistants can confidently point to a section or quote that doesn’t exist, or get it wrong. Open the document and read that part yourself before you rely on the answer.
A note on privacy: Don’t paste sensitive information, such as account numbers, ID numbers, health details, confidential work documents or other people’s private data, until you’ve checked the service’s data and privacy settings.
A note on legal documents: An AI’s reading of a contract or lease clause isn’t legal advice. For important decisions, check with a qualified professional.
What’s settled and what’s still open
Well established:
- AI language models process text as tokens, not letters or whole words.
- Every model has a limit on how much it can consider at once, and both your input and its replies count toward that limit.
- Text that falls outside that limit can’t influence the reply, except through a summary or saved memory that keeps a condensed version of it.
Still changing or debated:
- How big context windows will get. They have grown a lot, and the limits differ between products and plans.
- How well models actually use very long inputs. A bigger desk doesn’t automatically mean the model pays equal attention to everything on it.
- How best to handle a full desk. Dropping, summarizing and storing memories each have trade-offs, and companies keep changing their approach.
The short version
Tokens are the small chunks of text an AI reads. The context window is how many of those chunks it can look at in one go. Long chats and big files can go past that limit, and the AI loses sight of what fell off. Even inside the limit, it can miss details buried in a lot of text. You can manage both by keeping conversations focused, repeating key details, and sending large documents in pieces.
Related
What AI Agents Really Are, and Where They Help or Fail
A plain-English guide to AI agents: how they differ from chatbots, what tools they use, and where they're useful or still unreliable.
What Is a Large Language Model? How AI Chatbots Work
A plain-English guide to the technology behind AI chatbots: how they learn, why they sound so sure of themselves, and where they fall short.
How AI Image Generators Work: From Noise to Picture
A plain-English look at how AI turns a sentence into a picture, why it sometimes gets things wrong, and where the copyright debate stands.