The Long Run
← Back to blog

Data Science

Context Windows Demystified: How Much Can Your AI Actually Remember?

The finale of our LLM tokenization series: what context windows are, how they relate to token limits, and how real AI products manage them.

·6 min read

Part 4 of 4 in the "LLM & Tokenization" series

This is the last in series of the 4 part where we described LLM & Tokenisation.

In this final article we will connect all of that to a term you've almost certainly heard thrown around: the context window.

What Is a Context Window, Really?

The context window is the total number of tokens an LLM can "see" at any one time, your entire conversation history, any documents that was pasted, its own previous replies, plus the response it's about to generate, all combined.

If token limits (from Part 3 - Token Limits: Why Your AI Chatbot Suddenly Forgets Everything) are like the individual variable slots in a regression model, the context window is the entire spreadsheet the model can look at in one sitting. Once that spreadsheet is full, nothing new can be added without something old falling off the edge.

Context Window vs. Token Limit: What's the Difference?

These two terms get mixed up constantly, please find the difference below:

ConceptWhat It MeasuresAnalogy
Token LimitThe max tokens for one input or one output individuallyA single email's max attachment size
Context WindowThe combined total of prompt + conversation history + response the model can handle at onceThe total storage capacity of your email inbox

You could have a generous input limit but a small context window, or vice versa, though in most modern LLMs, the context window is really just the umbrella term for "everything the model can hold in memory during this interaction," and the input/output limits are how that total gets divided up.

Why Context Windows Have Grown So Fast

Early LLMs could only "remember" a few thousand tokens, roughly a few pages of text. Newer models can hold entire books' worth of context. This growth matters because a bigger context window means:

  • Longer conversations without the model "forgetting" earlier messages
  • The ability to paste in entire documents, spreadsheets, or codebases for analysis
  • Fewer awkward moments where you have to re-explain something you already said

Line chart showing fictional context window sizes growing across model versions This chart illustrates (using fictional version numbers, not any real product) how context window sizes have trended sharply upward over successive model generations, from a couple thousand tokens to hundreds of thousands.

Why Bigger isn't Automatically Better

A huge context window sounds great, but it comes with trade-offs worth knowing:

  • Cost: More tokens processed generally means higher compute cost, since the model has to attend to every token in the window.
  • Latency: Bigger context windows can mean slower responses, since there's simply more text to process.
  • "Needle in a haystack" risk: Even with a huge context window, models can sometimes struggle to give equal attention to something buried in the middle of a massive input versus something near the start or end.

So a smart AI product isn't necessarily the one with the biggest context window. it's the one that uses its context window wisely, feeding the model only what's actually relevant.

Connecting all the Dots Together: The Full Picture

Let's connect everything from all four parts into one clean mental model:

  1. Tokens (Part 1) are the basic chunks of text an LLM processes, words, sub-words, characters, or punctuation.
  2. Tokenization algorithms (Part 2) like BPE, WordPiece, SentencePiece, and Unigram decide exactly how text gets broken into those chunks.
  3. Token limits (Part 3) cap how many tokens can go into a single input or come out in a single output.
  4. Context windows (Part 4) are the total token budget across an entire interaction i.e. prompt, history, and response combined.

Think of it as building a house: tokens are the bricks, the tokenization algorithm is the mason deciding how to cut each brick, the token limit is how many bricks fit on one truck, and the context window is the total size of the building site where all the bricks get used.

Practical Example at "FreshCart"

Let's close the series with one fictional example that ties every concept together: a grocery delivery app called FreshCart, which uses an LLM-powered assistant to help customers plan weekly meals.

  • A customer's message ("I want a low-carb dinner plan for 5 nights, and I am allergic to peanuts") gets tokenized using FreshCart's chosen model's algorithm, turning the sentence into a mix of word and sub-word tokens.
  • FreshCart's assistant also needs to include the customer's order history and saved dietary preferences in the prompt, adding a few hundred more tokens.
  • All of this, the new message, the order history, and the preferences, must fit inside the model's input token limit.
  • As the conversation continues across many back-and-forth messages ("Can you swap Tuesday's recipe?", "Add more protein to Thursday"), the growing conversation has to fit inside the overall context window. Once FreshCart's chat history gets too long, the oldest messages are quietly dropped to make room for new ones, which is why, after a very long session, the assistant might "forget" that the customer mentioned a peanut allergy 50 messages ago unless FreshCart's engineers explicitly re-inject that detail into every new prompt.

That last point is a real design lesson: smart AI products often re-insert critical facts (like allergies) into every prompt rather than relying on the context window to remember them forever.

Key Takeaways (For the Whole Series)

  • A token is the basic building block of text an LLM understands not the same as a word or a character.
  • Tokenization algorithms (BPE, WordPiece, SentencePiece, Unigram) each decide differently how to chop text into tokens.
  • Token limits cap the size of a single input or output, just like a regression model has a fixed number of variable slots.
  • The context window is the total token budget for an entire interaction i.e. conversation history, prompt, and response combined.
  • Bigger context windows unlock longer, richer conversations but come with cost, latency, and attention trade-offs.
  • Well-designed AI products manage these limits deliberately, batching large inputs and re-injecting critical facts rather than assuming the model will "just remember."

This is one of the exhaustive 4 article series to explain. Thanks for following along through all 4 parts!

Next time the AI assistant seems to "run out of memory" mid-conversation, you will know exactly why and how it happened.

In case of any queries or feedback feel free to drop comments or reach out using the Contact secion.

Comments