In plain English
It is working memory, measured in tokens. Everything the model can take into account has to fit inside it, and when it fills, the earliest content falls out of view.
Larger windows are genuinely useful, but attention is not uniform across a long context. Material at the start and end is used more reliably than material buried in the middle.
What to know
Why it matters
Most complaints that a model forgot something are context window problems. Knowing the limit changes how you structure long work: summarise and restate rather than assuming a fifty-message thread is all still in view.
Common mistakes
FAQs
Is a bigger context window always better?
It helps, but relevance beats volume. Retrieval usually outperforms pasting everything.
What happens when I exceed it?
Earlier content is dropped or the request fails. Either way, quietly.
I'm a marketing consultant, entrepreneur and content creator. I help businesses grow through practical marketing, websites, SEO, content and AI.
More About Tariq →