Reading · 6 min · Lesson 2 of 8
Tokens and the context window
Why AI forgets the start of a long conversation and how to avoid it.
Models do not read words the way we do. They split text into tokens, small pieces that are often only part of a word. In English one token is roughly three quarters of a word, while Croatian, with its rich word endings, usually needs more tokens for the same message.
Everything the model can take into account at once, meaning your instructions, pasted documents, earlier messages and its own replies, has to fit into the context window. Once a conversation outgrows it, the oldest parts drop out or get compressed, and the model starts to lose details you gave it at the start.
Practical rule: for a new task, start a new conversation, and restate the key facts when a thread gets long. Paste only the part of a document that matters instead of the whole file. Shorter, focused context gives more accurate answers and costs less when you use paid APIs.
Key takeaways
- Text is processed as tokens, and Croatian usually needs more of them than English.
- The context window limits how much the model can consider at once.
- Start fresh for new tasks and paste only what matters.