3. Context Windows and Tokens
Budget prompt size, attachments, and long conversations.
By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.
The lesson
A token is a small chunk of text. Context window is the model's working memory for the current request.
Long context is useful, but it costs money, time, and memory. Put the highest-value facts first and remove old noise.
For coding, include file names, exact error text, acceptance criteria, and what must not change.
A paralegal splits a long lease into themed blocks, feeding the rent terms, the maintenance clauses and the termination conditions as separate context rather than pasting the whole document.
A hospital clinical coder supplies only the discharge summary and the diagnosis list, removing pages of unrelated observation charts so the request stays focused on the coding question.
A freight forwarding coordinator gives the model the shipment reference, the carrier conditions and the current delay note, leaving older email chains out so the context budget is preserved.
A laboratory scientist trims a procedure review to the method section and the instrument settings, discarding boilerplate that would consume space without adding information.
A software developer pastes only the failing function, the exact error text and the acceptance criteria, avoiding a whole-project dump that would exhaust the context window.
A production engineer chunks a machine manual by inspection type, sending only the lubrication and calibration sections needed for the current question.
A course lecturer summarises a topic in three short blocks, keeping the key facts at the top of the prompt and deleting earlier drafts that no longer matter.
A credit analyst feeds the statement lines that matter plus a short account history, trimming commentary so the model reads the numbers first.
A property manager condenses a building file into a summary plus exact excerpts from the current inspection report instead of attaching the entire history.
A sports journalist gives the model the match summary and two or three direct quotes, keeping the prompt short enough to remain inside working memory.
Check yourself
Question 1: What is a context window?
- A GPU fan setting
- The model working memory for the current request — correct
- A Windows desktop panel
- A license file
Answer: The model working memory for the current request
The context window is the amount of prompt and conversation the model can consider.
Question 2: Why can very long prompts be expensive or slow?
- They use more input tokens and processing — correct
- They make the keyboard slower
- They disable Markdown
- They shrink the model
Answer: They use more input tokens and processing
More tokens usually mean more cost, latency, and attention burden.
Question 3: For coding help, which context is most valuable?
- Only the word fix
- A random screenshot without explanation
- No project details
- Exact file paths, errors, constraints, and acceptance criteria — correct
Answer: Exact file paths, errors, constraints, and acceptance criteria
Specific technical context helps the model solve the actual problem.
← Previous lesson · All 91 lessons · Next lesson →
The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.