3. Context Windows and Tokens

Budget prompt size, attachments, and long conversations.

By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.

The lesson

A token is a small chunk of text. Context window is the model's working memory for the current request.

Long context is useful, but it costs money, time, and memory. Put the highest-value facts first and remove old noise.

For coding, include file names, exact error text, acceptance criteria, and what must not change.

A paralegal splits a long lease into themed blocks, feeding the rent terms, the maintenance clauses and the termination conditions as separate context rather than pasting the whole document.

A hospital clinical coder supplies only the discharge summary and the diagnosis list, removing pages of unrelated observation charts so the request stays focused on the coding question.

A freight forwarding coordinator gives the model the shipment reference, the carrier conditions and the current delay note, leaving older email chains out so the context budget is preserved.

A laboratory scientist trims a procedure review to the method section and the instrument settings, discarding boilerplate that would consume space without adding information.

A software developer pastes only the failing function, the exact error text and the acceptance criteria, avoiding a whole-project dump that would exhaust the context window.

A production engineer chunks a machine manual by inspection type, sending only the lubrication and calibration sections needed for the current question.

A course lecturer summarises a topic in three short blocks, keeping the key facts at the top of the prompt and deleting earlier drafts that no longer matter.

A credit analyst feeds the statement lines that matter plus a short account history, trimming commentary so the model reads the numbers first.

A property manager condenses a building file into a summary plus exact excerpts from the current inspection report instead of attaching the entire history.

A sports journalist gives the model the match summary and two or three direct quotes, keeping the prompt short enough to remain inside working memory.

Check yourself

Question 1: What is a context window?
  1. A GPU fan setting
  2. The model working memory for the current request — correct
  3. A Windows desktop panel
  4. A license file

Answer: The model working memory for the current request

The context window is the amount of prompt and conversation the model can consider.

Question 2: Why can very long prompts be expensive or slow?
  1. They use more input tokens and processing — correct
  2. They make the keyboard slower
  3. They disable Markdown
  4. They shrink the model

Answer: They use more input tokens and processing

More tokens usually mean more cost, latency, and attention burden.

Question 3: For coding help, which context is most valuable?
  1. Only the word fix
  2. A random screenshot without explanation
  3. No project details
  4. Exact file paths, errors, constraints, and acceptance criteria — correct

Answer: Exact file paths, errors, constraints, and acceptance criteria

Specific technical context helps the model solve the actual problem.

← Previous lesson · All 91 lessons · Next lesson →

The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.