21. Quantisation and Model Size

Choose between FP16, INT8, INT4, quality, speed, and memory use.

By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.

The lesson

Quantisation stores model weights with fewer bits. This lowers memory use and can increase speed, but may reduce quality.

A 7B or 13B quantised model can be very useful locally. Larger models may need more VRAM or careful offloading.

The right choice depends on task risk: drafting and summaries tolerate smaller models; deep reasoning and complex coding often need stronger models.

A maintenance planner at a textile mill picks a low-bit model for summarising machine logs because it fits the office computer with memory to spare.

A junior support analyst at an insurance firm records that an eight-bit model gave clearer answers than a four-bit one on policy wording questions.

A team lead at a pharmaceutical warehouse documents which quantised model fits the available graphics memory on the stockroom computer.

A firmware tester at an electronics workshop chooses a smaller quantised model for routine code comments and a larger one for difficult fault analysis.

A veterinary practice manager notes that a four-bit model runs quickly on the clinic laptop but occasionally muddles the wording of dosage instructions.

A reliability engineer at a mining operator compares an eight-bit and a four-bit model on equipment failure notes before deciding which one to keep.

A junior archivist at a museum records that a small quantised model handles catalogue descriptions well enough for a first draft.

A construction site supervisor notes that a heavily compressed model answers faster on a rugged tablet but loses detail on long specifications.

A customer service team leader at an online retailer documents which model size suits routine reply drafts against complex complaint histories.

A tools programmer at a game studio records the memory saved by a four-bit model when running local checks on a developer workstation.

Check yourself

Question 1: What does quantisation reduce?
  1. Number of keyboard keys
  2. Monitor size
  3. All need for review
  4. Bits used to store model weights — correct

Answer: Bits used to store model weights

Lower precision reduces memory use and may improve speed.

Question 2: What is the trade-off of aggressive quantisation?
  1. More VRAM required always
  2. No local support
  3. Possible quality loss — correct
  4. Guaranteed perfect reasoning

Answer: Possible quality loss

Compression can reduce accuracy or reasoning quality.

Question 3: When might a stronger or less-quantised model be needed?
  1. Opening the About dialog
  2. Deep reasoning or complex coding — correct
  3. Changing font color
  4. Saving a file name

Answer: Deep reasoning or complex coding

Higher-risk tasks may need better model quality.

← Previous lesson · All 91 lessons · Next lesson →

The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.