21. Quantisation and Model Size
Choose between FP16, INT8, INT4, quality, speed, and memory use.
By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.
The lesson
Quantisation stores model weights with fewer bits. This lowers memory use and can increase speed, but may reduce quality.
A 7B or 13B quantised model can be very useful locally. Larger models may need more VRAM or careful offloading.
The right choice depends on task risk: drafting and summaries tolerate smaller models; deep reasoning and complex coding often need stronger models.
A maintenance planner at a textile mill picks a low-bit model for summarising machine logs because it fits the office computer with memory to spare.
A junior support analyst at an insurance firm records that an eight-bit model gave clearer answers than a four-bit one on policy wording questions.
A team lead at a pharmaceutical warehouse documents which quantised model fits the available graphics memory on the stockroom computer.
A firmware tester at an electronics workshop chooses a smaller quantised model for routine code comments and a larger one for difficult fault analysis.
A veterinary practice manager notes that a four-bit model runs quickly on the clinic laptop but occasionally muddles the wording of dosage instructions.
A reliability engineer at a mining operator compares an eight-bit and a four-bit model on equipment failure notes before deciding which one to keep.
A junior archivist at a museum records that a small quantised model handles catalogue descriptions well enough for a first draft.
A construction site supervisor notes that a heavily compressed model answers faster on a rugged tablet but loses detail on long specifications.
A customer service team leader at an online retailer documents which model size suits routine reply drafts against complex complaint histories.
A tools programmer at a game studio records the memory saved by a four-bit model when running local checks on a developer workstation.
Check yourself
Question 1: What does quantisation reduce?
- Number of keyboard keys
- Monitor size
- All need for review
- Bits used to store model weights — correct
Answer: Bits used to store model weights
Lower precision reduces memory use and may improve speed.
Question 2: What is the trade-off of aggressive quantisation?
- More VRAM required always
- No local support
- Possible quality loss — correct
- Guaranteed perfect reasoning
Answer: Possible quality loss
Compression can reduce accuracy or reasoning quality.
Question 3: When might a stronger or less-quantised model be needed?
- Opening the About dialog
- Deep reasoning or complex coding — correct
- Changing font color
- Saving a file name
Answer: Deep reasoning or complex coding
Higher-risk tasks may need better model quality.
← Previous lesson · All 91 lessons · Next lesson →
The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.