46. Local Model Strategy

Choose the right local models (Llama, Qwen, DeepSeek, Mistral) for your hardware and needs.

By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.

The lesson

Llama 4 Scout (109B active) fits in 24GB VRAM when quantised to 4-bit, making it accessible for high-end consumer hardware.

Llama 4 Maverick (400B) offers deeper reasoning but requires significantly more VRAM or heavy quantisation.

DeepSeek V4 Pro offers 1M context and strong coding/reasoning at dramatically lower cost than Western frontier models.

Qwen 3.7 Max (June 2026) is the new agentic coding flagship with 1M context. Qwen 3.6 models remain excellent for multilingual tasks.

Mistral Large 3 (May 2026) is Apache 2.0 open-weight with 256K context - a strong Western open-weight option.

Match model size to your VRAM: under 12GB use 7B-13B models; 24GB handles 35B-70B quantised; 40GB+ for larger models.

A junior IT support officer at a small clinic tests a 7B local model on an 8 GB laptop to draft simple ticket replies, keeping patient-adjacent text on the machine.

An engineering lead at a manufacturing plant writes a model comparison note in TPEE before choosing a quantised size that fits a workstation GPU for maintenance-log summaries.

A veterinary practice manager drafts prompts in TPEE for a local model that turns consultation notes into follow-up reminders, copying the text across by hand.

A library technician runs a small local model on a public access computer to draft catalogue descriptions and records in TPEE which model size suits the older hardware.

A mining safety specialist keeps a model comparison note in TPEE, matching a mid-size local model to the site office computer because the network is unreliable.

The owner of a family bakery chooses a small local model for routine recipe-cost summaries, noting in TPEE that it runs comfortably on an ordinary laptop.

A legal clerk at a small practice keeps confidential case notes on a local model and uses TPEE to record which model size fits the office machines.

A logistics planner at a freight depot drafts dispatch-note prompts in TPEE for an offline model, because the handheld terminals have limited memory.

A secondary school teacher documents a local model choice in TPEE for classroom material drafts, picking a size that runs on the school laboratory computers.

A cybersecurity analyst at a software company writes a TPEE note comparing local model sizes for reviewing log excerpts without sending data offsite.

Check yourself

Question 1: When should you prefer Llama 4 Scout over Llama 4 Maverick?
  1. When you need maximum accuracy
  2. When you have limited VRAM (4-bit quantised 17B fits in 24GB) — correct
  3. Only for image tasks
  4. Never

Answer: When you have limited VRAM (4-bit quantised 17B fits in 24GB)

Llama 4 Scout is smaller and fits better on consumer hardware, while Maverick needs more VRAM.

Question 2: What is DeepSeek V4 Pro best suited for locally?
  1. Real-time voice chat
  2. Reasoning, coding, and math with 1M context at lower cost — correct
  3. 8K video editing
  4. Spreadsheet macros only

Answer: Reasoning, coding, and math with 1M context at lower cost

DeepSeek V4 Pro offers strong reasoning and 1M context locally at dramatically lower cost than Western frontier models.

Question 3: Why might you choose Qwen 3.7 Max or 3.6 for a multilingual local application?
  1. It only supports English
  2. Strong multilingual capabilities in a runnable local size — correct
  3. It is the smallest model available
  4. It requires no GPU

Answer: Strong multilingual capabilities in a runnable local size

Qwen models (3.7 Max for coding, 3.6 for multilingual) excel at multilingual tasks while remaining runnable on consumer hardware.

← Previous lesson · All 91 lessons · Next lesson →

The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.