40. Small vs Large Models
Choose the right model size based on task complexity, cost, and hardware constraints.
By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.
The lesson
Small models (7B-13B parameters) run well on consumer hardware and are excellent for: classification, simple extraction, formatting, and fast local tasks.
Medium models (20B-40B) offer better reasoning while remaining runnable on high-end consumer GPUs with quantisation.
Large models (70B+) provide deep reasoning but require significant VRAM (40GB+) unless heavily quantised.
Think of model size like resource allocation: match the model to the task. Use small fast models for simple work, large powerful models for complex reasoning.
Cloud frontier models (hundreds of billions of parameters) offer the best reasoning but at higher cost and latency.
A library cataloguer runs a small local model to classify incoming donations by subject, keeping a larger model for questions that need careful reasoning.
A hotel reservations supervisor uses a small model to extract booking details from enquiries and a larger one only for complex group-booking questions.
A recycling operations supervisor picks a small fast model to sort weighbridge entries into material categories, reserving a larger model for unusual contamination cases.
A pharmacy technician runs a small model to format repeat-script reminders, while a larger model is kept for questions that need careful checking against guidelines.
A drafter at an architecture practice uses a small model to standardise drawing notes and a large quantised model for checking conflicting specification clauses.
A club administrator at a sports association classifies membership forms with a small model and uses a larger one for planning questions about the coming season.
A workshop foreman at an automotive garage chooses a small local model for formatting job notes, because it runs quickly on the modest shop computer.
A support engineer at a software company classifies incoming tickets with a small model and escalates architecture questions to a larger model.
A livestock records clerk extracts animal tag numbers from handwritten notes with a small model, keeping a larger one for herd planning questions.
A fulfilment coordinator at an online store uses a small model to format order-exception lists and a larger one for questions about stock planning.
Check yourself
Question 1: When is a small model (7B-13B) the better choice?
- Deep philosophical reasoning
- Safety-critical code review
- Fast classification, simple extraction, or local deployment — correct
- Creative fiction writing
Answer: Fast classification, simple extraction, or local deployment
Small models are excellent for fast, simple tasks and local deployment where speed matters.
Question 2: What is the main trade-off when using a 70B model versus a 7B model locally?
- Price only
- VRAM requirements and inference speed versus reasoning depth — correct
- Color output
- File format support
Answer: VRAM requirements and inference speed versus reasoning depth
70B models need more VRAM and are slower but provide deeper reasoning than 7B models.
Question 3: Why might you use a cheaper cloud model for initial drafting?
- They are always more accurate
- To save cost before reviewing with a premium model — correct
- They have larger context
- They run offline
Answer: To save cost before reviewing with a premium model
A cost-effective workflow uses cheaper models for drafts and premium models for final review.
← Previous lesson · All 91 lessons · Next lesson →
The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.